Google
Operational· 100% uptime, 7 days

Gemini 3.1 Flash-Lite API

The Gemini 3.1 Flash-Lite API for high-volume work, 61% below Google's list price

Price per 1M tokens61%off
Input
$0.10$0.25
Output
$0.59$1.50
Cached input
$0.01
Model IDgemini-3.1-flash-lite
chat.py
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key="YOUR_ZURELAY_KEY",
)
stream = client.chat.completions.create(
model="gemini-3.1-flash-lite",
messages=[{"role": "user", "content": "Explain quicksort in two sentences."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")

Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.

Context window
1M tokens
Max output
66K tokens
Input
Text, Image, Video, Audio
Output
Text
Released
Mar 3, 2026
Uptime, 7 days
100%

Pricing

Gemini 3.1 Flash-Lite API pricing. A fraction of the list price.

Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.

RatezurelayGoogle listYou save
Input
per 1M tokens
$0.10$0.2560%off
Output
per 1M tokens
$0.59$1.5061%off
Cached input
per 1M tokens
$0.01——

One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.

Savings calculator

Move the sliders to your monthly usage.

Input tokens per month50M
Output tokens per month10M
zurelay
$10.90
Google list price
$27.50
You save every month$16.60

$199.20 a year

Overview

What is Gemini 3.1 Flash-Lite?

Gemini 3.1 Flash-Lite is Google's low-latency, low-cost Gemini 3 model for high-volume work such as translation, moderation, extraction and model routing. The zurelay Gemini 3.1 Flash-Lite API serves it through one OpenAI-compatible endpoint with a 1M-token context window, at 61% below Google's list price.

Google introduced Gemini 3.1 Flash-Lite in preview on March 3, 2026 as the fastest and most cost-efficient model in the Gemini 3 series, and made it generally available on May 7, 2026. The Gemini 3.1 Flash-Lite API model reads text, images, video, audio and PDFs and writes text, with up to 1,048,576 input tokens and 65,536 output tokens. On zurelay the model ID is gemini-3.1-flash-lite, and it costs $0.10 per 1M input tokens and $0.59 per 1M output tokens.

At launch Google reported an Elo of 1432 on the Arena.ai leaderboard, 86.9% on GPQA Diamond and 76.8% on MMMU Pro. Against Gemini 2.5 Flash, Google reports a 2.5x faster time to first answer token and 45% higher output speed. The model supports the minimal, low, medium and high thinking levels, with minimal as the default, plus function calling and structured outputs.

Google shut down the preview ID, gemini-3.1-flash-lite-preview, on May 25, 2026. It has set May 7, 2027 as the shutdown date for the stable model on its API and names Gemini 3.5 Flash-Lite as the replacement. Until then, 3.1 Flash-Lite has a lower list price than 3.5 Flash-Lite, which makes it a good fit for simple, very high-volume jobs.

Strengths

Where Gemini 3.1 Flash-Lite shines. And what teams build with it.

01

Low cost per call

Its list price is below Gemini 3.5 Flash-Lite's, and zurelay charges 61% less than that list price.

02

Fast first token

Google reports a 2.5x faster time to first answer token than Gemini 2.5 Flash and 45% higher output speed.

03

Thinking when you need it

Defaults to minimal thinking and scales to low, medium or high for tasks like UI generation, simulations or complex instructions.

04

Long, multimodal context

Takes up to 1,048,576 tokens of text, images, video, audio or PDF in one request.

Use cases

  • Translation and transcription

    High-volume translation of text and transcription of audio, two of the core uses Google lists for the model.

  • Content moderation

    Screen large volumes of user content where cost per call is the main constraint.

  • Data extraction

    Pull fields from documents, emails and forms into structured output at low cost.

  • Model routing

    Classify incoming requests and send each one to the right model or tool.

Get started

Call Gemini 3.1 Flash-Lite in three steps.

No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.

  1. 1

    Create a key

    Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.

  2. 2

    Point your SDK at zurelay

    Change the base URL. Everything else in your code stays the same.

    https://api.zurelay.com/v1
  3. 3

    Use gemini-3.1-flash-lite

    Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.

    gemini-3.1-flash-lite

FAQ

Gemini 3.1 Flash-Lite API questions.

Anything else, write to sales@zurelay.com and an engineer will answer.

How much does the Gemini 3.1 Flash-Lite API cost?

On zurelay, gemini-3.1-flash-lite costs $0.10 per 1M input tokens and $0.59 per 1M output tokens, with cached input at $0.01. Google's list price is $0.25 input and $1.50 output per 1M, so you save 61%. Thinking tokens bill as output.

Is there a cheaper Gemini 3.1 Flash-Lite API than Google's?

Yes. zurelay serves gemini-3.1-flash-lite at 61% below Google's list price, with one rate at every prompt length. You pay from prepaid credit, with no subscription.

Is Gemini 3.1 Flash-Lite still in preview?

No. Google made gemini-3.1-flash-lite generally available on May 7, 2026 and shut down the gemini-3.1-flash-lite-preview ID on May 25, 2026. zurelay also accepts gemini-3.1-flash-lite-preview as a model name, so older code keeps working.

How do I call the Gemini 3.1 Flash-Lite API with the OpenAI SDK?

Install the official openai package, set base_url to https://api.zurelay.com/v1 and use your zurelay API key. Then send a Chat Completions request with model set to "gemini-3.1-flash-lite". Messages, tools and streaming use the standard OpenAI request format.

What is the Gemini 3.1 Flash-Lite context window?

Gemini 3.1 Flash-Lite takes up to 1,048,576 input tokens and writes up to 65,536 output tokens per response. The model accepts text, image, video, audio and PDF input and writes text.

When will Google shut down Gemini 3.1 Flash-Lite?

Google lists May 7, 2027 as the shutdown date for gemini-3.1-flash-lite on its API and recommends gemini-3.5-flash-lite as the replacement. Both models work with the same zurelay key, so switching is a one-line change.

What is Gemini 3.1 Flash-Lite best at?

High-volume, latency-sensitive work where cost per call matters most: translation, transcription, moderation, data extraction and model routing. Raise the thinking level for heavier tasks such as generating user interfaces or following complex instructions.

Does the Gemini 3.1 Flash-Lite API support streaming and rate limits?

Yes. Set stream: true, and add stream_options.include_usage to get token counts in the final chunk. Each zurelay API key can have its own monthly budget and requests-per-minute cap. Failed requests are retried on another route and never billed.

Is it the same model as Google's Gemini 3.1 Flash-Lite?

Yes, it is Google's Gemini 3.1 Flash-Lite model. Responses are sampled, so wording varies between runs, and zurelay does not promise byte-identical output. zurelay is independent and is not affiliated with or endorsed by Google.

More models on zurelay

All models

Stop paying list price.
Start saving today.

Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.