Built for throughput
Measured at 350 output tokens per second by Artificial Analysis, according to Google, the fastest model in the 3.5 series.
The Gemini 3.5 Flash-Lite API: Google's fastest 3.5 model, 58% below list price
gemini-3.5-flash-litefrom openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key="YOUR_ZURELAY_KEY",)stream = client.chat.completions.create( model="gemini-3.5-flash-lite", messages=[{"role": "user", "content": "Explain quicksort in two sentences."}], stream=True,)for chunk in stream: if chunk.choices: print(chunk.choices[0].delta.content or "", end="")Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.
Pricing
Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.
| Rate | zurelay | Google list | You save |
|---|---|---|---|
Input per 1M tokens | $0.12 | 60%off | |
Output per 1M tokens | $1.04 | 58%off | |
Cached input per 1M tokens | $0.012 | — | — |
One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.
Move the sliders to your monthly usage.
$283.20 a year
Overview
Gemini 3.5 Flash-Lite is the fastest and most cost-effective model in Google's Gemini 3.5 series, released on July 21, 2026. The zurelay Gemini 3.5 Flash-Lite API serves it through one OpenAI-compatible endpoint for high-volume agent steps, search and document processing, at 58% below Google's list price.
Google released Gemini 3.5 Flash-Lite on July 21, 2026, alongside Gemini 3.6 Flash. Google describes the Gemini 3.5 Flash-Lite API model as low-latency and cost-effective, built for high-throughput subagent tasks and document parsing. It reads text, images, video, audio and PDFs and writes text, with up to 1,048,576 input tokens and 65,536 output tokens. On zurelay the model ID is gemini-3.5-flash-lite, and it costs $0.12 per 1M input tokens and $1.04 per 1M output tokens.
Speed is the point. Google cites Artificial Analysis measuring it at 350 output tokens per second, the fastest model in the 3.5 series. It is also a large step up from 3.1 Flash-Lite: 54% against 31% on Terminal-Bench 2.1, 72.2% against 60.1% on the GDM-MRCR v2 long-context test, and 1140 against 642 on GDPval-AA v2. Google reports it ahead of Gemini 3 Flash on SWE-Bench Pro (54.2% against 49.6%) and OSWorld-Verified (74.0% against 65.1%).
Thinking runs at the minimal level by default, which keeps latency low. Raise it to low, medium or high when a step needs multi-step reasoning, on the same model ID. The model supports function calling and structured outputs, so it works well as a fast worker inside a larger agent. Google lists it as a stable model with no shutdown date announced, and names it as the replacement for Gemini 3.1 Flash-Lite.
Strengths
Measured at 350 output tokens per second by Artificial Analysis, according to Google, the fastest model in the 3.5 series.
Defaults to minimal thinking for fast answers. Raise it to low, medium or high for harder steps without switching models.
Google reports 54% against 31% on Terminal-Bench 2.1 and 72.2% against 60.1% on the GDM-MRCR v2 long-context test.
Takes up to 1,048,576 tokens of text, images, video, audio or PDF in one request.
Use cases
Hand search, lookup and tool-calling steps from a larger planner model to a fast, low-cost worker.
Parse PDFs and long documents and extract fields into structured output.
Run many quick search-and-summarize loops, where per-step latency adds up.
Translate, clean and reformat large volumes of text with thinking left at minimal.
Get started
No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.
Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.
Change the base URL. Everything else in your code stays the same.
https://api.zurelay.com/v1Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.
gemini-3.5-flash-liteWorks with the tools you already use
Compare
$0.32 / $1.59 per 1M tokens
Pick Gemini 3.7 Flash for harder coding, debugging and multi-step agent work where quality matters more than speed.
Gemini 3.7 Flash API$0.10 / $0.59 per 1M tokens
Pick Gemini 3.1 Flash-Lite for simple, high-volume jobs where its lower list price counts for more than the quality gains of 3.5.
Gemini 3.1 Flash-Lite API$0.13 / $0.51 per 1M tokens
Pick MiniMax M2.7 for text-only coding and agent work on an open-weight model with a lower list output price.
MiniMax M2.7 APIFAQ
Anything else, write to sales@zurelay.com and an engineer will answer.
On zurelay, gemini-3.5-flash-lite costs $0.12 per 1M input tokens and $1.04 per 1M output tokens, with cached input at $0.012. Google's list price is $0.30 input and $2.50 output per 1M, so you save 58%. Thinking tokens bill as output.
Yes. zurelay serves gemini-3.5-flash-lite at 58% below Google's list price, with one rate at every prompt length. You pay from prepaid credit, with no subscription.
Install the official openai package, set base_url to https://api.zurelay.com/v1 and use your zurelay API key. Then send a Chat Completions request with model set to "gemini-3.5-flash-lite". Messages, tools and streaming use the standard OpenAI request format.
Gemini 3.5 Flash-Lite takes up to 1,048,576 input tokens and writes up to 65,536 output tokens per response. The model accepts text, image, video, audio and PDF input and writes text.
Google cites Artificial Analysis measuring it at 350 output tokens per second. Speed through any API also depends on load, prompt length and thinking level. The default minimal thinking level gives the fastest responses.
It scores much higher in Google's tests: 54% against 31% on Terminal-Bench 2.1, 72.2% against 60.1% on GDM-MRCR v2 and 1140 against 642 on GDPval-AA v2. Its list price is higher, so 3.1 Flash-Lite can still make sense for simple, very high-volume jobs.
Fast, high-volume work: subagent steps, agentic search, document processing and translation. With thinking raised, it also handles coding and multi-step agent tasks, where Google reports it ahead of Gemini 3 Flash on SWE-Bench Pro and OSWorld-Verified.
Yes. Set stream: true, and add stream_options.include_usage to get token counts in the final chunk. Each zurelay API key can have its own monthly budget and requests-per-minute cap. Failed requests are retried on another route and never billed.
Yes, it is Google's Gemini 3.5 Flash-Lite model. Responses are sampled, so wording varies between runs, and zurelay does not promise byte-identical output. zurelay is independent and is not affiliated with or endorsed by Google.
Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.