Google
Operational· 100% uptime, 7 days

Gemini 3.7 Flash API

The Gemini 3.7 Flash API for coding and agents, 58% below Google's list price

Price per 1M tokens58%off
Input
$0.32$0.75
Output
$1.59$3.75
Cached input
$0.032
Model IDgemini-3.7-flash
chat.py
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key="YOUR_ZURELAY_KEY",
)
stream = client.chat.completions.create(
model="gemini-3.7-flash",
messages=[{"role": "user", "content": "Explain quicksort in two sentences."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")

Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.

Context window
1M tokens
Max output
66K tokens
Input
Text, Image, Video, Audio
Output
Text
Released
Aug 13, 2026
Uptime, 7 days
100%

Pricing

Gemini 3.7 Flash API pricing. A fraction of the list price.

Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.

RatezurelayGoogle listYou save
Input
per 1M tokens
$0.32$0.7557%off
Output
per 1M tokens
$1.59$3.7558%off
Cached input
per 1M tokens
$0.032——

One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.

Savings calculator

Move the sliders to your monthly usage.

Input tokens per month50M
Output tokens per month10M
zurelay
$31.90
Google list price
$75.00
You save every month$43.10

$517.20 a year

Overview

What is Gemini 3.7 Flash?

Gemini 3.7 Flash is Google's Flash model for coding, agentic tool use and multi-step work, released on August 13, 2026. The zurelay Gemini 3.7 Flash API serves it through one OpenAI-compatible endpoint with a 1M-token context window, at 58% below Google's list price.

Google released Gemini 3.7 Flash on August 13, 2026, three weeks after Gemini 3.6 Flash, and called it its most intelligent workhorse model yet for coding and agents. The Gemini 3.7 Flash API model reads text, images, video, audio and PDFs and writes text. It takes up to 1,048,576 input tokens and writes up to 65,536 output tokens per response. On zurelay the model ID is gemini-3.7-flash, and it costs $0.32 per 1M input tokens and $1.59 per 1M output tokens.

Google's numbers show a clear step up from 3.6 Flash. It scores 65.3% on DeepSWE v1.1 against 49.0%, and 43.6% on FrontierCode 1.1 Main against 34.4%. On Arena.ai's WebDev Arena it has an Elo of 1588 against 1538. For document-heavy knowledge work it scores 34.0% on GDP.pdf against 22.0%, and 30.4% on AutomationBench against 17.0%. Google also says it adapts better to roadblocks, asks for clarification when needed and follows instructions more closely, which means fewer retries in agent loops.

Thinking is always on. Gemini 3.7 Flash supports the low, medium and high thinking levels, with medium as the default. The minimal level is not supported and returns an error. The model also supports function calling and structured outputs. Thinking tokens bill as output, so a lower level cuts both cost and latency on simple steps.

Google's list price for 3.7 Flash is an introductory rate that runs through December 31, 2026. Google doubles it on January 1, 2027. Gemini 3.8 Flash followed on September 2, 2026, and Google says 3.7 Flash remains fully supported for efficiency-first workloads. Google has not announced a shutdown date for it.

Strengths

Where Gemini 3.7 Flash shines. And what teams build with it.

01

Stronger coding than 3.6 Flash

Google reports 65.3% on DeepSWE v1.1 and 43.6% on FrontierCode 1.1 Main, up from 49.0% and 34.4% for 3.6 Flash.

02

Steadier agent runs

Google says it adapts to roadblocks and follows instructions more closely. Multi-step tool loops need fewer retries and less oversight.

03

Reads long, complex documents

A 1M-token context with PDF, image, video and audio input. It scores 34.0% on GDP.pdf, a complex-document benchmark, against 22.0% for 3.6 Flash.

04

Front-end and UI generation

Builds more functional layouts in fewer prompts and follows a reference screenshot or design system closely. Its WebDev Arena Elo is 1588.

Use cases

  • Coding agents

    Debugging, issue resolution and code generation in IDE and CLI agents, where better first-pass accuracy saves a round trip.

  • Web and UI generation

    Turn a screenshot, mockup or design system into working front-end code.

  • Document-heavy knowledge work

    Read long contracts, filings and research papers in fields like finance, law and biosciences, then extract, summarize or answer questions.

  • Business workflow automation

    Run multi-step, tool-calling workflows with structured outputs, such as ticket triage or moving data between systems.

Get started

Call Gemini 3.7 Flash in three steps.

No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.

  1. 1

    Create a key

    Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.

  2. 2

    Point your SDK at zurelay

    Change the base URL. Everything else in your code stays the same.

    https://api.zurelay.com/v1
  3. 3

    Use gemini-3.7-flash

    Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.

    gemini-3.7-flash

FAQ

Gemini 3.7 Flash API questions.

Anything else, write to sales@zurelay.com and an engineer will answer.

How much does the Gemini 3.7 Flash API cost?

On zurelay, gemini-3.7-flash costs $0.32 per 1M input tokens and $1.59 per 1M output tokens, with cached input at $0.032. Google's list price is $0.75 input and $3.75 output per 1M, so you save 58%. Thinking tokens bill as output, and the rate is the same at every prompt length.

Is there a cheaper Gemini 3.7 Flash API than Google's?

Yes. zurelay serves gemini-3.7-flash at 58% below Google's list price, paid from prepaid credit with no subscription. Google's own price is an introductory rate that runs through December 31, 2026, and Google doubles it on January 1, 2027.

How do I call the Gemini 3.7 Flash API with the OpenAI SDK?

Install the official openai package, set base_url to https://api.zurelay.com/v1 and use your zurelay API key. Then send a Chat Completions request with model set to "gemini-3.7-flash". Messages, tools and streaming use the standard OpenAI request format.

What is the Gemini 3.7 Flash context window?

Gemini 3.7 Flash takes up to 1,048,576 input tokens and writes up to 65,536 output tokens per response. The model accepts text, image, video, audio and PDF input and writes text.

How do Gemini 3.7 Flash thinking levels work?

Thinking is always on. The model supports low, medium and high, with medium as the default, and returns an error for minimal. Google's OpenAI-compatible API maps the reasoning_effort field to Gemini 3 thinking levels. Use low for quick tool steps and high for hard debugging.

What is Gemini 3.7 Flash best at?

Coding and agent work: debugging, issue resolution, front-end generation and multi-step tool use. It is also strong on long, complex documents, where Google reports clear gains over 3.6 Flash on GDP.pdf and AutomationBench.

Is Gemini 3.7 Flash still supported now that 3.8 Flash is out?

Yes. Google released Gemini 3.8 Flash on September 2, 2026 and says 3.7 Flash remains fully supported for efficiency-first workloads. Google has not announced a shutdown date for gemini-3.7-flash.

Does the Gemini 3.7 Flash API support streaming and rate limits?

Yes. Set stream: true, and add stream_options.include_usage to get token counts in the final chunk. Each zurelay API key can have its own monthly budget and requests-per-minute cap. Failed requests are retried on another route and never billed.

Is it the same model as Google's Gemini 3.7 Flash?

Yes, it is Google's Gemini 3.7 Flash model. Responses are sampled, so wording varies between runs, and zurelay does not promise byte-identical output. zurelay is independent and is not affiliated with or endorsed by Google.

More models on zurelay

All models

Stop paying list price.
Start saving today.

Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.