Google
Operational· checked every few minutes

Gemini 3.8 Flash API

The Gemini 3.8 Flash API for long-horizon coding and agents, 58% below Google's list price

Price per 1M tokens58%off
Input
$0.32$0.75
Output
$1.59$3.75
Cached input
$0.032
Model IDgemini-3.8-flash
chat.py
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key="YOUR_ZURELAY_KEY",
)
stream = client.chat.completions.create(
model="gemini-3.8-flash",
messages=[{"role": "user", "content": "Explain quicksort in two sentences."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")

Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK. Chat API docs

Context window
1M tokens
Max output
64K tokens
Reasoning
low · medium · high
Input
Text, Image, Video, Audio
Released
Sep 2, 2026
Status
Operational

Pricing

Gemini 3.8 Flash API pricing. A fraction of the list price.

Pay as you go from prepaid credit, with no subscription. Top up from $10. Requests that fail are never billed.

RateZurelayGoogle list
Input
per 1M tokens
$0.32$0.75
Output
per 1M tokens
$1.59$3.75
Cached input
per 1M tokens
$0.032—

One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.

Savings calculator

Move the sliders to your monthly usage.

Input tokens per month50M
Output tokens per month10M
Zurelay
$31.90
Google list price
$75.00
You save every month$43.10

$517.20 a year

Overview

What is Gemini 3.8 Flash?

Gemini 3.8 Flash is Google's most intelligent Flash model, built for long-horizon software engineering, autonomous agents and complex enterprise workflows, and released on September 2, 2026. The Zurelay Gemini 3.8 Flash API serves it through one OpenAI-compatible endpoint with a 1M-token context window, at 58% below Google's list price.

Google released Gemini 3.8 Flash on September 2, 2026, three weeks after Gemini 3.7 Flash and its third Flash release in six weeks. Google calls it its most intelligent workhorse model, with significant improvements over 3.7 Flash in software engineering, agentic tasks and multi-step reasoning in specialized fields. The Gemini 3.8 Flash API model reads text, images, video, audio and PDFs and writes text. It takes up to 1,048,576 input tokens and writes up to 65,536 output tokens per response. On Zurelay the model ID is gemini-3.8-flash, and it costs $0.32 per 1M input tokens and $1.59 per 1M output tokens.

Google's model card shows where the gains land. It scores 73.7% on DeepSWE v1.1, a long-horizon software engineering test, against 65.3% for 3.7 Flash and 74.0% for Claude Opus 5. Among the six models in Google's table it has the top score on Terminal-bench 2.1 (89.4%), HLE-Verified (54.9%), Vals Finance Agent v2 (61.4%) and Harvey's Legal Agent Benchmark (10.0%). Claude Opus 5 stays well ahead on Terminal-bench 4.0, a test of general agent skills, and on OSWorld-2.0 for computer use.

Thinking is always on. Gemini 3.8 Flash supports the low, medium and high thinking levels, with medium as the default; the minimal level is not supported. The model also supports function calling and structured outputs. Google says 3.8 Flash works harder on complex tasks, taking extra reasoning steps and calling tools iteratively. Thinking tokens bill as output, so a lower level cuts both cost and latency on simple steps.

Google lists 3.8 Flash at the same introductory price as 3.7 Flash, a rate that runs through December 31, 2026. Google doubles it on January 1, 2027. Where compute efficiency matters most, Google suggests a lower thinking level, or Gemini 3.7 Flash, which it says remains fully supported for efficiency-first workloads. Google has not announced a shutdown date for either model.

Strengths

Where Gemini 3.8 Flash shines. And what teams build with it.

01

Long-horizon software engineering

Google reports 73.7% on DeepSWE v1.1, up from 65.3% for 3.7 Flash and close to Claude Opus 5 at 74.0%.

02

Finance and legal agents

It posts the top scores in Google's six-model table on Vals Finance Agent v2 (61.4%) and Harvey's Legal Agent Benchmark (10.0%).

03

Terminal and computer use

89.4% on Terminal-bench 2.1, the best in Google's table, and 59.0% on OSWorld-2.0, up from 50.6% for 3.7 Flash.

04

Charts, PDFs and long video

A 1M-token context with PDF, image, video and audio input. It scores 86.2% on CharXiv Reasoning and 87.8% on LVBench, a long-video test.

Use cases

  • Coding agents

    Long-horizon work in IDE and CLI agents: multi-file features, debugging and issue resolution carried through to a finished result.

  • Finance and legal work

    Analyst and legal workflows over filings, contracts and case files, where Google's table puts it ahead of every other model it lists.

  • Science and research

    Bioinformatics and lab research tasks, where Google reports gains over 3.7 Flash on BioMysteryBench and LABBench2.

  • Enterprise workflow automation

    Multi-step, tool-calling workflows with structured outputs, such as ticket triage or moving data between systems.

Get started

Call Gemini 3.8 Flash in three steps.

No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to Zurelay.

  1. 1

    Create a key

    Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.

  2. 2

    Point your SDK at Zurelay

    Change the base URL. Everything else in your code stays the same.

    https://api.zurelay.com/v1
  3. 3

    Use gemini-3.8-flash

    Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.

    gemini-3.8-flash

FAQ

Gemini 3.8 Flash API questions.

Anything else? Ask our team and we’ll answer by email.

How much does the Gemini 3.8 Flash API cost?

On Zurelay, gemini-3.8-flash costs $0.32 per 1M input tokens and $1.59 per 1M output tokens, with cached input at $0.032. Google's list price is $0.75 input and $3.75 output per 1M, so you save 58%. Thinking tokens bill as output, and the rate is the same at every prompt length.

Is there a cheaper Gemini 3.8 Flash API than Google's?

Yes. Zurelay serves gemini-3.8-flash at 58% below Google's list price, paid from prepaid credit with no subscription. Google's own price is an introductory rate that runs through December 31, 2026, and Google doubles it on January 1, 2027.

How do I call the Gemini 3.8 Flash API with the OpenAI SDK?

Install the official openai package, set base_url to https://api.zurelay.com/v1 and use your Zurelay API key. Then send a Chat Completions request with model set to "gemini-3.8-flash". Messages, tools and streaming use the standard OpenAI request format.

What is the Gemini 3.8 Flash context window?

Gemini 3.8 Flash takes up to 1,048,576 input tokens and writes up to 65,536 output tokens per response. The model accepts text, image, video, audio and PDF input and writes text.

How do Gemini 3.8 Flash thinking levels work?

Thinking is always on. The model supports low, medium and high, with medium as the default, and does not support minimal. On Zurelay you pick the level with reasoning_effort. Google says 3.8 Flash takes extra reasoning steps on complex tasks, so use low for quick tool steps and save high for the hardest problems.

How is Gemini 3.8 Flash different from 3.7 Flash?

It scores higher across Google's model card: 73.7% against 65.3% on DeepSWE v1.1, 89.4% against 85.8% on Terminal-bench 2.1 and 59.0% against 50.6% on OSWorld-2.0. Both have a 1M-token context, a 65,536-token output limit and the same Google list price. Google says 3.8 Flash works harder on complex tasks and points efficiency-first workloads to lower thinking levels or to 3.7 Flash, which remains fully supported.

What is Gemini 3.8 Flash best at?

Long-horizon coding, autonomous agents and specialist knowledge work. In Google's table it leads every listed model on finance and legal agent tasks, Terminal-bench 2.1, chart reasoning and long-video understanding. Claude Opus 5 still scores higher on general agent tasks and computer use.

Can I use Gemini 3.8 Flash Cyber on Zurelay?

No. Google offers Gemini 3.8 Flash Cyber, its cybersecurity variant, only to trusted defenders through its Fairwind Program. Zurelay serves Gemini 3.8 Flash, which Google ships with safeguards against misuse such as cyber offense.

Does the Gemini 3.8 Flash API support streaming and rate limits?

Yes. Set stream: true, and add stream_options.include_usage to get token counts in the final chunk. Each Zurelay API key can have its own monthly budget and requests-per-minute cap. Failed requests are retried on another route and never billed.

Is it the same model as Google's Gemini 3.8 Flash?

Yes, it is Google's Gemini 3.8 Flash model. Responses are sampled, so wording varies between runs, and Zurelay does not promise byte-identical output. Zurelay is independent and is not affiliated with or endorsed by Google.

More models on Zurelay

All models

Stop paying list price.
Start saving today.

Create a key in seconds, point the SDK you already use at Zurelay, and every request costs up to 90% less from the first token.