Abstract reasoning
A verified 77.1% on ARC-AGI-2, more than double Gemini 3 Pro's 31.1%, on puzzles that test entirely new logic patterns.
The Gemini 3.1 Pro API, in preview, for complex reasoning and coding, 60% below Google's list price
gemini-3.1-profrom openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key="YOUR_ZURELAY_KEY",)stream = client.chat.completions.create( model="gemini-3.1-pro", messages=[{"role": "user", "content": "Explain quicksort in two sentences."}], stream=True,)for chunk in stream: if chunk.choices: print(chunk.choices[0].delta.content or "", end="")Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK. Chat API docs
Pricing
Pay as you go from prepaid credit, with no subscription. Top up from $10. Requests that fail are never billed.
| Rate | Zurelay | Google list | You save |
|---|---|---|---|
Input per 1M tokens | $0.80 | 60%off | |
Output per 1M tokens | $4.80 | 60%off | |
Cached input per 1M tokens | $0.08 | — | — |
One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.
Move the sliders to your monthly usage.
$1,584 a year
Overview
Gemini 3.1 Pro is Google's Pro model for complex problem-solving and agentic coding, released as a preview on February 19, 2026 and still in preview at Google. The Zurelay Gemini 3.1 Pro API serves it through one OpenAI-compatible endpoint with a 1M-token context window, at 60% below Google's list price and one rate at every prompt length.
Google released Gemini 3.1 Pro as a preview on February 19, 2026, and it is still a preview model: Google serves it as gemini-3.1-pro-preview and has not announced a date for general availability. The Gemini 3.1 Pro API model reads text, images, video, audio and PDFs and writes text, with up to 1,048,576 input tokens and 65,536 output tokens. On Zurelay the model ID is gemini-3.1-pro, Google's ID gemini-3.1-pro-preview works as an alias, and it costs $0.80 per 1M input tokens and $4.80 per 1M output tokens.
Google built it to refine the Gemini 3 Pro series, with better thinking, improved token efficiency and more grounded, factually consistent answers, and optimized it for software engineering and for agentic workflows that need precise tool use and reliable multi-step execution. Its headline result is a verified 77.1% on ARC-AGI-2, a test of entirely new logic patterns, more than double Gemini 3 Pro's 31.1%. Google's model card also lists 94.3% on GPQA Diamond, 44.4% on Humanity's Last Exam without tools, 80.6% on SWE-Bench Verified, 68.5% on Terminal-Bench 2.0 and 85.9% on BrowseComp.
Thinking is always on. Gemini 3.1 Pro supports the low, medium and high thinking levels, with high as the default; minimal is not supported. The model also supports function calling, structured outputs and context caching. Thinking tokens bill as output, and Google recommends leaving temperature at its default of 1.0 for Gemini 3 models.
Google's list price rises with prompt length: it charges double for input and 1.5 times for output once a prompt passes 200K tokens. Zurelay charges $0.80 and $4.80 per 1M tokens at every prompt length, so long prompts save even more. Google says preview models may be used in production, might come with more restrictive rate limits and are deprecated with at least two weeks' notice. It has not announced a shutdown date for gemini-3.1-pro-preview.
Strengths
A verified 77.1% on ARC-AGI-2, more than double Gemini 3 Pro's 31.1%, on puzzles that test entirely new logic patterns.
Google's model card lists 80.6% on SWE-Bench Verified, 68.5% on Terminal-Bench 2.0 and an Elo of 2887 on LiveCodeBench Pro.
69.2% on MCP Atlas, a test of multi-step workflows over MCP tools, and 85.9% on BrowseComp, a test of agentic search.
Google charges more once a prompt passes 200K tokens. Zurelay charges $0.80 and $4.80 per 1M tokens at every length, up to the full 1M-token window.
Use cases
Multi-step analysis, math and science questions where a simple answer isn't enough and deeper thinking pays off.
Repository-scale changes, debugging and terminal work in coding agents that call tools over many steps.
Animated SVGs, interactive 3D scenes and live dashboards that visualize real telemetry, all examples Google showed at launch.
Read long reports, codebases, recordings or video in one request of up to 1,048,576 tokens and synthesize what they say.
Get started
No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to Zurelay.
Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.
Change the base URL. Everything else in your code stays the same.
https://api.zurelay.com/v1Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.
gemini-3.1-proWorks with the tools you already use
Compare
$0.32 / $1.59 per 1M tokens
Pick Gemini 3.8 Flash for a stable model at a lower list price; Google calls Gemini 3.8 its best reasoning and coding model yet.
Gemini 3.8 Flash API$0.32 / $1.59 per 1M tokens
Pick Gemini 3.7 Flash for stable, lower-cost coding and agent work with the same 1M-token context and input types.
Gemini 3.7 Flash API$1.40 / $7.00 per 1M tokens
Pick Claude Opus 5.5 to compare Anthropic's Opus model on long-running agentic coding and knowledge work.
Claude Opus 5.5 APIOn Zurelay, gemini-3.1-pro costs $0.80 per 1M input tokens and $4.80 per 1M output tokens, with cached input at $0.08. Google's list price is $2.00 input and $12.00 output per 1M for prompts up to 200K tokens, so you save 60%. Zurelay's rate is the same at every prompt length, and thinking tokens bill as output.
Yes. Zurelay serves gemini-3.1-pro at 60% below Google's list price for prompts up to 200K tokens. Above that, Google charges double for input and 1.5 times for output while Zurelay's rate stays the same, so the saving grows. You pay from prepaid credit, with no subscription.
Yes. Google released it as a preview on February 19, 2026, still serves it as gemini-3.1-pro-preview and has not announced a date for general availability. Google says preview models may be used in production, might come with more restrictive rate limits and are deprecated with at least two weeks' notice.
Install the official openai package, set base_url to https://api.zurelay.com/v1 and use your Zurelay API key. Then send a Chat Completions request with model set to "gemini-3.1-pro". Zurelay also accepts Google's ID, gemini-3.1-pro-preview. Leave temperature at its default of 1.0, as Google recommends for Gemini 3 models.
Gemini 3.1 Pro takes up to 1,048,576 input tokens and writes up to 65,536 output tokens per response. The model accepts text, image, video, audio and PDF input and writes text.
Thinking is always on. The model supports low, medium and high, with high as the default, and does not support minimal. On Zurelay you pick the level with reasoning_effort: low or medium for quicker, cheaper answers, and high for the hardest problems.
Complex problem-solving, agentic coding and tool use. Google reports a verified 77.1% on ARC-AGI-2, 80.6% on SWE-Bench Verified and 85.9% on BrowseComp, and at launch showed it building animated SVGs, live dashboards and interactive 3D scenes from prompts.
Gemini 3.8 Flash is newer, stable and has a lower list price, and Google calls Gemini 3.8 its best reasoning and coding model yet. Gemini 3.1 Pro is Google's Pro model and is still in preview. Both take the same input types with a 1M-token context and 65,536 output tokens, and both work with the same Zurelay key, so test them side by side on your own prompts.
Yes. Set stream: true, and add stream_options.include_usage to get token counts in the final chunk. Each Zurelay API key can have its own monthly budget and requests-per-minute cap. Failed requests are retried on another route and never billed.
Yes, it is Google's Gemini 3.1 Pro Preview model. Responses are sampled, so wording varies between runs, and Zurelay does not promise byte-identical output. Zurelay is independent and is not affiliated with or endorsed by Google.
Create a key in seconds, point the SDK you already use at Zurelay, and every request costs up to 90% less from the first token.