Built for long agent runs
OpenAI designed Sol for frontier reasoning and long-horizon agentic work. At launch it scored 88.8% on Terminal-Bench 2.1, a state-of-the-art result for multi-step terminal tasks.
GPT-5.6 Sol API for hard coding and agent work, 72% below OpenAI's list price.
gpt-5.6-solfrom openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key="YOUR_ZURELAY_KEY",)stream = client.chat.completions.create( model="gpt-5.6-sol", messages=[{"role": "user", "content": "Explain quicksort in two sentences."}], stream=True,)for chunk in stream: if chunk.choices: print(chunk.choices[0].delta.content or "", end="")Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.
Pricing
Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.
| Rate | zurelay | OpenAI list | You save |
|---|---|---|---|
Input per 1M tokens | $1.14 | 72%off | |
Output per 1M tokens | $5.70 | 72%off | |
Cached input per 1M tokens | $0.114 | — | — |
One rate at every prompt length, up to the full 1.05M context window. Streaming, tool calls and structured outputs cost nothing extra.
Move the sliders to your monthly usage.
$3,432 a year
Overview
GPT-5.6 Sol is the flagship of OpenAI's GPT-5.6 family, built for frontier reasoning and long-horizon agentic work. The GPT-5.6 Sol API on zurelay runs the same model through one OpenAI-compatible endpoint, below OpenAI's list price. You pay one flat rate per token, even on prompts over 272K tokens.
GPT-5.6 Sol is the flagship of OpenAI's GPT-5.6 family, released in the API on July 9, 2026, above GPT-5.6 Terra and GPT-5.6 Luna. OpenAI built it for frontier reasoning and long-horizon agentic work, and its gpt-5.6 alias points to Sol. At launch, OpenAI reported a state-of-the-art 88.8% on Terminal-Bench 2.1, a test of multi-step command-line work. The GPT-5.6 Sol API on zurelay uses the same model ID, gpt-5.6-sol, on an OpenAI-compatible chat completions endpoint.
Sol reads text and images and writes text. It has a 1,050,000-token context window, takes up to 922,000 input tokens and writes up to 128,000 output tokens per request. Its knowledge cutoff is February 16, 2026. On chat completions, reasoning_effort can be none, low, medium (the default), high or xhigh. OpenAI's max effort level and Pro mode are Responses API features.
Function calling, structured outputs, streaming and prompt caching all work on chat completions, with one rule for tools. OpenAI accepts function tools with GPT-5.6 models on chat completions only when reasoning_effort is none. A request that sends tools at any other level, including the default medium, returns an error. The same goes for temperature, top_p and logprobs: they work only with reasoning_effort set to none.
OpenAI charges double for input and 1.5 times for output once a prompt passes 272K input tokens. zurelay bills one flat rate at every prompt length, so a 600K-token codebase review costs the same per token as a short question. OpenAI released GPT-6 Sol on September 22, 2026 at half of GPT-5.6 Sol's list price, and says it makes about half as many mistakes. Keep GPT-5.6 Sol for workloads whose prompts and evals are already tuned to it, and test GPT-6 Sol alongside.
Strengths
OpenAI designed Sol for frontier reasoning and long-horizon agentic work. At launch it scored 88.8% on Terminal-Bench 2.1, a state-of-the-art result for multi-step terminal tasks.
Sol takes up to 922,000 input tokens. OpenAI raises its price past 272K input tokens; zurelay keeps one flat rate at every length.
Set reasoning_effort from none for quick, direct answers up to xhigh for the hardest problems, on the same model ID.
Send screenshots, diagrams or scanned pages together with text, and get text back.
Use cases
Multi-step changes across a repository, terminal tasks and bug hunts, with reasoning effort raised for the hard parts.
Load large parts of a codebase or long logs into one request and ask for a review, a risk list or a root cause.
Work through long reports, papers or filings and return findings as structured outputs your code can parse.
Agents that call your functions on chat completions with reasoning_effort none, next to separate reasoning-heavy planning calls.
Get started
No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.
Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.
Change the base URL. Everything else in your code stays the same.
https://api.zurelay.com/v1Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.
gpt-5.6-solWorks with the tools you already use
Compare
$0.54 / $3.26 per 1M tokens
Pick it for everyday production work; OpenAI says it is competitive with GPT-5.5 at a lower price than Sol.
GPT-5.6 Terra API$0.50 / $2.50 per 1M tokens
OpenAI's newer Sol model; pick it for fewer factual mistakes at half the list price of GPT-5.6 Sol.
GPT-6 Sol API$2.50 / $12.50 per 1M tokens
Pick it for the hardest work that needs no function tools, since OpenAI supports Astra's tool calling only in the Responses API.
GPT-6 Astra APIFAQ
Anything else, write to sales@zurelay.com and an engineer will answer.
On zurelay, GPT-5.6 Sol costs $1.14 per 1M input tokens and $5.70 per 1M output tokens. OpenAI's list price is $4.00 input and $20.00 output, so you save 72%. Cached input bills at $0.114 per 1M, reasoning tokens count as output, and the rate stays the same above 272K input tokens. OpenAI says its current Sol list price is a promotion that runs at least through November 21, 2026.
zurelay serves gpt-5.6-sol at 72% below OpenAI's list price, with the same model ID and request format. You buy prepaid credit and pay per token, with no subscription. If you need a lower price per token, GPT-5.6 Terra and GPT-6 Sol both list below GPT-5.6 Sol.
Install the official openai package and create a client with base_url set to https://api.zurelay.com/v1 and your zurelay API key. Send a chat completions request with model set to "gpt-5.6-sol" and pick a reasoning_effort. Cap output with max_completion_tokens. Only send temperature, top_p or logprobs when reasoning_effort is none.
GPT-5.6 Sol has a 1,050,000-token context window. One request can include up to 922,000 input tokens, and the model can write up to 128,000 output tokens. Input can be text or images; output is text. The knowledge cutoff is February 16, 2026.
Yes, when reasoning_effort is set to none. OpenAI rejects chat completions requests that send function tools to GPT-5.6 models at any other effort level, including the default medium. Requests without tools can use any effort from none to xhigh.
OpenAI built GPT-5.6 Sol for frontier reasoning and long-horizon agentic work. It fits multi-step coding, terminal tasks, research and analysis of large inputs. For everyday production traffic GPT-5.6 Terra costs less, and for high-volume classification or summaries use GPT-5.6 Luna.
GPT-6 Sol is newer, released September 22, 2026, with a later knowledge cutoff of April 20, 2026. OpenAI says it makes about half as many mistakes as GPT-5.6 Sol, and its list price is half. Stay on GPT-5.6 Sol if your prompts and evals are tuned to it, and compare both on your own tasks before switching.
Yes. Set stream: true to receive tokens as they are generated, and add stream_options.include_usage to get token counts in the final chunk. Each zurelay API key can have its own monthly budget and requests-per-minute cap. Failed requests are retried on another route, and you are never billed for a request that fails.
Yes. Requests to gpt-5.6-sol run on OpenAI's GPT-5.6 Sol, not a smaller or different model. Two responses to the same prompt can still differ because of sampling, as with any call to OpenAI. zurelay is independent and is not affiliated with or endorsed by OpenAI.
Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.