Astra's methods at a lower tier
OpenAI trained Sol with methods similar to Astra's, aimed at professional work, factuality, coding and computer use. At list price it costs one-fifth of Astra per token.
GPT-6 Sol API for coding and agent workflows, below OpenAI's list price.
gpt-6-solfrom openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key="YOUR_ZURELAY_KEY",)stream = client.chat.completions.create( model="gpt-6-sol", messages=[{"role": "user", "content": "Explain quicksort in two sentences."}], stream=True,)for chunk in stream: if chunk.choices: print(chunk.choices[0].delta.content or "", end="")Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.
Pricing
Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.
| Rate | zurelay | OpenAI list | You save |
|---|---|---|---|
Input per 1M tokens | $0.50 | 75%off | |
Output per 1M tokens | $2.50 | 75%off | |
Cached input per 1M tokens | $0.05 | — | — |
One rate at every prompt length, up to the full 1.05M context window. Streaming, tool calls and structured outputs cost nothing extra.
Move the sliders to your monthly usage.
$1,800 a year
Overview
GPT-6 Sol is the middle model in OpenAI's GPT-6 family, built for complex coding and agentic workflows. The GPT-6 Sol API on zurelay runs the same model through one OpenAI-compatible endpoint, below OpenAI's list price. It suits teams that want Sol's full reasoning range, from none to max.
GPT-6 Sol is the middle model in OpenAI's GPT-6 family, released on September 22, 2026, below GPT-6 Astra and above GPT-6 Luna. OpenAI describes it as built for complex coding and agentic workflows, and says it was trained with methods similar to Astra's to improve professional work, factuality, coding, computer use and alignment. The GPT-6 Sol API on zurelay uses the same model ID, gpt-6-sol, on an OpenAI-compatible Chat Completions endpoint.
Sol reads text and images and writes text. It has a 1,050,000-token context window, accepts up to 922,000 input tokens and can write up to 128,000 output tokens per request. Its knowledge cutoff is April 20, 2026. Reasoning effort runs from none through low, medium (the default), high and xhigh to max. OpenAI's docs say Chat Completions supports function calling on GPT-6 Sol when reasoning_effort is none.
OpenAI says GPT-6 Sol makes about half as many mistakes as GPT-5.6 Sol in its factuality testing. At list price it costs one-fifth of GPT-6 Astra per token, which makes it a practical model for work that repeats every day: building features, reviewing code, debugging and analyzing data.
One week after launch, OpenAI released GPT-6.1 Sol, and its docs now point to 6.1 Sol as the newer Sol model. 6.1 Sol has the same list price per token and cheaper cached input, but according to OpenAI it drops the none reasoning effort and needs OpenAI's Responses API for tool calling. Keep GPT-6 Sol if your app depends on none-effort answers, function calling on Chat Completions, or behavior you have already tested.
Strengths
OpenAI trained Sol with methods similar to Astra's, aimed at professional work, factuality, coding and computer use. At list price it costs one-fifth of Astra per token.
OpenAI says GPT-6 Sol makes about half as many mistakes as GPT-5.6 Sol in its factuality testing.
Reasoning effort goes from none to max. Use none for fast, direct answers and raise it for hard problems.
With reasoning_effort set to none, GPT-6 Sol supports function calling on Chat Completions, according to OpenAI's docs.
Use cases
Write and change code across a repository, with reasoning effort raised for the harder parts.
Review pull requests and trace bugs through logs and source files in one long context.
Work through data sets, reports and long exports with up to 922,000 input tokens per request.
Agent steps that call your functions with reasoning_effort none, where speed matters more than deep reasoning.
Get started
No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.
Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.
Change the base URL. Everything else in your code stays the same.
https://api.zurelay.com/v1Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.
gpt-6-solWorks with the tools you already use
Compare
$0.50 / $2.50 per 1M tokens
OpenAI's newer Sol model; pick it for stronger coding and professional work at the same list price per token.
GPT-6.1 Sol API$2.50 / $12.50 per 1M tokens
Pick it for the hardest reasoning, coding and research tasks, where quality matters more than cost.
GPT-6 Astra API$0.025 / $0.125 per 1M tokens
Pick it for focused, high-volume tasks where speed and price matter more than depth.
GPT-6 Luna APIFAQ
Anything else, write to sales@zurelay.com and an engineer will answer.
On zurelay, GPT-6 Sol costs $0.50 per 1M input tokens and $2.50 per 1M output tokens. OpenAI's list price is $2.00 input and $10.00 output, so you save 75%. Cached input tokens bill at $0.05 per 1M, and reasoning tokens count as output tokens.
zurelay serves gpt-6-sol at 75% below OpenAI's list price, with the same model ID and request format. If you need a lower price per token than that, GPT-6 Luna is the lowest-cost GPT-6 model and fits focused, high-volume tasks.
Install the official openai package and create a client with base_url set to https://api.zurelay.com/v1 and your zurelay API key. Then send a Chat Completions request with model set to "gpt-6-sol". Set reasoning_effort to control how much the model thinks; if it is anything other than none, leave out temperature, top_p and logprobs.
GPT-6 Sol has a 1,050,000-token context window. One request can include up to 922,000 input tokens, and the model can write up to 128,000 output tokens. Input can be text or images; output is text.
OpenAI built GPT-6 Sol for complex coding and agentic workflows. It fits work that repeats every day, such as building features, reviewing code, debugging and analyzing data. For summaries and extraction at high volume, GPT-6 Luna costs less per token.
Yes, with one condition. OpenAI's docs say Chat Completions supports function calling on GPT-6 Sol when reasoning_effort is set to none. Send tools in the standard Chat Completions format with that setting.
GPT-6.1 Sol is OpenAI's newer Sol model, released a week later, with stronger coding and professional-work results and cheaper cached input at the same list price per token. According to OpenAI, it drops the none reasoning effort and needs the Responses API for tool calling. GPT-6 Sol still supports none and function calling on Chat Completions.
Yes. Set stream: true to receive tokens as they are generated, and add stream_options.include_usage to get token counts in the final chunk. You can give each zurelay API key its own requests-per-minute limit and monthly budget in the dashboard. If an upstream call errors or times out, zurelay retries it on another route before returning an error.
Yes. Requests to gpt-6-sol run on OpenAI's GPT-6 Sol, and zurelay never swaps in a smaller or cheaper model. As with any call to OpenAI, two responses to the same prompt can differ because of sampling. zurelay is independent and is not affiliated with or endorsed by OpenAI.
Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.