OpenAI
Degraded· 96.6% uptime, 7 days

GPT-6 Sol API

GPT-6 Sol API for coding and agent workflows, below OpenAI's list price.

Price per 1M tokens75%off
Input
$0.50$2.00
Output
$2.50$10.00
Cached input
$0.05
Model IDgpt-6-sol
chat.py
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key="YOUR_ZURELAY_KEY",
)
stream = client.chat.completions.create(
model="gpt-6-sol",
messages=[{"role": "user", "content": "Explain quicksort in two sentences."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")

Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.

Context window
1.05M tokens
Max output
128K tokens
Input
Text, Image
Output
Text
Released
Sep 22, 2026
Uptime, 7 days
96.6%

Pricing

GPT-6 Sol API pricing. A fraction of the list price.

Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.

RatezurelayOpenAI listYou save
Input
per 1M tokens
$0.50$2.0075%off
Output
per 1M tokens
$2.50$10.0075%off
Cached input
per 1M tokens
$0.05——

One rate at every prompt length, up to the full 1.05M context window. Streaming, tool calls and structured outputs cost nothing extra.

Savings calculator

Move the sliders to your monthly usage.

Input tokens per month50M
Output tokens per month10M
zurelay
$50.00
OpenAI list price
$200.00
You save every month$150.00

$1,800 a year

Overview

What is GPT-6 Sol?

GPT-6 Sol is the middle model in OpenAI's GPT-6 family, built for complex coding and agentic workflows. The GPT-6 Sol API on zurelay runs the same model through one OpenAI-compatible endpoint, below OpenAI's list price. It suits teams that want Sol's full reasoning range, from none to max.

GPT-6 Sol is the middle model in OpenAI's GPT-6 family, released on September 22, 2026, below GPT-6 Astra and above GPT-6 Luna. OpenAI describes it as built for complex coding and agentic workflows, and says it was trained with methods similar to Astra's to improve professional work, factuality, coding, computer use and alignment. The GPT-6 Sol API on zurelay uses the same model ID, gpt-6-sol, on an OpenAI-compatible Chat Completions endpoint.

Sol reads text and images and writes text. It has a 1,050,000-token context window, accepts up to 922,000 input tokens and can write up to 128,000 output tokens per request. Its knowledge cutoff is April 20, 2026. Reasoning effort runs from none through low, medium (the default), high and xhigh to max. OpenAI's docs say Chat Completions supports function calling on GPT-6 Sol when reasoning_effort is none.

OpenAI says GPT-6 Sol makes about half as many mistakes as GPT-5.6 Sol in its factuality testing. At list price it costs one-fifth of GPT-6 Astra per token, which makes it a practical model for work that repeats every day: building features, reviewing code, debugging and analyzing data.

One week after launch, OpenAI released GPT-6.1 Sol, and its docs now point to 6.1 Sol as the newer Sol model. 6.1 Sol has the same list price per token and cheaper cached input, but according to OpenAI it drops the none reasoning effort and needs OpenAI's Responses API for tool calling. Keep GPT-6 Sol if your app depends on none-effort answers, function calling on Chat Completions, or behavior you have already tested.

Strengths

Where GPT-6 Sol shines. And what teams build with it.

01

Astra's methods at a lower tier

OpenAI trained Sol with methods similar to Astra's, aimed at professional work, factuality, coding and computer use. At list price it costs one-fifth of Astra per token.

02

Fewer factual mistakes

OpenAI says GPT-6 Sol makes about half as many mistakes as GPT-5.6 Sol in its factuality testing.

03

Full reasoning range

Reasoning effort goes from none to max. Use none for fast, direct answers and raise it for hard problems.

04

Function calling on Chat Completions

With reasoning_effort set to none, GPT-6 Sol supports function calling on Chat Completions, according to OpenAI's docs.

Use cases

  • Feature development

    Write and change code across a repository, with reasoning effort raised for the harder parts.

  • Code review and debugging

    Review pull requests and trace bugs through logs and source files in one long context.

  • Data analysis

    Work through data sets, reports and long exports with up to 922,000 input tokens per request.

  • Fast tool-calling steps

    Agent steps that call your functions with reasoning_effort none, where speed matters more than deep reasoning.

Get started

Call GPT-6 Sol in three steps.

No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.

  1. 1

    Create a key

    Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.

  2. 2

    Point your SDK at zurelay

    Change the base URL. Everything else in your code stays the same.

    https://api.zurelay.com/v1
  3. 3

    Use gpt-6-sol

    Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.

    gpt-6-sol

FAQ

GPT-6 Sol API questions.

Anything else, write to sales@zurelay.com and an engineer will answer.

How much does the GPT-6 Sol API cost?

On zurelay, GPT-6 Sol costs $0.50 per 1M input tokens and $2.50 per 1M output tokens. OpenAI's list price is $2.00 input and $10.00 output, so you save 75%. Cached input tokens bill at $0.05 per 1M, and reasoning tokens count as output tokens.

Is there a cheaper GPT-6 Sol API?

zurelay serves gpt-6-sol at 75% below OpenAI's list price, with the same model ID and request format. If you need a lower price per token than that, GPT-6 Luna is the lowest-cost GPT-6 model and fits focused, high-volume tasks.

How do I call the GPT-6 Sol API with the OpenAI SDK?

Install the official openai package and create a client with base_url set to https://api.zurelay.com/v1 and your zurelay API key. Then send a Chat Completions request with model set to "gpt-6-sol". Set reasoning_effort to control how much the model thinks; if it is anything other than none, leave out temperature, top_p and logprobs.

What is the GPT-6 Sol context window?

GPT-6 Sol has a 1,050,000-token context window. One request can include up to 922,000 input tokens, and the model can write up to 128,000 output tokens. Input can be text or images; output is text.

What is GPT-6 Sol best used for?

OpenAI built GPT-6 Sol for complex coding and agentic workflows. It fits work that repeats every day, such as building features, reviewing code, debugging and analyzing data. For summaries and extraction at high volume, GPT-6 Luna costs less per token.

Does GPT-6 Sol support function calling?

Yes, with one condition. OpenAI's docs say Chat Completions supports function calling on GPT-6 Sol when reasoning_effort is set to none. Send tools in the standard Chat Completions format with that setting.

GPT-6 Sol vs GPT-6.1 Sol: what's the difference?

GPT-6.1 Sol is OpenAI's newer Sol model, released a week later, with stronger coding and professional-work results and cheaper cached input at the same list price per token. According to OpenAI, it drops the none reasoning effort and needs the Responses API for tool calling. GPT-6 Sol still supports none and function calling on Chat Completions.

Does the GPT-6 Sol API support streaming, and what are the rate limits?

Yes. Set stream: true to receive tokens as they are generated, and add stream_options.include_usage to get token counts in the final chunk. You can give each zurelay API key its own requests-per-minute limit and monthly budget in the dashboard. If an upstream call errors or times out, zurelay retries it on another route before returning an error.

Is it the same GPT-6 Sol model OpenAI serves?

Yes. Requests to gpt-6-sol run on OpenAI's GPT-6 Sol, and zurelay never swaps in a smaller or cheaper model. As with any call to OpenAI, two responses to the same prompt can differ because of sampling. zurelay is independent and is not affiliated with or endorsed by OpenAI.

More models on zurelay

All models

Stop paying list price.
Start saving today.

Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.