OpenAI
Operational· 100% uptime, 7 days

GPT-6.1 Sol API

GPT-6.1 Sol API: near-Astra results on coding and professional work, for less.

Price per 1M tokens75%off
Input
$0.50$2.00
Output
$2.50$10.00
Cached input
$0.025
Model IDgpt-6.1-sol
chat.py
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key="YOUR_ZURELAY_KEY",
)
stream = client.chat.completions.create(
model="gpt-6.1-sol",
messages=[{"role": "user", "content": "Explain quicksort in two sentences."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")

Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.

Context window
1.05M tokens
Max output
128K tokens
Input
Text, Image
Output
Text
Released
Sep 29, 2026
Uptime, 7 days
100%

Pricing

GPT-6.1 Sol API pricing. A fraction of the list price.

Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.

RatezurelayOpenAI listYou save
Input
per 1M tokens
$0.50$2.0075%off
Output
per 1M tokens
$2.50$10.0075%off
Cached input
per 1M tokens
$0.025——

One rate at every prompt length, up to the full 1.05M context window. Streaming, tool calls and structured outputs cost nothing extra.

Savings calculator

Move the sliders to your monthly usage.

Input tokens per month50M
Output tokens per month10M
zurelay
$50.00
OpenAI list price
$200.00
You save every month$150.00

$1,800 a year

Overview

What is GPT-6.1 Sol?

GPT-6.1 Sol is OpenAI's newest Sol model, released on September 29, 2026. OpenAI says it nearly matches GPT-6 Astra on agentic coding, computer use and professional work at one-fifth of Astra's list price per token. The GPT-6.1 Sol API on zurelay serves it through one OpenAI-compatible endpoint, below OpenAI's list price.

GPT-6.1 Sol is the upgrade to GPT-6 Sol that OpenAI released at DevDay on September 29, 2026, one week after GPT-6 Sol. OpenAI describes it as near-Astra performance at a lower cost for complex coding, computer use and professional work, and lists it next to GPT-6 Astra and GPT-6 Luna as the balanced choice in the GPT-6 family. The GPT-6.1 Sol API on zurelay uses the model ID gpt-6.1-sol on an OpenAI-compatible Chat Completions endpoint.

The limits match the rest of the family: text and image input, text output, a 1,050,000-token context window, up to 922,000 input tokens and up to 128,000 output tokens. The knowledge cutoff is April 30, 2026, ten days later than GPT-6 Sol's. Reasoning effort accepts low, medium (the default), high, xhigh and max. Unlike GPT-6 Sol, it does not accept none, so use low when you want the fastest answers.

OpenAI reports gains over GPT-6 Sol in programming, debugging, document understanding and multi-step workflows. By OpenAI's figures, the factual error rate at low reasoning effort fell from 11.4% to 7.7%. In OpenAI's system card, GPT-6.1 Sol scored 64.2 on HealthBench Professional, against 64.7 for GPT-6 Astra.

OpenAI's list price per token is the same as GPT-6 Sol's, and cached input is priced at 5% of the input rate, half of GPT-6 Sol's cached rate. That helps agents that resend long, stable prompts. OpenAI's own advice is to compare 6.1 Sol with Astra on your tasks and weigh quality against cost; for most GPT-6 coding and agent work, 6.1 Sol is the place to start.

Strengths

Where GPT-6.1 Sol shines. And what teams build with it.

01

Near-Astra results

OpenAI says GPT-6.1 Sol nearly matches GPT-6 Astra on agentic coding, computer use and professional work.

02

One-fifth of Astra's list price

OpenAI lists GPT-6.1 Sol at one-fifth of Astra's input and output price per token, and zurelay charges less than that list rate.

03

More factual than GPT-6 Sol

OpenAI reports that its factual error rate at low reasoning effort fell from 11.4% to 7.7% compared with GPT-6 Sol.

04

Cheaper cached input

Cached input costs half of GPT-6 Sol's cached rate at list price, which lowers the bill for agents that resend long prompts.

Use cases

  • Coding agents

    Multi-step code changes, debugging and refactors where you want close to Astra's quality at a lower price.

  • Document-heavy work

    Read contracts, reports and specs in full inside a 1,050,000-token context and draft from them.

  • Business workflow automation

    Multi-step workflows that move through several documents, checks and decisions.

  • Default GPT-6 model

    Start new GPT-6 projects here, then move specific tasks up to Astra or down to Luna based on your own evals.

Get started

Call GPT-6.1 Sol in three steps.

No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.

  1. 1

    Create a key

    Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.

  2. 2

    Point your SDK at zurelay

    Change the base URL. Everything else in your code stays the same.

    https://api.zurelay.com/v1
  3. 3

    Use gpt-6.1-sol

    Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.

    gpt-6.1-sol

FAQ

GPT-6.1 Sol API questions.

Anything else, write to sales@zurelay.com and an engineer will answer.

How much does the GPT-6.1 Sol API cost?

On zurelay, GPT-6.1 Sol costs $0.50 per 1M input tokens and $2.50 per 1M output tokens. OpenAI's list price is $2.00 input and $10.00 output, so you save 75%. Cached input tokens bill at $0.025 per 1M, and reasoning tokens count as output tokens.

Is there a cheaper GPT-6.1 Sol API?

zurelay serves gpt-6.1-sol at 75% below OpenAI's list price, with the same model ID and request format. Cached input at $0.025 per 1M lowers the cost further when your prompts share a long prefix. For simpler, high-volume tasks, GPT-6 Luna costs less per token.

How do I call the GPT-6.1 Sol API with the OpenAI SDK?

Install the official openai package and create a client with base_url set to https://api.zurelay.com/v1 and your zurelay API key. Then send a Chat Completions request with model set to "gpt-6.1-sol". Set reasoning_effort to low, medium, high, xhigh or max, and leave out temperature, top_p and logprobs, which OpenAI does not accept while reasoning is on.

What is the GPT-6.1 Sol context window?

GPT-6.1 Sol has a 1,050,000-token context window. One request can include up to 922,000 input tokens, and the model can write up to 128,000 output tokens. Input can be text or images; output is text.

What is GPT-6.1 Sol best used for?

OpenAI built GPT-6.1 Sol for complex coding, computer use and professional work. It suits coding agents, document-heavy tasks and multi-step business workflows where you want results close to Astra's at a lower price.

GPT-6.1 Sol vs GPT-6 Sol: what changed?

OpenAI reports gains in programming, debugging, document understanding and multi-step workflows, and a lower factual error rate. The list price per token is unchanged, and cached input costs half as much. 6.1 Sol has a later knowledge cutoff, April 30, 2026, and no longer accepts the none reasoning effort.

GPT-6.1 Sol vs GPT-6 Astra: which should I use?

OpenAI says GPT-6.1 Sol nearly matches Astra on agentic coding, computer use and professional work, at one-fifth of Astra's list price per token. Astra remains OpenAI's pick for its most demanding work. Run both on a sample of your real prompts and pay for Astra only where it wins.

Does the GPT-6.1 Sol API support streaming, and what are the rate limits?

Yes. Set stream: true to receive tokens as they are generated, and add stream_options.include_usage to get token counts in the final chunk. You can give each zurelay API key its own requests-per-minute limit and monthly budget in the dashboard. If an upstream call errors or times out, zurelay retries it on another route before returning an error.

Is it the same GPT-6.1 Sol model OpenAI serves?

Yes. Requests to gpt-6.1-sol run on OpenAI's GPT-6.1 Sol, and zurelay never swaps in a smaller or cheaper model. As with any call to OpenAI, two responses to the same prompt can differ because of sampling. zurelay is independent and is not affiliated with or endorsed by OpenAI.

More models on zurelay

All models

Stop paying list price.
Start saving today.

Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.