GPT-5.5-level work for less
OpenAI says Terra is competitive with GPT-5.5 at half the cost, which makes it a practical default for daily traffic.
GPT-5.6 Terra API for everyday production work, 73% below OpenAI's list price.
gpt-5.6-terrafrom openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key="YOUR_ZURELAY_KEY",)stream = client.chat.completions.create( model="gpt-5.6-terra", messages=[{"role": "user", "content": "Explain quicksort in two sentences."}], stream=True,)for chunk in stream: if chunk.choices: print(chunk.choices[0].delta.content or "", end="")Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.
Pricing
Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.
| Rate | zurelay | OpenAI list | You save |
|---|---|---|---|
Input per 1M tokens | $0.54 | 73%off | |
Output per 1M tokens | $3.26 | 73%off | |
Cached input per 1M tokens | $0.054 | — | — |
One rate at every prompt length, up to the full 1.05M context window. Streaming, tool calls and structured outputs cost nothing extra.
Move the sliders to your monthly usage.
$1,925 a year
Overview
GPT-5.6 Terra is the middle tier of OpenAI's GPT-5.6 family, built to balance intelligence and cost. The GPT-5.6 Terra API on zurelay runs the same model through one OpenAI-compatible endpoint, below OpenAI's list price. It fits production traffic that needs solid reasoning without paying for Sol.
GPT-5.6 Terra is the middle tier of OpenAI's GPT-5.6 family, released in the API on July 9, 2026, between GPT-5.6 Sol and GPT-5.6 Luna. OpenAI calls it a balanced everyday model with performance competitive with GPT-5.5 at half the cost. It fills the mini slot of earlier GPT-5 releases. The GPT-5.6 Terra API on zurelay uses the same model ID, gpt-5.6-terra, on an OpenAI-compatible chat completions endpoint.
Terra shares Sol's limits. It reads text and images and writes text, with a 1,050,000-token context window, up to 922,000 input tokens and up to 128,000 output tokens per request. Its knowledge cutoff is February 16, 2026. On chat completions, reasoning_effort can be none, low, medium (the default), high or xhigh.
OpenAI cut Terra's list price by 20% on July 30, 2026. It still charges double for input and 1.5 times for output once a prompt passes 272K input tokens, while zurelay bills one flat rate at every length. Function calling, structured outputs, streaming and prompt caching all work on chat completions. As with every GPT-5.6 model, function tools there need reasoning_effort set to none.
OpenAI's GPT-6 family has no Terra tier. At OpenAI list prices, GPT-6 Sol costs the same per input token as GPT-5.6 Terra and less per output token, with a later knowledge cutoff. OpenAI also names Terra as the replacement for the gpt-5-mini snapshot it removes from its API on December 11, 2026. Terra is a natural next step for apps moving off GPT-5 mini or already tuned to GPT-5.6.
Strengths
OpenAI says Terra is competitive with GPT-5.5 at half the cost, which makes it a practical default for daily traffic.
Terra keeps the full 1,050,000-token window and 128,000-token output limit, so you can move down from Sol without trimming prompts.
OpenAI raises the price past 272K input tokens. On zurelay, a 700K-token prompt bills at the same rate per token as a short one.
Run with none for fast replies and raise effort up to xhigh for harder requests.
Use cases
Customer-facing assistants and internal copilots that need good reasoning at a steady cost per request.
Code generation, pull request review and test writing, with Sol kept for the hardest changes.
Extract fields, compare contracts or summarize long reports into structured outputs you can parse.
OpenAI names Terra as the replacement for gpt-5-mini, whose snapshot leaves OpenAI's API on December 11, 2026.
Get started
No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.
Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.
Change the base URL. Everything else in your code stays the same.
https://api.zurelay.com/v1Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.
gpt-5.6-terraWorks with the tools you already use
Compare
$1.14 / $5.70 per 1M tokens
Pick it for the hardest multi-step coding and agent tasks, where quality matters more than price.
GPT-5.6 Sol API$0.08 / $0.49 per 1M tokens
Pick it for high-volume, cost-sensitive tasks such as classification, routing and short summaries.
GPT-5.6 Luna API$0.50 / $2.50 per 1M tokens
Pick it for a newer, higher-tier model with a later knowledge cutoff at a similar OpenAI list price.
GPT-6 Sol APIFAQ
Anything else, write to sales@zurelay.com and an engineer will answer.
On zurelay, GPT-5.6 Terra costs $0.54 per 1M input tokens and $3.26 per 1M output tokens. OpenAI's list price is $2.00 input and $12.00 output, so you save 73%. Cached input bills at $0.054 per 1M, reasoning tokens count as output, and the rate stays the same above 272K input tokens.
zurelay serves gpt-5.6-terra at 73% below OpenAI's list price, with the same model ID and request format. You buy prepaid credit and pay per token, with no subscription. For simpler high-volume work, GPT-5.6 Luna costs much less per token.
Install the official openai package and create a client with base_url set to https://api.zurelay.com/v1 and your zurelay API key. Send a chat completions request with model set to "gpt-5.6-terra" and pick a reasoning_effort. Cap output with max_completion_tokens. Only send temperature, top_p or logprobs when reasoning_effort is none.
GPT-5.6 Terra has a 1,050,000-token context window. One request can include up to 922,000 input tokens, and the model can write up to 128,000 output tokens. Input can be text or images; output is text. The knowledge cutoff is February 16, 2026.
They share the context window, output limit, input types and knowledge cutoff. Sol is the flagship for frontier reasoning and long agent runs; Terra trades some capability for a lower price. OpenAI says Terra is competitive with GPT-5.5, so you can run Terra by default and send only the hardest requests to Sol.
Yes. Structured outputs work through response_format with a JSON schema. On chat completions, OpenAI accepts function tools with GPT-5.6 models only when reasoning_effort is none. Requests with tools at other levels, including the default medium, return an error.
OpenAI designed GPT-5.6 Terra for workloads that balance intelligence and cost. It fits production assistants, everyday coding, document processing and teams moving off GPT-5 mini. For the hardest agent tasks use GPT-5.6 Sol, and for bulk classification use GPT-5.6 Luna.
Yes. Set stream: true to receive tokens as they are generated, and add stream_options.include_usage to get token counts in the final chunk. Each zurelay API key can have its own monthly budget and requests-per-minute cap. Failed requests are retried on another route, and you are never billed for a request that fails.
Yes. Requests to gpt-5.6-terra run on OpenAI's GPT-5.6 Terra, not a smaller or different model. Two responses to the same prompt can still differ because of sampling, as with any call to OpenAI. zurelay is independent and is not affiliated with or endorsed by OpenAI.
Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.