OpenAI
Operational· 100% uptime, 7 days

GPT-5.6 Terra API

GPT-5.6 Terra API for everyday production work, 73% below OpenAI's list price.

Price per 1M tokens73%off
Input
$0.54$2.00
Output
$3.26$12.00
Cached input
$0.054
Model IDgpt-5.6-terra
chat.py
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key="YOUR_ZURELAY_KEY",
)
stream = client.chat.completions.create(
model="gpt-5.6-terra",
messages=[{"role": "user", "content": "Explain quicksort in two sentences."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")

Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.

Context window
1.05M tokens
Max output
128K tokens
Input
Text, Image
Output
Text
Released
Jul 9, 2026
Uptime, 7 days
100%

Pricing

GPT-5.6 Terra API pricing. A fraction of the list price.

Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.

RatezurelayOpenAI listYou save
Input
per 1M tokens
$0.54$2.0073%off
Output
per 1M tokens
$3.26$12.0073%off
Cached input
per 1M tokens
$0.054——

One rate at every prompt length, up to the full 1.05M context window. Streaming, tool calls and structured outputs cost nothing extra.

Savings calculator

Move the sliders to your monthly usage.

Input tokens per month50M
Output tokens per month10M
zurelay
$59.60
OpenAI list price
$220.00
You save every month$160.40

$1,925 a year

Overview

What is GPT-5.6 Terra?

GPT-5.6 Terra is the middle tier of OpenAI's GPT-5.6 family, built to balance intelligence and cost. The GPT-5.6 Terra API on zurelay runs the same model through one OpenAI-compatible endpoint, below OpenAI's list price. It fits production traffic that needs solid reasoning without paying for Sol.

GPT-5.6 Terra is the middle tier of OpenAI's GPT-5.6 family, released in the API on July 9, 2026, between GPT-5.6 Sol and GPT-5.6 Luna. OpenAI calls it a balanced everyday model with performance competitive with GPT-5.5 at half the cost. It fills the mini slot of earlier GPT-5 releases. The GPT-5.6 Terra API on zurelay uses the same model ID, gpt-5.6-terra, on an OpenAI-compatible chat completions endpoint.

Terra shares Sol's limits. It reads text and images and writes text, with a 1,050,000-token context window, up to 922,000 input tokens and up to 128,000 output tokens per request. Its knowledge cutoff is February 16, 2026. On chat completions, reasoning_effort can be none, low, medium (the default), high or xhigh.

OpenAI cut Terra's list price by 20% on July 30, 2026. It still charges double for input and 1.5 times for output once a prompt passes 272K input tokens, while zurelay bills one flat rate at every length. Function calling, structured outputs, streaming and prompt caching all work on chat completions. As with every GPT-5.6 model, function tools there need reasoning_effort set to none.

OpenAI's GPT-6 family has no Terra tier. At OpenAI list prices, GPT-6 Sol costs the same per input token as GPT-5.6 Terra and less per output token, with a later knowledge cutoff. OpenAI also names Terra as the replacement for the gpt-5-mini snapshot it removes from its API on December 11, 2026. Terra is a natural next step for apps moving off GPT-5 mini or already tuned to GPT-5.6.

Strengths

Where GPT-5.6 Terra shines. And what teams build with it.

01

GPT-5.5-level work for less

OpenAI says Terra is competitive with GPT-5.5 at half the cost, which makes it a practical default for daily traffic.

02

Sol's context at a lower tier

Terra keeps the full 1,050,000-token window and 128,000-token output limit, so you can move down from Sol without trimming prompts.

03

One rate at any length

OpenAI raises the price past 272K input tokens. On zurelay, a 700K-token prompt bills at the same rate per token as a short one.

04

Reasoning when it pays off

Run with none for fast replies and raise effort up to xhigh for harder requests.

Use cases

  • Production assistants

    Customer-facing assistants and internal copilots that need good reasoning at a steady cost per request.

  • Everyday coding

    Code generation, pull request review and test writing, with Sol kept for the hardest changes.

  • Document processing

    Extract fields, compare contracts or summarize long reports into structured outputs you can parse.

  • Moving off GPT-5 mini

    OpenAI names Terra as the replacement for gpt-5-mini, whose snapshot leaves OpenAI's API on December 11, 2026.

Get started

Call GPT-5.6 Terra in three steps.

No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.

  1. 1

    Create a key

    Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.

  2. 2

    Point your SDK at zurelay

    Change the base URL. Everything else in your code stays the same.

    https://api.zurelay.com/v1
  3. 3

    Use gpt-5.6-terra

    Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.

    gpt-5.6-terra

FAQ

GPT-5.6 Terra API questions.

Anything else, write to sales@zurelay.com and an engineer will answer.

How much does the GPT-5.6 Terra API cost?

On zurelay, GPT-5.6 Terra costs $0.54 per 1M input tokens and $3.26 per 1M output tokens. OpenAI's list price is $2.00 input and $12.00 output, so you save 73%. Cached input bills at $0.054 per 1M, reasoning tokens count as output, and the rate stays the same above 272K input tokens.

Is there a cheaper GPT-5.6 Terra API?

zurelay serves gpt-5.6-terra at 73% below OpenAI's list price, with the same model ID and request format. You buy prepaid credit and pay per token, with no subscription. For simpler high-volume work, GPT-5.6 Luna costs much less per token.

How do I call the GPT-5.6 Terra API with the OpenAI SDK?

Install the official openai package and create a client with base_url set to https://api.zurelay.com/v1 and your zurelay API key. Send a chat completions request with model set to "gpt-5.6-terra" and pick a reasoning_effort. Cap output with max_completion_tokens. Only send temperature, top_p or logprobs when reasoning_effort is none.

What is the GPT-5.6 Terra context window?

GPT-5.6 Terra has a 1,050,000-token context window. One request can include up to 922,000 input tokens, and the model can write up to 128,000 output tokens. Input can be text or images; output is text. The knowledge cutoff is February 16, 2026.

GPT-5.6 Terra vs GPT-5.6 Sol: what's the difference?

They share the context window, output limit, input types and knowledge cutoff. Sol is the flagship for frontier reasoning and long agent runs; Terra trades some capability for a lower price. OpenAI says Terra is competitive with GPT-5.5, so you can run Terra by default and send only the hardest requests to Sol.

Does GPT-5.6 Terra support function calling and structured outputs?

Yes. Structured outputs work through response_format with a JSON schema. On chat completions, OpenAI accepts function tools with GPT-5.6 models only when reasoning_effort is none. Requests with tools at other levels, including the default medium, return an error.

What is GPT-5.6 Terra best used for?

OpenAI designed GPT-5.6 Terra for workloads that balance intelligence and cost. It fits production assistants, everyday coding, document processing and teams moving off GPT-5 mini. For the hardest agent tasks use GPT-5.6 Sol, and for bulk classification use GPT-5.6 Luna.

Does the GPT-5.6 Terra API support streaming, and what are the limits?

Yes. Set stream: true to receive tokens as they are generated, and add stream_options.include_usage to get token counts in the final chunk. Each zurelay API key can have its own monthly budget and requests-per-minute cap. Failed requests are retried on another route, and you are never billed for a request that fails.

Is it the same GPT-5.6 Terra model OpenAI serves?

Yes. Requests to gpt-5.6-terra run on OpenAI's GPT-5.6 Terra, not a smaller or different model. Two responses to the same prompt can still differ because of sampling, as with any call to OpenAI. zurelay is independent and is not affiliated with or endorsed by OpenAI.

More models on zurelay

All models

Stop paying list price.
Start saving today.

Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.