OpenAI
Operational· 100% uptime, 7 days

GPT-5.6 Luna API

GPT-5.6 Luna API for fast, high-volume work, 59% below OpenAI's list price.

Price per 1M tokens59%off
Input
$0.08$0.20
Output
$0.49$1.20
Cached input
$0.008
Model IDgpt-5.6-luna
chat.py
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key="YOUR_ZURELAY_KEY",
)
stream = client.chat.completions.create(
model="gpt-5.6-luna",
messages=[{"role": "user", "content": "Explain quicksort in two sentences."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")

Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.

Context window
1.05M tokens
Max output
128K tokens
Input
Text, Image
Output
Text
Released
Jul 9, 2026
Uptime, 7 days
100%

Pricing

GPT-5.6 Luna API pricing. A fraction of the list price.

Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.

RatezurelayOpenAI listYou save
Input
per 1M tokens
$0.08$0.2060%off
Output
per 1M tokens
$0.49$1.2059%off
Cached input
per 1M tokens
$0.008——

One rate at every prompt length, up to the full 1.05M context window. Streaming, tool calls and structured outputs cost nothing extra.

Savings calculator

Move the sliders to your monthly usage.

Input tokens per month50M
Output tokens per month10M
zurelay
$8.90
OpenAI list price
$22.00
You save every month$13.10

$157.20 a year

Overview

What is GPT-5.6 Luna?

GPT-5.6 Luna is the fastest and lowest-cost tier of OpenAI's GPT-5.6 family, built for cost-sensitive, high-volume workloads. The GPT-5.6 Luna API on zurelay runs the same model through one OpenAI-compatible endpoint, below OpenAI's list price. It suits classification, routing, extraction and short summaries at scale.

GPT-5.6 Luna is the entry tier of OpenAI's GPT-5.6 family, released in the API on July 9, 2026. OpenAI calls it the fastest and most affordable GPT-5.6 model and designed it for cost-sensitive, high-volume workloads. It fills the nano slot of earlier GPT-5 releases. The GPT-5.6 Luna API on zurelay uses the same model ID, gpt-5.6-luna, on an OpenAI-compatible chat completions endpoint.

Luna keeps the family's full limits. It reads text and images and writes text, with a 1,050,000-token context window, up to 922,000 input tokens and up to 128,000 output tokens per request. Its knowledge cutoff is February 16, 2026. On chat completions, reasoning_effort can be none, low, medium (the default), high or xhigh.

OpenAI cut Luna's list price by 80% on July 30, 2026, three weeks after launch. It charges more once a prompt passes 272K input tokens, while zurelay bills one flat rate at every length. Structured outputs, streaming and prompt caching work on chat completions. Function tools work there too, but only with reasoning_effort set to none, as with every GPT-5.6 model.

OpenAI names GPT-5.6 Luna as the replacement for the gpt-5-nano snapshot it removes from its API on December 11, 2026. GPT-6 Luna, released September 22, 2026, is the newer Luna, with half the list price and a later knowledge cutoff of May 18, 2026. Keep GPT-5.6 Luna for pipelines you have already tested on it, and compare GPT-6 Luna on your own data.

Strengths

Where GPT-5.6 Luna shines. And what teams build with it.

01

Lowest-cost GPT-5.6 tier

Luna is the cheapest model in the GPT-5.6 family, after OpenAI cut its list price by 80% in July 2026.

02

Made for volume

OpenAI designed Luna for cost-sensitive, high-volume work, where cost and speed per request matter most.

03

Full-size context window

Luna keeps the family's 1,050,000-token window and 128,000-token output limit, so a small model can still read long inputs.

04

Reasoning is optional

Set reasoning_effort to none for the fastest answers, or raise it when a task needs more thought.

Use cases

  • Classification and routing

    Label tickets, detect intent or decide which model should handle a request, at a low cost per call.

  • Extraction to JSON

    Pull fields from emails, invoices or forms into a fixed schema with structured outputs.

  • Bulk summaries

    Summarize support threads, reviews or logs in large batches.

  • Image triage

    Read screenshots or photos and return short labels, captions or extracted text.

Get started

Call GPT-5.6 Luna in three steps.

No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.

  1. 1

    Create a key

    Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.

  2. 2

    Point your SDK at zurelay

    Change the base URL. Everything else in your code stays the same.

    https://api.zurelay.com/v1
  3. 3

    Use gpt-5.6-luna

    Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.

    gpt-5.6-luna

FAQ

GPT-5.6 Luna API questions.

Anything else, write to sales@zurelay.com and an engineer will answer.

How much does the GPT-5.6 Luna API cost?

On zurelay, GPT-5.6 Luna costs $0.08 per 1M input tokens and $0.49 per 1M output tokens. OpenAI's list price is $0.20 input and $1.20 output, so you save 59%. Cached input bills at $0.008 per 1M, reasoning tokens count as output, and the rate stays the same above 272K input tokens.

Is there a cheaper GPT-5.6 Luna API?

zurelay serves gpt-5.6-luna at 59% below OpenAI's list price, with the same model ID and request format. You buy prepaid credit and pay per token, with no subscription. GPT-6 Luna and GPT-5 nano both have lower OpenAI list prices per token, if they meet your quality bar.

How do I call the GPT-5.6 Luna API with the OpenAI SDK?

Install the official openai package and create a client with base_url set to https://api.zurelay.com/v1 and your zurelay API key. Send a chat completions request with model set to "gpt-5.6-luna" and pick a reasoning_effort; none gives the fastest replies. Cap output with max_completion_tokens. Only send temperature, top_p or logprobs when reasoning_effort is none.

What is the GPT-5.6 Luna context window?

GPT-5.6 Luna has a 1,050,000-token context window. One request can include up to 922,000 input tokens, and the model can write up to 128,000 output tokens. Input can be text or images; output is text. The knowledge cutoff is February 16, 2026.

Is GPT-5.6 Luna a good replacement for GPT-5 nano?

OpenAI recommends gpt-5.6-luna as the replacement for the gpt-5-nano-2025-08-07 snapshot, which leaves OpenAI's API on December 11, 2026. Luna has a larger context window, 1,050,000 tokens against 400,000, and a much newer knowledge cutoff, February 2026 against May 2024. One difference to plan for: on chat completions, Luna accepts function tools only with reasoning_effort set to none.

GPT-5.6 Luna vs GPT-6 Luna: which should I use?

GPT-6 Luna is newer, released September 22, 2026, with a knowledge cutoff of May 18, 2026 and half the list price. Both have a 1,050,000-token window and accept function tools on chat completions only with reasoning_effort none. Stay on GPT-5.6 Luna if your pipeline is tuned to it, and test GPT-6 Luna on the same data before switching.

What is GPT-5.6 Luna best used for?

OpenAI designed GPT-5.6 Luna for cost-sensitive, high-volume workloads. It fits classification, routing, extraction to JSON, bulk summaries and image triage. For harder reasoning or coding, step up to GPT-5.6 Terra.

Does the GPT-5.6 Luna API support streaming, and what are the limits?

Yes. Set stream: true to receive tokens as they are generated, and add stream_options.include_usage to get token counts in the final chunk. Each zurelay API key can have its own monthly budget and requests-per-minute cap. Failed requests are retried on another route, and you are never billed for a request that fails.

Is it the same GPT-5.6 Luna model OpenAI serves?

Yes. Requests to gpt-5.6-luna run on OpenAI's GPT-5.6 Luna, not a smaller or different model. Two responses to the same prompt can still differ because of sampling, as with any call to OpenAI. zurelay is independent and is not affiliated with or endorsed by OpenAI.

More models on zurelay

All models

Stop paying list price.
Start saving today.

Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.