Lowest-cost GPT-5.6 tier
Luna is the cheapest model in the GPT-5.6 family, after OpenAI cut its list price by 80% in July 2026.
GPT-5.6 Luna API for fast, high-volume work, 59% below OpenAI's list price.
gpt-5.6-lunafrom openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key="YOUR_ZURELAY_KEY",)stream = client.chat.completions.create( model="gpt-5.6-luna", messages=[{"role": "user", "content": "Explain quicksort in two sentences."}], stream=True,)for chunk in stream: if chunk.choices: print(chunk.choices[0].delta.content or "", end="")Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.
Pricing
Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.
| Rate | zurelay | OpenAI list | You save |
|---|---|---|---|
Input per 1M tokens | $0.08 | 60%off | |
Output per 1M tokens | $0.49 | 59%off | |
Cached input per 1M tokens | $0.008 | — | — |
One rate at every prompt length, up to the full 1.05M context window. Streaming, tool calls and structured outputs cost nothing extra.
Move the sliders to your monthly usage.
$157.20 a year
Overview
GPT-5.6 Luna is the fastest and lowest-cost tier of OpenAI's GPT-5.6 family, built for cost-sensitive, high-volume workloads. The GPT-5.6 Luna API on zurelay runs the same model through one OpenAI-compatible endpoint, below OpenAI's list price. It suits classification, routing, extraction and short summaries at scale.
GPT-5.6 Luna is the entry tier of OpenAI's GPT-5.6 family, released in the API on July 9, 2026. OpenAI calls it the fastest and most affordable GPT-5.6 model and designed it for cost-sensitive, high-volume workloads. It fills the nano slot of earlier GPT-5 releases. The GPT-5.6 Luna API on zurelay uses the same model ID, gpt-5.6-luna, on an OpenAI-compatible chat completions endpoint.
Luna keeps the family's full limits. It reads text and images and writes text, with a 1,050,000-token context window, up to 922,000 input tokens and up to 128,000 output tokens per request. Its knowledge cutoff is February 16, 2026. On chat completions, reasoning_effort can be none, low, medium (the default), high or xhigh.
OpenAI cut Luna's list price by 80% on July 30, 2026, three weeks after launch. It charges more once a prompt passes 272K input tokens, while zurelay bills one flat rate at every length. Structured outputs, streaming and prompt caching work on chat completions. Function tools work there too, but only with reasoning_effort set to none, as with every GPT-5.6 model.
OpenAI names GPT-5.6 Luna as the replacement for the gpt-5-nano snapshot it removes from its API on December 11, 2026. GPT-6 Luna, released September 22, 2026, is the newer Luna, with half the list price and a later knowledge cutoff of May 18, 2026. Keep GPT-5.6 Luna for pipelines you have already tested on it, and compare GPT-6 Luna on your own data.
Strengths
Luna is the cheapest model in the GPT-5.6 family, after OpenAI cut its list price by 80% in July 2026.
OpenAI designed Luna for cost-sensitive, high-volume work, where cost and speed per request matter most.
Luna keeps the family's 1,050,000-token window and 128,000-token output limit, so a small model can still read long inputs.
Set reasoning_effort to none for the fastest answers, or raise it when a task needs more thought.
Use cases
Label tickets, detect intent or decide which model should handle a request, at a low cost per call.
Pull fields from emails, invoices or forms into a fixed schema with structured outputs.
Summarize support threads, reviews or logs in large batches.
Read screenshots or photos and return short labels, captions or extracted text.
Get started
No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.
Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.
Change the base URL. Everything else in your code stays the same.
https://api.zurelay.com/v1Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.
gpt-5.6-lunaWorks with the tools you already use
Compare
$0.025 / $0.125 per 1M tokens
OpenAI's newer Luna; pick it for a later knowledge cutoff at half the list price of GPT-5.6 Luna.
GPT-6 Luna API$0.54 / $3.26 per 1M tokens
Pick it when Luna falls short on harder reasoning, coding or long-document work.
GPT-5.6 Terra API$2.50 / $12.50 per 1M tokens
GPT-6 Astra by OpenAI, 75% below list price.
GPT-6 Astra APIFAQ
Anything else, write to sales@zurelay.com and an engineer will answer.
On zurelay, GPT-5.6 Luna costs $0.08 per 1M input tokens and $0.49 per 1M output tokens. OpenAI's list price is $0.20 input and $1.20 output, so you save 59%. Cached input bills at $0.008 per 1M, reasoning tokens count as output, and the rate stays the same above 272K input tokens.
zurelay serves gpt-5.6-luna at 59% below OpenAI's list price, with the same model ID and request format. You buy prepaid credit and pay per token, with no subscription. GPT-6 Luna and GPT-5 nano both have lower OpenAI list prices per token, if they meet your quality bar.
Install the official openai package and create a client with base_url set to https://api.zurelay.com/v1 and your zurelay API key. Send a chat completions request with model set to "gpt-5.6-luna" and pick a reasoning_effort; none gives the fastest replies. Cap output with max_completion_tokens. Only send temperature, top_p or logprobs when reasoning_effort is none.
GPT-5.6 Luna has a 1,050,000-token context window. One request can include up to 922,000 input tokens, and the model can write up to 128,000 output tokens. Input can be text or images; output is text. The knowledge cutoff is February 16, 2026.
OpenAI recommends gpt-5.6-luna as the replacement for the gpt-5-nano-2025-08-07 snapshot, which leaves OpenAI's API on December 11, 2026. Luna has a larger context window, 1,050,000 tokens against 400,000, and a much newer knowledge cutoff, February 2026 against May 2024. One difference to plan for: on chat completions, Luna accepts function tools only with reasoning_effort set to none.
GPT-6 Luna is newer, released September 22, 2026, with a knowledge cutoff of May 18, 2026 and half the list price. Both have a 1,050,000-token window and accept function tools on chat completions only with reasoning_effort none. Stay on GPT-5.6 Luna if your pipeline is tuned to it, and test GPT-6 Luna on the same data before switching.
OpenAI designed GPT-5.6 Luna for cost-sensitive, high-volume workloads. It fits classification, routing, extraction to JSON, bulk summaries and image triage. For harder reasoning or coding, step up to GPT-5.6 Terra.
Yes. Set stream: true to receive tokens as they are generated, and add stream_options.include_usage to get token counts in the final chunk. Each zurelay API key can have its own monthly budget and requests-per-minute cap. Failed requests are retried on another route, and you are never billed for a request that fails.
Yes. Requests to gpt-5.6-luna run on OpenAI's GPT-5.6 Luna, not a smaller or different model. Two responses to the same prompt can still differ because of sampling, as with any call to OpenAI. zurelay is independent and is not affiliated with or endorsed by OpenAI.
Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.