Lowest GPT-6 price
At OpenAI's list prices, Luna costs one-twentieth of GPT-6 Sol per token, and zurelay charges less than that list rate.
GPT-6 Luna API: OpenAI's fastest, lowest-cost GPT-6 model for high-volume work.
gpt-6-lunafrom openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key="YOUR_ZURELAY_KEY",)stream = client.chat.completions.create( model="gpt-6-luna", messages=[{"role": "user", "content": "Explain quicksort in two sentences."}], stream=True,)for chunk in stream: if chunk.choices: print(chunk.choices[0].delta.content or "", end="")Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.
Pricing
Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.
| Rate | zurelay | OpenAI list | You save |
|---|---|---|---|
Input per 1M tokens | $0.025 | 75%off | |
Output per 1M tokens | $0.125 | 75%off | |
Cached input per 1M tokens | $0.0025 | — | — |
One rate at every prompt length, up to the full 1.05M context window. Streaming, tool calls and structured outputs cost nothing extra.
Move the sliders to your monthly usage.
$90.00 a year
Overview
GPT-6 Luna is the fastest and lowest-cost model in OpenAI's GPT-6 family, built for focused, high-volume tasks like summaries, extraction and quick answers. The GPT-6 Luna API on zurelay runs the same model through one OpenAI-compatible endpoint, below OpenAI's list price.
GPT-6 Luna is OpenAI's most efficient GPT-6 model, released on September 22, 2026 alongside GPT-6 Sol. OpenAI's model guide calls it the fastest and most cost-effective option in the family, for focused, high-volume tasks. The GPT-6 Luna API on zurelay uses the model ID gpt-6-luna on an OpenAI-compatible Chat Completions endpoint.
Luna keeps the family's limits: text and image input, text output, a 1,050,000-token context window, up to 922,000 input tokens and up to 128,000 output tokens. Its knowledge cutoff is May 18, 2026, the latest in the GPT-6 family. Reasoning effort runs from none to max, with medium as the default. OpenAI's docs say Chat Completions supports function calling on Luna when reasoning_effort is none.
Luna fits jobs with a clear goal: summarizing documents, extracting fields, routing requests and answering quick questions. At OpenAI's list prices it costs one-twentieth of GPT-6 Sol and one-hundredth of GPT-6 Astra per token, so it is the GPT-6 model to use when you run calls by the million. Structured outputs return clean JSON from extraction jobs.
When a task needs multi-step reasoning, longer coding work or computer use, move it up to GPT-6.1 Sol. A common pattern uses Luna to triage or route incoming requests and sends only the hard ones to a larger model.
Strengths
At OpenAI's list prices, Luna costs one-twentieth of GPT-6 Sol per token, and zurelay charges less than that list rate.
OpenAI designed Luna for fast responses and high-volume tasks with a clear goal.
Luna keeps the 1,050,000-token context window and 128,000-token output limit of the larger GPT-6 models, so long documents still fit.
Reasoning effort runs from none to max. Use none for the lowest latency, including function calling on Chat Completions.
Use cases
Condense documents, threads and transcripts at high volume.
Pull fields from emails, forms and scanned pages into JSON with structured outputs.
Classify incoming requests and send each one to the right model, queue or team.
Answer short, well-defined questions in chat and support flows where response time matters.
Get started
No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.
Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.
Change the base URL. Everything else in your code stays the same.
https://api.zurelay.com/v1Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.
gpt-6-lunaWorks with the tools you already use
Compare
$0.50 / $2.50 per 1M tokens
Pick it when the task needs deeper reasoning, multi-step coding or computer use.
GPT-6.1 Sol API$2.50 / $12.50 per 1M tokens
Pick it for the hardest reasoning and research work, where quality matters far more than price.
GPT-6 Astra API$0.037 / $0.15 per 1M tokens
Another low-cost model; test it side by side with Luna on your own high-volume prompts.
DeepSeek V4.1 Flash APIFAQ
Anything else, write to sales@zurelay.com and an engineer will answer.
On zurelay, GPT-6 Luna costs $0.025 per 1M input tokens and $0.125 per 1M output tokens. OpenAI's list price is $0.10 input and $0.50 output, so you save 75%. Cached input tokens bill at $0.0025 per 1M, and reasoning tokens count as output tokens.
GPT-6 Luna is already the lowest-priced GPT-6 model, and zurelay serves it at 75% below OpenAI's list price: $0.025 input and $0.125 output per 1M tokens. Cached input at $0.0025 per 1M cuts the cost further when prompts reuse the same prefix.
Install the official openai package and create a client with base_url set to https://api.zurelay.com/v1 and your zurelay API key. Then send a Chat Completions request with model set to "gpt-6-luna". Set reasoning_effort to none for the fastest answers; with any other value, leave out temperature, top_p and logprobs.
GPT-6 Luna has a 1,050,000-token context window, the same as the larger GPT-6 models. One request can include up to 922,000 input tokens, and the model can write up to 128,000 output tokens.
OpenAI built Luna for focused, high-volume tasks with a clear goal. Use it for summaries, data extraction, request routing and quick answers. For multi-step coding or agent work, GPT-6.1 Sol is the better fit.
Yes. Luna accepts text and image input and supports structured outputs, so you can send a scanned form or screenshot and get back JSON that matches your schema.
At OpenAI's list prices, Luna costs one-twentieth of Sol per token. OpenAI positions Sol for complex coding and agentic workflows and Luna for fast, focused, high-volume tasks. Start with Luna for simple jobs and move a task to Sol or GPT-6.1 Sol only when Luna's answers fall short.
Yes. Set stream: true to receive tokens as they are generated, and add stream_options.include_usage to get token counts in the final chunk. You can give each zurelay API key its own requests-per-minute limit and monthly budget in the dashboard. If an upstream call errors or times out, zurelay retries it on another route before returning an error.
Yes. Requests to gpt-6-luna run on OpenAI's GPT-6 Luna, and zurelay never swaps in a smaller or cheaper model. As with any call to OpenAI, two responses to the same prompt can differ because of sampling. zurelay is independent and is not affiliated with or endorsed by OpenAI.
Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.