OpenAI
Operational· 100% uptime, 7 days

GPT-6 Luna API

GPT-6 Luna API: OpenAI's fastest, lowest-cost GPT-6 model for high-volume work.

Price per 1M tokens75%off
Input
$0.025$0.10
Output
$0.125$0.50
Cached input
$0.0025
Model IDgpt-6-luna
chat.py
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key="YOUR_ZURELAY_KEY",
)
stream = client.chat.completions.create(
model="gpt-6-luna",
messages=[{"role": "user", "content": "Explain quicksort in two sentences."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")

Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.

Context window
1.05M tokens
Max output
128K tokens
Input
Text, Image
Output
Text
Released
Sep 22, 2026
Uptime, 7 days
100%

Pricing

GPT-6 Luna API pricing. A fraction of the list price.

Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.

RatezurelayOpenAI listYou save
Input
per 1M tokens
$0.025$0.1075%off
Output
per 1M tokens
$0.125$0.5075%off
Cached input
per 1M tokens
$0.0025——

One rate at every prompt length, up to the full 1.05M context window. Streaming, tool calls and structured outputs cost nothing extra.

Savings calculator

Move the sliders to your monthly usage.

Input tokens per month50M
Output tokens per month10M
zurelay
$2.50
OpenAI list price
$10.00
You save every month$7.50

$90.00 a year

Overview

What is GPT-6 Luna?

GPT-6 Luna is the fastest and lowest-cost model in OpenAI's GPT-6 family, built for focused, high-volume tasks like summaries, extraction and quick answers. The GPT-6 Luna API on zurelay runs the same model through one OpenAI-compatible endpoint, below OpenAI's list price.

GPT-6 Luna is OpenAI's most efficient GPT-6 model, released on September 22, 2026 alongside GPT-6 Sol. OpenAI's model guide calls it the fastest and most cost-effective option in the family, for focused, high-volume tasks. The GPT-6 Luna API on zurelay uses the model ID gpt-6-luna on an OpenAI-compatible Chat Completions endpoint.

Luna keeps the family's limits: text and image input, text output, a 1,050,000-token context window, up to 922,000 input tokens and up to 128,000 output tokens. Its knowledge cutoff is May 18, 2026, the latest in the GPT-6 family. Reasoning effort runs from none to max, with medium as the default. OpenAI's docs say Chat Completions supports function calling on Luna when reasoning_effort is none.

Luna fits jobs with a clear goal: summarizing documents, extracting fields, routing requests and answering quick questions. At OpenAI's list prices it costs one-twentieth of GPT-6 Sol and one-hundredth of GPT-6 Astra per token, so it is the GPT-6 model to use when you run calls by the million. Structured outputs return clean JSON from extraction jobs.

When a task needs multi-step reasoning, longer coding work or computer use, move it up to GPT-6.1 Sol. A common pattern uses Luna to triage or route incoming requests and sends only the hard ones to a larger model.

Strengths

Where GPT-6 Luna shines. And what teams build with it.

01

Lowest GPT-6 price

At OpenAI's list prices, Luna costs one-twentieth of GPT-6 Sol per token, and zurelay charges less than that list rate.

02

Built for speed and volume

OpenAI designed Luna for fast responses and high-volume tasks with a clear goal.

03

Full 1M-token context

Luna keeps the 1,050,000-token context window and 128,000-token output limit of the larger GPT-6 models, so long documents still fit.

04

Reasoning when you need it

Reasoning effort runs from none to max. Use none for the lowest latency, including function calling on Chat Completions.

Use cases

  • Summarization

    Condense documents, threads and transcripts at high volume.

  • Data extraction

    Pull fields from emails, forms and scanned pages into JSON with structured outputs.

  • Request routing

    Classify incoming requests and send each one to the right model, queue or team.

  • Quick answers

    Answer short, well-defined questions in chat and support flows where response time matters.

Get started

Call GPT-6 Luna in three steps.

No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.

  1. 1

    Create a key

    Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.

  2. 2

    Point your SDK at zurelay

    Change the base URL. Everything else in your code stays the same.

    https://api.zurelay.com/v1
  3. 3

    Use gpt-6-luna

    Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.

    gpt-6-luna

FAQ

GPT-6 Luna API questions.

Anything else, write to sales@zurelay.com and an engineer will answer.

How much does the GPT-6 Luna API cost?

On zurelay, GPT-6 Luna costs $0.025 per 1M input tokens and $0.125 per 1M output tokens. OpenAI's list price is $0.10 input and $0.50 output, so you save 75%. Cached input tokens bill at $0.0025 per 1M, and reasoning tokens count as output tokens.

Is there a cheaper GPT-6 Luna API?

GPT-6 Luna is already the lowest-priced GPT-6 model, and zurelay serves it at 75% below OpenAI's list price: $0.025 input and $0.125 output per 1M tokens. Cached input at $0.0025 per 1M cuts the cost further when prompts reuse the same prefix.

How do I call the GPT-6 Luna API with the OpenAI SDK?

Install the official openai package and create a client with base_url set to https://api.zurelay.com/v1 and your zurelay API key. Then send a Chat Completions request with model set to "gpt-6-luna". Set reasoning_effort to none for the fastest answers; with any other value, leave out temperature, top_p and logprobs.

What is the GPT-6 Luna context window?

GPT-6 Luna has a 1,050,000-token context window, the same as the larger GPT-6 models. One request can include up to 922,000 input tokens, and the model can write up to 128,000 output tokens.

What is GPT-6 Luna best used for?

OpenAI built Luna for focused, high-volume tasks with a clear goal. Use it for summaries, data extraction, request routing and quick answers. For multi-step coding or agent work, GPT-6.1 Sol is the better fit.

Can GPT-6 Luna read images and return JSON?

Yes. Luna accepts text and image input and supports structured outputs, so you can send a scanned form or screenshot and get back JSON that matches your schema.

GPT-6 Luna vs GPT-6 Sol: which should I use?

At OpenAI's list prices, Luna costs one-twentieth of Sol per token. OpenAI positions Sol for complex coding and agentic workflows and Luna for fast, focused, high-volume tasks. Start with Luna for simple jobs and move a task to Sol or GPT-6.1 Sol only when Luna's answers fall short.

Does the GPT-6 Luna API support streaming, and what are the rate limits?

Yes. Set stream: true to receive tokens as they are generated, and add stream_options.include_usage to get token counts in the final chunk. You can give each zurelay API key its own requests-per-minute limit and monthly budget in the dashboard. If an upstream call errors or times out, zurelay retries it on another route before returning an error.

Is it the same GPT-6 Luna model OpenAI serves?

Yes. Requests to gpt-6-luna run on OpenAI's GPT-6 Luna, and zurelay never swaps in a smaller or cheaper model. As with any call to OpenAI, two responses to the same prompt can differ because of sampling. zurelay is independent and is not affiliated with or endorsed by OpenAI.

More models on zurelay

All models

Stop paying list price.
Start saving today.

Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.