Zurelay
Operationalยท checked every few minutes

Zurelay Auto API

One model name for every request. Zurelay Auto picks the model that fits each one and bills you that model's price, from $0.025 input and $0.125 output per 1M tokens.

The price of the model it picks, per 1M tokens14 models
Input
from $0.025
to $3.50
Output
from $0.125
to $17.50
Model IDzurelay-auto
chat.py
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key="YOUR_ZURELAY_KEY",
)
stream = client.chat.completions.create(
model="zurelay-auto",
messages=[{"role": "user", "content": "Explain quicksort in two sentences."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")

Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK, and the Anthropic SDK. Zurelay Auto docs

Picks from
14 models
Billed at
The price of its pick
Added latency
None
Context window
Up to 1.05M tokens
Reasoning
none to max
Status
Operational

Pricing

Zurelay Auto API pricing. The price of the model it picks.

Each request is billed at the price of the model that answered it, with nothing added for the pick. Pay as you go from prepaid credit, with no subscription. Requests that fail are never billed.

ModelInputOutput
Everyday requestsShort chat, questions, extraction, summaries.
DeepSeek V4.1 Flash$0.037$0.15
Gemini 3.1 Flash-Lite$0.10$0.59
Gemini 3.5 Flash-Lite$0.12$1.04
GPT-6 Luna$0.025$0.125
Qwen 3.8 Flash$0.07$0.22
Claude Haiku 4.5$0.38$1.90
Tools, code and structureTools, a JSON schema, code, long prompts, a low or medium reasoning level.
Claude Sonnet 5.5$0.80$4.00
GPT-6 Sol$0.50$2.50
Gemini 3.7 Flash$0.32$1.59
GPT-6.1 Sol$0.50$2.50
DeepSeek V4 Pro 0813$0.23$0.62
Deep reasoningA reasoning level of high or more.
Claude Opus 5.5$1.40$7.00
GPT-6.1 Sol$0.50$2.50
GPT-6 Astra$2.50$12.50
Claude Fable 5.1$3.50$17.50

Prices per 1M tokens. Each tier lists its models in the order Auto tries them; it skips any that can't read what the request sends.

How it picks

  1. 1
    It reads the request

    The media it sends, its length, tools, a JSON schema, code, and the reasoning level you set. No model call, so no added delay.

  2. 2
    It picks the model that fits

    The first model in the request's tier that reads what you sent and whose context window holds it.

  3. 3
    That model answers, and heals

    Retries and its backup routes, as for any model on Zurelay.

  4. 4
    Then the next best models

    If it still can't answer, the rest of its tier, then the tier above, then the one below. Your smart routing setting doesn't turn that off.

The answer's model field and the x-zurelay-model header name the model that answered, and you pay its price.

Overview

What is Zurelay Auto?

Zurelay Auto is a model you call like any other, zurelay-auto, that answers each request with the model that fits it. Everyday requests go to fast, low-cost models, tools, code and structured output to stronger ones, and requests that ask for deep reasoning to the strongest. You pay the price of the model that answered, and the answer tells you which one it was.

Zurelay Auto reads what each request carries and picks a model for it without an extra model call, so it adds no delay and costs nothing on top. It looks at the images, files, audio or video you send, how long the request is, whether it gives the model tools or a JSON schema, whether it's about code, and the reasoning level you set.

The pick follows three tiers. Short, everyday requests go to fast, low-cost models such as DeepSeek V4.1 Flash and Gemini 3.1 Flash-Lite. Requests with tools, a JSON schema, code, a long prompt or a low or medium reasoning level go to stronger models such as Claude Sonnet 5.5 and GPT-6 Sol. A reasoning level of high or more goes to the strongest, such as Claude Opus 5.5 and GPT-6.1 Sol. It only picks models that read what you send and whose context window holds it.

Each request is billed at the price of the model that answered it: from $0.025 to $3.50 per 1M input tokens and from $0.125 to $17.50 per 1M output tokens, never more than that model costs on its own. The answer's model field and the x-zurelay-model header name the model, and your Requests page shows it too.

Zurelay Auto is built to answer. The model it picks heals like any model on Zurelay: it retries, then uses its backup routes. If that model still can't answer, Auto moves on to its next best models, whatever your smart routing setting. Streaming, tool calls, structured outputs and reasoning levels work as with any model, through /v1/chat/completions or /v1/messages.

Strengths

Where Zurelay Auto shines. And what teams build with it.

01

The right model for each request

Fast, low-cost models for everyday requests, stronger ones for tools, code and structured output, and the strongest when you ask for deep reasoning.

02

You pay the model's price

Nothing extra for the pick. Each request costs what the model that answered costs on Zurelay, and the answer names it.

03

Built to answer

Retries and backup routes first, then its next best models. A key's smart routing setting doesn't turn that off.

04

One name in your code

Set zurelay-auto once. As better or cheaper models arrive, the models Auto picks from change without a line of your code changing.

Use cases

  • Apps with mixed traffic

    Chat, extraction and summaries next to agent steps and hard questions, each on a model that fits, from one model name.

  • Prototypes

    Start building before choosing a model. See which models Auto picks for your traffic, then pin one if you like.

  • Cost control

    Everyday requests stay on low-cost models, so a busy app's bill follows the work it does.

  • Agents

    Tool calls go to models that call tools reliably, and steps that ask for deep reasoning go to the strongest.

Get started

Call Zurelay Auto in three steps.

No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to Zurelay.

  1. 1

    Create a key

    Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.

  2. 2

    Point your SDK at Zurelay

    Change the base URL. Everything else in your code stays the same.

    https://api.zurelay.com/v1
  3. 3

    Use zurelay-auto

    Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.

    zurelay-auto

FAQ

Zurelay Auto API questions.

Anything else? Ask our team and weโ€™ll answer by email.

What does Zurelay Auto cost?

Each request costs what the model that answered it costs on Zurelay: from $0.025 to $3.50 per 1M input tokens and from $0.125 to $17.50 per 1M output tokens. There's no extra charge for the pick, and requests that fail are never billed.

How does Zurelay Auto pick a model?

From what the request carries, without an extra model call. Everyday requests go to fast, low-cost models. Tools, a JSON schema, code, a long prompt or a low or medium reasoning level go to stronger models, and a reasoning level of high or more to the strongest. It only picks models that read the media you send and whose context window holds the request.

How do I know which model answered?

The answer's model field and the x-zurelay-model header name it, and x-zurelay-routed-from says zurelay-auto. Your Requests page shows the same.

What happens if the model it picks is down?

It heals like any model: it retries and uses that model's backup routes. If the model still can't answer, Auto moves on to its next best models, whatever your smart routing setting.

Does Zurelay Auto work with the Anthropic SDK?

Yes. Through /v1/messages it picks among Claude models, the ones that take Anthropic's format.

Can I limit an API key to Zurelay Auto?

Yes. A key limited to zurelay-auto can use whichever model Auto picks for each request.

More models on Zurelay

All models

Stop paying list price.
Start saving today.

Create a key in seconds, point the SDK you already use at Zurelay, and every request costs up to 88% less from the first token.