The right model for each request
Fast, low-cost models for everyday requests, stronger ones for tools, code and structured output, and the strongest when you ask for deep reasoning.
One model name for every request. Zurelay Auto picks the model that fits each one and bills you that model's price, from $0.025 input and $0.125 output per 1M tokens.
zurelay-autofrom openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key="YOUR_ZURELAY_KEY",)stream = client.chat.completions.create( model="zurelay-auto", messages=[{"role": "user", "content": "Explain quicksort in two sentences."}], stream=True,)for chunk in stream: if chunk.choices: print(chunk.choices[0].delta.content or "", end="")Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK, and the Anthropic SDK. Zurelay Auto docs
Pricing
Each request is billed at the price of the model that answered it, with nothing added for the pick. Pay as you go from prepaid credit, with no subscription. Requests that fail are never billed.
| Model | Input | Output |
|---|---|---|
| Everyday requestsShort chat, questions, extraction, summaries. | ||
| DeepSeek V4.1 FlashFirst pick | $0.037 | $0.15 |
| Gemini 3.1 Flash-Lite | $0.10 | $0.59 |
| Gemini 3.5 Flash-Lite | $0.12 | $1.04 |
| GPT-6 Luna | $0.025 | $0.125 |
| Qwen 3.8 Flash | $0.07 | $0.22 |
| Claude Haiku 4.5 | $0.38 | $1.90 |
| Tools, code and structureTools, a JSON schema, code, long prompts, a low or medium reasoning level. | ||
| Claude Sonnet 5.5First pick | $0.80 | $4.00 |
| GPT-6 Sol | $0.50 | $2.50 |
| Gemini 3.7 Flash | $0.32 | $1.59 |
| GPT-6.1 Sol | $0.50 | $2.50 |
| DeepSeek V4 Pro 0813 | $0.23 | $0.62 |
| Deep reasoningA reasoning level of high or more. | ||
| Claude Opus 5.5First pick | $1.40 | $7.00 |
| GPT-6.1 Sol | $0.50 | $2.50 |
| GPT-6 Astra | $2.50 | $12.50 |
| Claude Fable 5.1 | $3.50 | $17.50 |
Prices per 1M tokens. Each tier lists its models in the order Auto tries them; it skips any that can't read what the request sends.
The media it sends, its length, tools, a JSON schema, code, and the reasoning level you set. No model call, so no added delay.
The first model in the request's tier that reads what you sent and whose context window holds it.
Retries and its backup routes, as for any model on Zurelay.
If it still can't answer, the rest of its tier, then the tier above, then the one below. Your smart routing setting doesn't turn that off.
The answer's model field and the x-zurelay-model header name the model that answered, and you pay its price.
Overview
Zurelay Auto is a model you call like any other, zurelay-auto, that answers each request with the model that fits it. Everyday requests go to fast, low-cost models, tools, code and structured output to stronger ones, and requests that ask for deep reasoning to the strongest. You pay the price of the model that answered, and the answer tells you which one it was.
Zurelay Auto reads what each request carries and picks a model for it without an extra model call, so it adds no delay and costs nothing on top. It looks at the images, files, audio or video you send, how long the request is, whether it gives the model tools or a JSON schema, whether it's about code, and the reasoning level you set.
The pick follows three tiers. Short, everyday requests go to fast, low-cost models such as DeepSeek V4.1 Flash and Gemini 3.1 Flash-Lite. Requests with tools, a JSON schema, code, a long prompt or a low or medium reasoning level go to stronger models such as Claude Sonnet 5.5 and GPT-6 Sol. A reasoning level of high or more goes to the strongest, such as Claude Opus 5.5 and GPT-6.1 Sol. It only picks models that read what you send and whose context window holds it.
Each request is billed at the price of the model that answered it: from $0.025 to $3.50 per 1M input tokens and from $0.125 to $17.50 per 1M output tokens, never more than that model costs on its own. The answer's model field and the x-zurelay-model header name the model, and your Requests page shows it too.
Zurelay Auto is built to answer. The model it picks heals like any model on Zurelay: it retries, then uses its backup routes. If that model still can't answer, Auto moves on to its next best models, whatever your smart routing setting. Streaming, tool calls, structured outputs and reasoning levels work as with any model, through /v1/chat/completions or /v1/messages.
Strengths
Fast, low-cost models for everyday requests, stronger ones for tools, code and structured output, and the strongest when you ask for deep reasoning.
Nothing extra for the pick. Each request costs what the model that answered costs on Zurelay, and the answer names it.
Retries and backup routes first, then its next best models. A key's smart routing setting doesn't turn that off.
Set zurelay-auto once. As better or cheaper models arrive, the models Auto picks from change without a line of your code changing.
Use cases
Chat, extraction and summaries next to agent steps and hard questions, each on a model that fits, from one model name.
Start building before choosing a model. See which models Auto picks for your traffic, then pin one if you like.
Everyday requests stay on low-cost models, so a busy app's bill follows the work it does.
Tool calls go to models that call tools reliably, and steps that ask for deep reasoning go to the strongest.
Get started
No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to Zurelay.
Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.
Change the base URL. Everything else in your code stays the same.
https://api.zurelay.com/v1Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.
zurelay-autoWorks with the tools you already use
Compare
$0.037 / $0.15 per 1M tokens
What Auto picks first for everyday requests.
DeepSeek V4.1 Flash API$0.80 / $4.00 per 1M tokens
What Auto picks first for tools, code and structured output.
Claude Sonnet 5.5 API$1.40 / $7.00 per 1M tokens
What Auto picks first when you ask for deep reasoning.
Claude Opus 5.5 APIEach request costs what the model that answered it costs on Zurelay: from $0.025 to $3.50 per 1M input tokens and from $0.125 to $17.50 per 1M output tokens. There's no extra charge for the pick, and requests that fail are never billed.
From what the request carries, without an extra model call. Everyday requests go to fast, low-cost models. Tools, a JSON schema, code, a long prompt or a low or medium reasoning level go to stronger models, and a reasoning level of high or more to the strongest. It only picks models that read the media you send and whose context window holds the request.
The answer's model field and the x-zurelay-model header name it, and x-zurelay-routed-from says zurelay-auto. Your Requests page shows the same.
It heals like any model: it retries and uses that model's backup routes. If the model still can't answer, Auto moves on to its next best models, whatever your smart routing setting.
Yes. Through /v1/messages it picks among Claude models, the ones that take Anthropic's format.
Yes. A key limited to zurelay-auto can use whichever model Auto picks for each request.
Create a key in seconds, point the SDK you already use at Zurelay, and every request costs up to 88% less from the first token.