Chat
Zurelay Auto
One model name, zurelay-auto, that answers each request with the model that fits it. You pay the price of the model that answered, and the answer names it. It works wherever a model ID does, with the same streaming, tools and reasoning levels.
import osfrom openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"],)response = client.chat.completions.create( model="zurelay-auto", messages=[ { "role": "user", "content": "Summarize this in one sentence: Zurelay is one API for many models." } ],)print(response.choices[0].message.content)How it picks
Auto reads what the request carries and picks without calling a model of its own, so it adds no delay and costs nothing on top. Each request lands in one of three tiers:
In its tier, it takes the first model that reads every image, PDF, audio clip or video the request sends, whose context window holds the prompt and its max_tokens, and that takes a reasoning level when you set one. Your reasoning_effort runs at the closest level that model takes.
The models it picks from
Each tier in the order Auto tries them, with their prices per 1M input and output tokens. The list changes as better or cheaper models arrive; your code doesn't have to.
What you pay
Each request is billed at the price of the model that answered it, at that model's rates on Zurelay, with nothing added for the pick: today from $0.025 to $3.50 per 1M input tokens and from $0.125 to $17.50 per 1M output tokens. Requests that fail are never billed.
Which model answered
The answer's model field and the x-zurelay-model header name the model that answered, and x-zurelay-routed-from: zurelay-auto says Auto picked it. Requests in the dashboard shows the same.
HTTP/1.1 200 OKx-zurelay-model: deepseek-v4.1-flashx-zurelay-routed-from: zurelay-auto{ "model": "deepseek-v4.1-flash", "choices": [{ "message": { "role": "assistant", "content": "…" } }], …}When a model can't answer
Auto is built to answer:
- The model it picks heals like any model on Zurelay: it retries, then uses that model's backup routes.
- If it still can't answer, Auto moves on to the rest of the request's tier, then the tier above, then the one below, up to five models in all.
- Only if every one of them fails do you get a
503withmodel_unavailableand aretry-afterheader.
This doesn't depend on your smart routing setting: Auto always does it, and the x-zurelay-fallback header and a request's models list don't change it.
Anthropic format
POST /v1/messages takes zurelay-auto too. There Auto picks among Claude models, the ones that take Anthropic's format, so the Anthropic SDKs and tools built on them work as they do with any Claude model.
API keys
A key limited to zurelay-auto (see key rules) can use whichever model Auto picks. Spend limits and rate limits apply as usual. The short ID auto works too.
Errors
- 400
unsupported_content: none of its models reads the media the request sends together, such as audio and a PDF in one request. Send them separately, or pick a model that reads both. - 400
context_length_exceeded: the request is longer than any of its models can read. Send a shorter one or ask for less output. - 503
model_unavailable: every model it tried failed. Retry after theretry-afterheader.
Want the same model every time?
Questions, or something missing? Ask our support team and we’ll answer by email, usually within a minute.