Chat

Zurelay Auto

One model name, zurelay-auto, that answers each request with the model that fits it. You pay the price of the model that answered, and the answer names it. It works wherever a model ID does, with the same streaming, tools and reasoning levels.

import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key=os.environ["ZURELAY_API_KEY"],
)
response = client.chat.completions.create(
model="zurelay-auto",
messages=[
{
"role": "user",
"content": "Summarize this in one sentence: Zurelay is one API for many models."
}
],
)
print(response.choices[0].message.content)

How it picks

Auto reads what the request carries and picks without calling a model of its own, so it adds no delay and costs nothing on top. Each request lands in one of three tiers:

The requestGoes to
Short chat, questions, extraction, summariesFast, low-cost models
Tools, a JSON schema to follow (response_format of type json_schema), code, a prompt over 32K tokens, or reasoning_effort low or mediumStronger models
reasoning_effort high, xhigh or max, or Anthropic thinking with a large budgetThe strongest models

In its tier, it takes the first model that reads every image, PDF, audio clip or video the request sends, whose context window holds the prompt and its max_tokens, and that takes a reasoning level when you set one. Your reasoning_effort runs at the closest level that model takes.

The models it picks from

Each tier in the order Auto tries them, with their prices per 1M input and output tokens. The list changes as better or cheaper models arrive; your code doesn't have to.

TierModelInput / output per 1M
EverydayDeepSeek V4.1 Flash$0.037 / $0.15
Gemini 3.1 Flash-Lite$0.10 / $0.59
Gemini 3.5 Flash-Lite$0.12 / $1.04
GPT-6 Luna$0.025 / $0.125
Qwen 3.8 Flash$0.07 / $0.22
Claude Haiku 4.5$0.38 / $1.90
Tools, code, structureClaude Sonnet 5.5$0.80 / $4.00
GPT-6 Sol$0.50 / $2.50
Gemini 3.7 Flash$0.32 / $1.59
GPT-6.1 Sol$0.50 / $2.50
DeepSeek V4 Pro 0813$0.23 / $0.62
Deep reasoningClaude Opus 5.5$1.40 / $7.00
GPT-6.1 Sol$0.50 / $2.50
GPT-6 Astra$2.50 / $12.50
Claude Fable 5.1$3.50 / $17.50

What you pay

Each request is billed at the price of the model that answered it, at that model's rates on Zurelay, with nothing added for the pick: today from $0.025 to $3.50 per 1M input tokens and from $0.125 to $17.50 per 1M output tokens. Requests that fail are never billed.

Which model answered

The answer's model field and the x-zurelay-model header name the model that answered, and x-zurelay-routed-from: zurelay-auto says Auto picked it. Requests in the dashboard shows the same.

Response
HTTP/1.1 200 OK
x-zurelay-model: deepseek-v4.1-flash
x-zurelay-routed-from: zurelay-auto
{
"model": "deepseek-v4.1-flash",
"choices": [{ "message": { "role": "assistant", "content": "…" } }],
…
}

When a model can't answer

Auto is built to answer:

  1. The model it picks heals like any model on Zurelay: it retries, then uses that model's backup routes.
  2. If it still can't answer, Auto moves on to the rest of the request's tier, then the tier above, then the one below, up to five models in all.
  3. Only if every one of them fails do you get a 503 with model_unavailable and a retry-after header.

This doesn't depend on your smart routing setting: Auto always does it, and the x-zurelay-fallback header and a request's models list don't change it.

Anthropic format

POST /v1/messages takes zurelay-auto too. There Auto picks among Claude models, the ones that take Anthropic's format, so the Anthropic SDKs and tools built on them work as they do with any Claude model.

API keys

A key limited to zurelay-auto (see key rules) can use whichever model Auto picks. Spend limits and rate limits apply as usual. The short ID auto works too.

Errors

  • 400 unsupported_content: none of its models reads the media the request sends together, such as audio and a PDF in one request. Send them separately, or pick a model that reads both.
  • 400 context_length_exceeded: the request is longer than any of its models can read. Send a shorter one or ask for less output.
  • 503 model_unavailable: every model it tried failed. Retry after the retry-after header.

Want the same model every time?

Use its own ID. Auto's choices can change as models improve, so pin a model where you need its exact behavior.

Questions, or something missing? Ask our support team and we’ll answer by email, usually within a minute.