Platform

Smart routing

When the model a request names is down, smart routing answers with a close alternative instead of an error. Your app keeps working through an outage without a line of retry code.

How it works

  1. Every request first goes to the model you asked for. Most problems end here: if one path to the model stalls or errors, the request is retried on another before you see anything.
  2. Only when the model itself can’t answer (every path failed, or none is available) does smart routing step in.
  3. It tries the alternatives in order, each with the same self-healing, until one answers.
  4. The answer tells you which model replied (model in the body, x-zurelay-model and x-zurelay-routed-from headers), and it’s billed at that model’s price.

When it’s on, the model you asked for gets about half of the usual time before the alternatives take over, so an outage costs seconds, not minutes. A model known to be down is skipped straight away.

Turning it on

WhereHowApplies to
Workspace defaultSettings → Smart routingEvery key that doesn’t set its own, and the Playground
Per keyAPI keys → Edit → Smart routingRequests with that key
Per requesta models list, or the x-zurelay-fallback headerThat request

A key’s setting is one of:

Workspacedefault
Follow the workspace setting (off unless you turned it on).
Offmode
If the model is down, the request fails with model_unavailable, which you can retry.
Automaticmode
Use our list of close alternatives for each model (below): the same family and a similar price first.
Custommode
Your own rules: for each model, up to three alternatives in order. A rule for “any other model” (*) covers the rest.

Per request

A list of alternatives

Pass models next to model: up to three alternatives, tried in order if model is down, whatever the key’s setting or the header. It’s never sent to the model. The OpenAI Python SDK sends it through extra_body.

import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key=os.environ["ZURELAY_API_KEY"],
)
response = client.chat.completions.create(
model="gpt-6-sol",
messages=[
{
"role": "user",
"content": "Summarize this ticket in one line: ..."
}
],
extra_body={
"models": [
"gpt-6.1-sol",
"gpt-6-astra"
]
},
)
print(response.choices[0].message.content)

The header

Headers
x-zurelay-fallback: off # never reroute this request
x-zurelay-fallback: auto # use the automatic alternatives for this request

What applies, first to last: the request’s models, then the header, then the key’s setting, then the workspace default.

What alternatives can answer

  • They read what you sent. A request with a PDF or a video only goes to alternatives that read it. If none does, you get the error rather than an answer that ignores your file.
  • The key may use them. Alternatives outside a key’s allowed models are skipped.
  • They speak the endpoint. On /v1/messages, alternatives are Claude models.
  • The request itself is fine. If the model refused the request (a bad parameter, too long a prompt), it isn’t rerouted: the alternatives would refuse it too.

Automatic alternatives

Live from the catalog: the alternatives each model falls back to in automatic mode.

If this model is downAutomatic smart routing tries, in order
Claude Opus 5.5Claude Opus 5→Claude Opus 4.8→Claude Sonnet 5.5
GPT-6 AstraGPT-5.6 Sol→GPT-6.1 Sol→Claude Opus 5.5
Claude Sonnet 5.5Claude Opus 5.5→Claude Opus 4.8→Claude Haiku 4.5
GPT-6.1 SolGPT-6 Sol→GPT-5.6 Terra→GPT-5.6 Sol
GPT-6 SolGPT-6.1 Sol→GPT-5.6 Terra→GPT-5.6 Sol
GPT-6 LunaGPT-5.6 Luna→Gemini 3.1 Flash-Lite→DeepSeek V4.1 Flash
Grok 4.7GPT-6.1 Sol→Gemini 3.7 Flash→GLM-5.3
MiMo V2.6 ProMiMo V2.6 Flash→DeepSeek V4 Pro 0813→GLM-5.3
MiMo V2.6 FlashMiMo V2.6 Pro→DeepSeek V4.1 Flash→Qwen 3.8 Flash
DeepSeek V4.1 FlashDeepSeek V4 Flash 0731→DeepSeek V4 Pro 0813→Qwen 3.8 Flash
DeepSeek V4 FlashDeepSeek V4 Flash 0731→DeepSeek V4 Pro 0813→Qwen 3.8 Flash
DeepSeek V4 Flash 0731DeepSeek V4.1 Flash→MiMo V2.6 Flash→Qwen 3.8 Flash
DeepSeek V4 Pro 0813DeepSeek V4.1 Flash→GLM-5.3→MiMo V2.6 Pro
Claude Fable 5.1Claude Fable 5→Claude Opus 5.5→Claude Opus 5
Claude Fable 5Claude Fable 5.1→Claude Opus 5.5→Claude Opus 5
Claude Opus 5Claude Opus 4.8→Claude Opus 5.5→Claude Opus 4.7
Claude Opus 4.8Claude Opus 5→Claude Opus 4.7→Claude Opus 5.5
Claude Opus 4.7Claude Opus 4.8→Claude Opus 4.6→Claude Opus 5
Claude Opus 4.6Claude Opus 4.7→Claude Opus 4.8→Claude Opus 5
Claude Haiku 4.5Claude Sonnet 5.5
GPT-5.6 SolGPT-6.1 Sol→GPT-5.6 Terra→GPT-6 Sol
GPT-5.6 TerraGPT-6.1 Sol→GPT-6 Sol→GPT-5.6 Sol
GPT-5.6 LunaGPT-6 Luna→Gemini 3.5 Flash-Lite→DeepSeek V4.1 Flash
Gemini 3.7 FlashGPT-6.1 Sol→Gemini 3.5 Flash-Lite→DeepSeek V4 Flash
Gemini 3.5 Flash-LiteGemini 3.1 Flash-Lite→Gemini 3.7 Flash→GPT-5.6 Luna
Gemini 3.1 Flash-LiteGemini 3.5 Flash-Lite→GPT-6 Luna→DeepSeek V4.1 Flash
GLM-5.3GLM-5.2→DeepSeek V4 Pro 0813→MiniMax M2.7
GLM-5.2GLM-5.3→DeepSeek V4 Pro 0813→MiniMax M2.7
Kimi K3Qwen 3.8 Max→GLM-5.3→DeepSeek V4 Pro 0813
Qwen 3.8 MaxKimi K3→GLM-5.3→DeepSeek V4 Pro 0813
Qwen 3.8 FlashQwen 3.8 Omni Flash→DeepSeek V4.1 Flash→MiMo V2.6 Flash
Qwen 3.8 Omni FlashQwen 3.8 Flash→Gemini 3.1 Flash-Lite→MiMo V2.6 Flash
HY4 PreviewGLM-5.3→DeepSeek V4 Pro 0813→Qwen 3.8 Flash
MiniMax M2.7GLM-5.2→DeepSeek V4 Pro 0813→MiMo V2.6 Pro

Things to know

  • An alternative may cost more or less than the model you asked for. Your request log shows the model that answered and its cost.
  • Streaming works the same: nothing is sent until a model is answering, and the chunks name that model.
  • Images and videos aren’t rerouted: a different image or video model makes a different picture.
Check x-zurelay-routed-from in your logs or monitoring: it’s set only when an alternative answered, so you can see how often a model you depend on was down.

Questions, or something missing? Ask support in your dashboard or email support@zurelay.com.