Platform
Reliability, retries and timeouts
Every request is watched from start to first token. If a path to the model stalls or fails, the request moves to another before anything reaches you.
Self-healing
- Most models are served over several paths. A request starts on the healthiest one.
- A path that errors, or is slower to start than it usually is, is dropped for that request and retried elsewhere.
- Paths that keep failing are moved to the back of the line for everyone, and tried again once they recover.
- Nothing reaches you until a real token has arrived, so all of this is invisible: you see a slightly slower answer, not an error.
The models page shows each model’s live status (operational, degraded or down), from real requests and our own checks: every chat model is tested at least every 10 minutes.
Timeouts
Before the first token nothing is sent, not even headers, and a non-streaming request sends nothing until it’s done. Set your client’s timeout to at least 10 minutes for long reasoning, long outputs and images. The OpenAI SDKs default to 10 minutes; some HTTP clients default to 30 or 60 seconds.
Rate limits
There’s no global rate limit on your account: keys are limited only if you set a rate limit on them. A model that’s overloaded is retried on its other paths; if none answers, you get 503 model_unavailable with retry-after. Back off and retry, or turn on smart routing.
When a model is down
Rarely, every path to a model fails. Without smart routing, the request ends with 503 model_unavailable and retry-after, within about 100 seconds rather than hanging. With smart routing, a close alternative answers instead.
Questions, or something missing? Ask support in your dashboard or email support@zurelay.com.