Platform

Reliability, retries and timeouts

Every request is watched from start to first token. If a path to the model stalls or fails, the request moves to another before anything reaches you.

Self-healing

  • Most models are served over several paths. A request starts on the healthiest one.
  • A path that errors, or is slower to start than it usually is, is dropped for that request and retried elsewhere.
  • Paths that keep failing are moved to the back of the line for everyone, and tried again once they recover.
  • Nothing reaches you until a real token has arrived, so all of this is invisible: you see a slightly slower answer, not an error.

The models page shows each model’s live status (operational, degraded or down), from real requests and our own checks: every chat model is tested at least every 10 minutes.

Timeouts

Limit
Waiting for the first tokenAbout 100 seconds, plus 1 second per 5,000 prompt tokens (up to 30), plus 2 minutes at high reasoning effort
A quiet stream (no data)90 seconds; once a stream has started, keep-alive comments arrive every 15 seconds
A whole streamNo limit while data keeps coming
A whole non-streaming request10 minutes
An image3 minutes per attempt, retried on another path if it stalls

Before the first token nothing is sent, not even headers, and a non-streaming request sends nothing until it’s done. Set your client’s timeout to at least 10 minutes for long reasoning, long outputs and images. The OpenAI SDKs default to 10 minutes; some HTTP clients default to 30 or 60 seconds.

Rate limits

There’s no global rate limit on your account: keys are limited only if you set a rate limit on them. A model that’s overloaded is retried on its other paths; if none answers, you get 503 model_unavailable with retry-after. Back off and retry, or turn on smart routing.

When a model is down

Rarely, every path to a model fails. Without smart routing, the request ends with 503 model_unavailable and retry-after, within about 100 seconds rather than hanging. With smart routing, a close alternative answers instead.

Questions, or something missing? Ask support in your dashboard or email support@zurelay.com.