# zurelay documentation > One OpenAI- and Anthropic-compatible API for Claude, GPT, Gemini, Grok, DeepSeek, Nano Banana, GPT Image and Seedance, at a fraction of each lab's price. Every page of https://zurelay.com/docs, as Markdown. The index is at https://zurelay.com/llms.txt. --- # zurelay API documentation > One OpenAI- and Anthropic-compatible API for Claude, GPT, Gemini, Grok, DeepSeek, Nano Banana, GPT Image and Seedance, at a fraction of each lab's price. Source: https://zurelay.com/docs zurelay is one API for the AI models you already use, at a fraction of each lab’s own price. Chat with Anthropic, OpenAI, xAI, Xiaomi, DeepSeek, Google, Z.ai, Moonshot AI, Alibaba, Tencent and MiniMax models, make and edit pictures with Nano Banana and GPT Image, and make videos with Seedance, all with one key and one balance. It speaks OpenAI’s API and Anthropic’s, so the SDKs, frameworks and tools you have work as they are: change the base URL and the key, keep your code. ## Base URLs | For | Base URL | | --- | --- | | OpenAI SDKs, and anything OpenAI-compatible | `https://api.zurelay.com/v1` | | Anthropic SDKs and Claude Code | `https://api.zurelay.com` | Send your key as `Authorization: Bearer zr_live_…` (or `x-api-key`, as Anthropic’s SDKs do). Create keys in the dashboard under [API keys](https://zurelay.com/app/keys). ## Your first request **Python** (`main.py`): ```python import os from openai import OpenAI client = OpenAI( base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"], ) response = client.chat.completions.create( model="claude-sonnet-5-5", messages=[ { "role": "user", "content": "Explain what an API gateway does in one paragraph." } ], ) print(response.choices[0].message.content) ``` **Node.js** (`index.mjs`): ```javascript import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.zurelay.com/v1", apiKey: process.env.ZURELAY_API_KEY, }); const response = await client.chat.completions.create({ "model": "claude-sonnet-5-5", "messages": [ { "role": "user", "content": "Explain what an API gateway does in one paragraph." } ] }); console.log(response.choices[0].message.content); ``` **cURL**: ```bash curl https://api.zurelay.com/v1/chat/completions \ -H "Authorization: Bearer $ZURELAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "claude-sonnet-5-5", "messages": [ { "role": "user", "content": "Explain what an API gateway does in one paragraph." } ] }' ``` ## What you can do - [Chat completions](https://zurelay.com/docs/chat): Messages, streaming, tools, structured output and reasoning. - [Images, PDFs, audio and video](https://zurelay.com/docs/multimodal): Send files to the models that read them. - [Image generation](https://zurelay.com/docs/images): Nano Banana and GPT Image, at 1K, 2K and 4K. - [Image editing](https://zurelay.com/docs/image-editing): Edit a picture or combine up to 16 reference images. - [Video generation](https://zurelay.com/docs/video): Seedance from a prompt, a first frame or reference images. - [Smart routing](https://zurelay.com/docs/smart-routing): A close alternative answers when a model is down. ## Endpoints | Method | Path | What it does | | --- | --- | --- | | POST | `/v1/chat/completions` | Chat with any text model (OpenAI format) | | POST | `/v1/messages` | Chat with Claude models (Anthropic format) | | POST | `/v1/messages/count_tokens` | Count a Messages request’s input tokens | | POST | `/v1/images/generations` | Make images, optionally from reference images | | POST | `/v1/images/edits` | Edit images (OpenAI’s multipart form, or JSON) | | POST | `/v1/videos` | Start a video | | GET | `/v1/videos/{id}` | A video’s status, and its link when done | | GET | `/v1/videos/{id}/content` | Download a finished video | | GET | `/v1/videos` | Your recent videos | | GET | `/v1/models` | Every model, with what it reads | ## How it works Every request is served by the model you name, at its published quality, through a fleet that heals itself: if one path to a model stalls or errors, the request is retried on another before anything reaches you, so a hiccup costs you a moment instead of an error. When a model is fully down, [smart routing](https://zurelay.com/docs/smart-routing) can answer with its closest alternative. You pay as you go from a prepaid balance, at the per-token, per-image or per-second prices on the [models page](https://zurelay.com/docs/models). Failed requests are free. ## For AI assistants and agents Building with ChatGPT, Claude, Cursor or an agent? Every page has **Copy page** at the top, which copies it as Markdown to paste into any AI, and the docs are published in the formats agents read: - [llms.txt](https://zurelay.com/llms.txt): what zurelay is, how to call it, and a link to every page. - [llms-full.txt](https://zurelay.com/llms-full.txt): all the docs in one Markdown file, with today’s models and prices. - Any page as Markdown: add `.md` to its address, as in [/docs/quickstart.md](https://zurelay.com/docs/quickstart.md). --- # Quickstart > Make your first request in two minutes: create a key, point an SDK at zurelay, send a message. Source: https://zurelay.com/docs/quickstart Three steps: a key, two settings, one request. ## 1. Create an account and a key 1. [Create an account](https://zurelay.com/signup) (no password: we email you a code) and add credit under [Billing](https://zurelay.com/app/billing). 2. Open [API keys](https://zurelay.com/app/keys), name a key and create it. Copy it right away: you see the full key once. ## 2. Set the key and base URL Keep the key out of your code. In your shell, or in a .env file your app loads: .env: ``` ZURELAY_API_KEY=zr_live_your_key_here ``` Then install an SDK if you don’t have one: `pip install openai` for Python, `npm install openai` for Node. Anthropic’s SDKs work too (see [Anthropic Messages](https://zurelay.com/docs/anthropic)). ## 3. Send a request **Python** (`main.py`): ```python import os from openai import OpenAI client = OpenAI( base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"], ) response = client.chat.completions.create( model="gpt-6-sol", messages=[ { "role": "system", "content": "You are a concise assistant." }, { "role": "user", "content": "Give me three names for a coffee shop on the moon." } ], ) print(response.choices[0].message.content) ``` **Node.js** (`index.mjs`): ```javascript import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.zurelay.com/v1", apiKey: process.env.ZURELAY_API_KEY, }); const response = await client.chat.completions.create({ "model": "gpt-6-sol", "messages": [ { "role": "system", "content": "You are a concise assistant." }, { "role": "user", "content": "Give me three names for a coffee shop on the moon." } ] }); console.log(response.choices[0].message.content); ``` **cURL**: ```bash curl https://api.zurelay.com/v1/chat/completions \ -H "Authorization: Bearer $ZURELAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-6-sol", "messages": [ { "role": "system", "content": "You are a concise assistant." }, { "role": "user", "content": "Give me three names for a coffee shop on the moon." } ] }' ``` The answer comes back in OpenAI’s format, with the model that answered and the tokens used: Response: ```json { "id": "chatcmpl-...", "object": "chat.completion", "created": 1790870400, "model": "gpt-6-sol", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "1. Lunar Grind\n2. ..." }, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 31, "completion_tokens": 24, "total_tokens": 55 } } ``` ## Stream it Add `stream: true` to get the answer as it’s written, the same way OpenAI streams it. **Python** (`main.py`): ```python import os from openai import OpenAI client = OpenAI( base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"], ) stream = client.chat.completions.create( model="gemini-3.7-flash", messages=[ { "role": "user", "content": "Write a haiku about latency." } ], stream=True, ) for chunk in stream: if chunk.choices and chunk.choices[0].delta.content: print(chunk.choices[0].delta.content, end="", flush=True) ``` **Node.js** (`index.mjs`): ```javascript import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.zurelay.com/v1", apiKey: process.env.ZURELAY_API_KEY, }); const stream = await client.chat.completions.create({ "model": "gemini-3.7-flash", "messages": [ { "role": "user", "content": "Write a haiku about latency." } ], "stream": true }); for await (const chunk of stream) { process.stdout.write(chunk.choices[0]?.delta?.content ?? ""); } ``` **cURL**: ```bash curl https://api.zurelay.com/v1/chat/completions \ -H "Authorization: Bearer $ZURELAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "gemini-3.7-flash", "messages": [ { "role": "user", "content": "Write a haiku about latency." } ], "stream": true }' ``` ## Next - [Pick a model](https://zurelay.com/docs/models): IDs, prices and what each model reads. - [Send a picture or a PDF](https://zurelay.com/docs/multimodal): Files go in the message, as data or a link. - [Make an image](https://zurelay.com/docs/images): One call to /v1/images/generations. - [Use your tools](https://zurelay.com/docs/integrations): Claude Code, Cursor, LangChain and more. --- # API keys and authentication > How to send your API key, keep it safe, and what each key can be limited to. Source: https://zurelay.com/docs/authentication Every request carries an API key. Keys belong to a workspace and spend its balance; each can be held to its own budgets, models and addresses. ## Sending the key Either header works on every endpoint: Headers: ``` Authorization: Bearer zr_live_your_key_here # or, as Anthropic's SDKs send it: x-api-key: zr_live_your_key_here ``` Keys start with `zr_live_`. You see the full key once, when you create it; we only keep a hash, so nobody (us included) can show it to you again. Lost it? Revoke it and make another. ## Keeping keys safe - Keep keys in a secret manager or environment variables, never in code or a repository. - Call zurelay from your server. A key in a browser or a mobile app can be read by anyone who has the app. - Use one key per app or environment (say, `production` and `staging`), so you can revoke one without touching the others. - Give each key only what it needs: a budget, the models it calls, the addresses it runs on. See [Key limits and rules](https://zurelay.com/docs/key-rules). - If a key leaks, revoke it under API keys. Requests with it stop within 15 seconds. ## What a key can be held to - `Monthly budget` (USD): Spend allowed per calendar month (UTC). - `Daily budget` (USD): Spend allowed per day (UTC). - `Rate limit` (requests / minute): Requests per minute; more are refused with 429. - `Expiry` (date): After it, the key stops working. - `Allowed models` (list): Only these models; anything else is refused with 403. - `Allowed IP addresses` (list): Only requests from these addresses or ranges. - `Smart routing` (setting): What happens when a model is down: see [Smart routing](https://zurelay.com/docs/smart-routing). ## Authentication errors | Status | Code | Means | | --- | --- | --- | | 401 | `missing_api_key` | No key in the request. | | 401 | `invalid_api_key` | The key doesn’t exist or was revoked. | | 401 | `expired_api_key` | The key passed its expiry date. | | 403 | `ip_not_allowed` | The key only works from other addresses. | | 403 | `model_not_allowed` | The key isn’t allowed to use this model. | | 402 | `insufficient_quota` | The workspace balance is empty. | > **Tip:** Every chat, image and video response carries an `x-request-id` header. Include it when you write to support and we can find the request in seconds. --- # Models > Every model on zurelay with its ID, what it reads (images, PDFs, audio, video), its context window and its price. Source: https://zurelay.com/docs/models Use the ID in the `model` field. Prices update live; every model has its own page with details, benchmarks and examples. ## Chat models Prices are per 1M input / output tokens. “Reads” is what you can send besides text: see [Images, PDFs, audio and video](https://zurelay.com/docs/multimodal). | Model | ID | Context | Reads | Price per 1M input / output tokens | | --- | --- | --- | --- | --- | | [Claude Opus 5.5](https://zurelay.com/models/claude-opus-5-5) | `claude-opus-5-5` | 1M | Text, Images, PDFs | $1.40 / $7.00 | | [GPT-6 Astra](https://zurelay.com/models/gpt-6-astra) | `gpt-6-astra` | 1.05M | Text, Images, PDFs | $2.50 / $12.50 | | [Claude Sonnet 5.5](https://zurelay.com/models/claude-sonnet-5-5) | `claude-sonnet-5-5` | 1M | Text, Images, PDFs | $0.80 / $4.00 | | [GPT-6.1 Sol](https://zurelay.com/models/gpt-6.1-sol) | `gpt-6.1-sol` | 1.05M | Text, Images, PDFs | $0.50 / $2.50 | | [GPT-6 Sol](https://zurelay.com/models/gpt-6-sol) | `gpt-6-sol` | 1.05M | Text, Images, PDFs | $0.50 / $2.50 | | [GPT-6 Luna](https://zurelay.com/models/gpt-6-luna) | `gpt-6-luna` | 1.05M | Text, Images, PDFs | $0.025 / $0.125 | | [Grok 4.7](https://zurelay.com/models/grok-4.7) | `grok-4.7` | 500K | Text, Images | $0.70 / $2.10 | | [MiMo V2.6 Pro](https://zurelay.com/models/mimo-v2.6-pro) | `mimo-v2.6-pro` | 1M | Text, Images, Audio | $0.18 / $0.36 | | [MiMo V2.6 Flash](https://zurelay.com/models/mimo-v2.6-flash) | `mimo-v2.6-flash` | 1M | Text, Images, Audio | $0.06 / $0.12 | | [DeepSeek V4.1 Flash](https://zurelay.com/models/deepseek-v4.1-flash) | `deepseek-v4.1-flash` | 1M | Text, Images | $0.037 / $0.15 | | [DeepSeek V4 Flash](https://zurelay.com/models/deepseek-v4-flash) | `deepseek-v4-flash` | 1M | Text, Images | $0.045 / $0.18 | | [DeepSeek V4 Flash 0731](https://zurelay.com/models/deepseek-v4-flash-0731) | `deepseek-v4-flash-0731` | 1M | Text | $0.022 / $0.06 | | [DeepSeek V4 Pro 0813](https://zurelay.com/models/deepseek-v4-pro-0813) | `deepseek-v4-pro-0813` | 1M | Text | $0.23 / $0.62 | | [Claude Fable 5.1](https://zurelay.com/models/claude-fable-5-1) | `claude-fable-5-1` | 1M | Text, Images, PDFs | $3.50 / $17.50 | | [Claude Fable 5](https://zurelay.com/models/claude-fable-5) | `claude-fable-5` | 1M | Text, Images | $3.50 / $17.50 | | [Claude Opus 5](https://zurelay.com/models/claude-opus-5) | `claude-opus-5` | 1M | Text, Images, PDFs | $1.75 / $8.75 | | [Claude Opus 4.8](https://zurelay.com/models/claude-opus-4-8) | `claude-opus-4-8` | 1M | Text, Images, PDFs | $1.75 / $8.75 | | [Claude Opus 4.7](https://zurelay.com/models/claude-opus-4-7) | `claude-opus-4-7` | 1M | Text, Images | $1.78 / $8.93 | | [Claude Opus 4.6](https://zurelay.com/models/claude-opus-4-6) | `claude-opus-4-6` | 1M | Text, Images | $1.75 / $8.75 | | [Claude Haiku 4.5](https://zurelay.com/models/claude-haiku-4-5) | `claude-haiku-4-5` | 200K | Text, Images, PDFs | $0.38 / $1.90 | | [GPT-5.6 Sol](https://zurelay.com/models/gpt-5.6-sol) | `gpt-5.6-sol` | 1.05M | Text, Images, PDFs | $1.14 / $5.70 | | [GPT-5.6 Terra](https://zurelay.com/models/gpt-5.6-terra) | `gpt-5.6-terra` | 1.05M | Text, Images, PDFs | $0.54 / $3.26 | | [GPT-5.6 Luna](https://zurelay.com/models/gpt-5.6-luna) | `gpt-5.6-luna` | 1.05M | Text, Images, PDFs | $0.08 / $0.49 | | [Gemini 3.7 Flash](https://zurelay.com/models/gemini-3.7-flash) | `gemini-3.7-flash` | 1M | Text, Images, PDFs, Audio, Video | $0.32 / $1.59 | | [Gemini 3.5 Flash-Lite](https://zurelay.com/models/gemini-3.5-flash-lite) | `gemini-3.5-flash-lite` | 1M | Text, Images, PDFs, Audio, Video | $0.12 / $1.04 | | [Gemini 3.1 Flash-Lite](https://zurelay.com/models/gemini-3.1-flash-lite) | `gemini-3.1-flash-lite` | 1M | Text, Images, PDFs, Audio, Video | $0.10 / $0.59 | | [GLM-5.3](https://zurelay.com/models/glm-5.3) | `glm-5.3` | 1M | Text | $0.38 / $1.21 | | [GLM-5.2](https://zurelay.com/models/glm-5.2) | `glm-5.2` | 1M | Text | $0.38 / $1.21 | | [Kimi K3](https://zurelay.com/models/kimi-k3) | `kimi-k3` | 1M | Text, Images | $1.27 / $6.37 | | [Qwen 3.8 Max](https://zurelay.com/models/qwen3.8-max) | `qwen3.8-max` | 1M | Text, Images, PDFs, Video | $1.20 / $3.60 | | [Qwen 3.8 Flash](https://zurelay.com/models/qwen3.8-flash) | `qwen3.8-flash` | 1M | Text, Images, PDFs, Video | $0.07 / $0.22 | | [Qwen 3.8 Omni Flash](https://zurelay.com/models/qwen3.8-omni-flash) | `qwen3.8-omni-flash` | 1M | Text, Images, Video | $0.075 / $0.235 | | [HY4 Preview](https://zurelay.com/models/hy4-preview) | `hy4-preview` | 1M | Text, Images | $0.40 / $1.20 | | [MiniMax M2.7](https://zurelay.com/models/minimax-m2.7) | `minimax-m2.7` | 205K | Text | $0.13 / $0.51 | ## Image models Prices are per image. See [Image generation](https://zurelay.com/docs/images) and [Image editing](https://zurelay.com/docs/image-editing). | Model | ID | Takes | Price per image | | --- | --- | --- | --- | | [Nano Banana Pro](https://zurelay.com/models/nano-banana-pro) | `nano-banana-pro` | Prompt, up to 14 reference images | from $0.035 | | [Nano Banana 2](https://zurelay.com/models/nano-banana-2) | `nano-banana-2` | Prompt, up to 14 reference images | from $0.025 | | [GPT Image 2.5 Flare](https://zurelay.com/models/gpt-image-2.5-flare) | `gpt-image-2.5-flare` | Prompt, up to 16 reference images | $0.02 | | [GPT Image 2.5 Sunburst](https://zurelay.com/models/gpt-image-2.5-sunburst) | `gpt-image-2.5-sunburst` | Prompt, up to 16 reference images | $0.02 | | [GPT Image 2](https://zurelay.com/models/gpt-image-2) | `gpt-image-2` | Prompt, up to 16 reference images | $0.02 | ## Video models Prices are per second of video, by resolution. See [Video generation](https://zurelay.com/docs/video). | Model | ID | Takes | Price per second | | --- | --- | --- | --- | | [Seedance 2.5](https://zurelay.com/models/seedance-2.5) | `seedance-2.5` | Prompt, frames or up to 9 reference images | from $0.14/s | | [Seedance 2.0](https://zurelay.com/models/seedance-2.0) | `seedance-2.0` | Prompt, frames or up to 9 reference images | from $0.097/s | | [Seedance 2.0 Fast](https://zurelay.com/models/seedance-2.0-fast) | `seedance-2.0-fast` | Prompt, frames or up to 9 reference images | from $0.074/s | | [Seedance 2.0 Mini](https://zurelay.com/models/seedance-2.0-mini) | `seedance-2.0-mini` | Prompt, frames or up to 9 reference images | from $0.047/s | ## Listing models from the API `GET /v1/models` Returns every model in OpenAI’s list format, with two extra fields: `input_modalities` (what the model reads) and, for image models, `max_input_images`. `GET /v1/models/{id}` returns one. Neither needs credit. Response: ```json { "object": "list", "data": [ { "id": "gemini-3.7-flash", "object": "model", "created": 1790000000, "owned_by": "google", "input_modalities": ["text", "image", "file", "audio", "video"] }, { "id": "nano-banana-pro", "object": "model", "created": 1790000000, "owned_by": "google", "input_modalities": ["text", "image"], "max_input_images": 14 } ] } ``` ### Retired models When we retire a model, requests for it get a clear `model_not_found` error naming the models page. We announce retirements ahead of time where we can. --- # Chat completions > POST /v1/chat/completions: messages, streaming, tool calling, structured output, reasoning effort and usage. Source: https://zurelay.com/docs/chat OpenAI’s Chat Completions API, for every text model: the same request and response, the same streaming, the same SDKs. `POST /v1/chat/completions` ## Request - `model` (string, required): A model ID from the [models page](https://zurelay.com/docs/models), such as `claude-opus-5-5` or `gpt-6-sol`. - `messages` (array, required): The conversation: `system` (or `developer`), `user`, `assistant` and `tool` messages. `content` is a string, or an array of parts to add [images, PDFs, audio or video](https://zurelay.com/docs/multimodal). - `stream` (boolean): Send the answer as server-sent events while it’s written. See Streaming below. - `stream_options` (object): `{ "include_usage": true }` adds a final chunk with token usage. - `max_completion_tokens` (integer): The most tokens to generate, reasoning included. `max_tokens` works too. - `temperature` (number): 0 to 2. Lower is more focused, higher more varied. Some reasoning models ignore it. - `top_p` (number): Nucleus sampling, 0 to 1. Change this or temperature, not both. - `stop` (string | array): Up to 4 sequences where generation stops. - `tools` (array): Functions the model may call. See Tool calling below. - `tool_choice` (string | object): `auto`, `none`, `required`, or a specific function. - `response_format` (object): JSON output: `{ "type": "json_object" }`, or `json_schema` with a schema. See Structured output below. - `reasoning_effort` (string): For reasoning models: `low`, `medium` or `high` (some take `minimal` or `xhigh`). Higher thinks longer and costs more output tokens. - `seed` (integer): Best-effort reproducibility, for models that support it. - `models` (array): Smart routing for this request: alternatives to try, in order, if the model is down. See [Smart routing](https://zurelay.com/docs/smart-routing). > **Note:** Other OpenAI parameters pass through to the model unchanged, except `user` and `metadata`, which stay with us. Which ones a model honors (tools, JSON schemas, reasoning) is up to the model, as it is with its own lab. With the OpenAI Python SDK, send `models` as `extra_body={"models": [...]}`. ## Response Response: ```json { "id": "chatcmpl-AbC123...", "object": "chat.completion", "created": 1790870400, "model": "claude-opus-5-5", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Here's the summary..." }, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 1204, "completion_tokens": 312, "total_tokens": 1516, "prompt_tokens_details": { "cached_tokens": 1024 }, "completion_tokens_details": { "reasoning_tokens": 128 } } } ``` - `model` is the model that answered. It’s the one you asked for, unless smart routing answered with an alternative. - Reasoning models may add `reasoning_content` to the message: the model’s thinking, when it shares it. - `usage` is what you’re billed for. `cached_tokens` are input tokens read from the model’s prompt cache, billed at the cached price where a model has one. Every response also has these headers: | Header | Value | | --- | --- | | `x-request-id` | The request’s ID, as it appears in your request log. Quote it to support. | | `x-zurelay-model` | The model that answered. | | `x-zurelay-routed-from` | Only after smart routing: the model you asked for. | ## Streaming With `stream: true`, the answer arrives as server-sent events in OpenAI’s chunk format, ending with `data: [DONE]`. Every SDK reads these for you. **Python** (`main.py`): ```python import os from openai import OpenAI client = OpenAI( base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"], ) stream = client.chat.completions.create( model="claude-sonnet-5-5", messages=[ { "role": "user", "content": "Write a limerick about databases." } ], stream=True, stream_options={ "include_usage": True }, ) for chunk in stream: if chunk.choices and chunk.choices[0].delta.content: print(chunk.choices[0].delta.content, end="", flush=True) ``` **Node.js** (`index.mjs`): ```javascript import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.zurelay.com/v1", apiKey: process.env.ZURELAY_API_KEY, }); const stream = await client.chat.completions.create({ "model": "claude-sonnet-5-5", "messages": [ { "role": "user", "content": "Write a limerick about databases." } ], "stream": true, "stream_options": { "include_usage": true } }); for await (const chunk of stream) { process.stdout.write(chunk.choices[0]?.delta?.content ?? ""); } ``` **cURL**: ```bash curl https://api.zurelay.com/v1/chat/completions \ -H "Authorization: Bearer $ZURELAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "claude-sonnet-5-5", "messages": [ { "role": "user", "content": "Write a limerick about databases." } ], "stream": true, "stream_options": { "include_usage": true } }' ``` - Nothing is sent until the model has produced its first token, so a retry or a switch to another path stays invisible to you. - Once the stream has started, quiet stretches (a model thinking between steps) are filled with SSE comment lines (`: keep-alive`) every 15 seconds. Clients ignore them; they keep proxies from closing the connection. Before the first token nothing is sent, not even headers: allow up to about 100 seconds, or 4 minutes at high reasoning effort. - With `include_usage`, the last chunk before `[DONE]` has an empty `choices` list and the `usage`. ## Tool calling Describe functions in `tools`; when the model wants one, it answers with `tool_calls` instead of text. Run the function, send the result back as a `tool` message with the call’s ID, and the model continues. **Python** (`main.py`): ```python import os from openai import OpenAI client = OpenAI( base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"], ) response = client.chat.completions.create( model="gpt-6-sol", messages=[ { "role": "user", "content": "What's the weather in Lisbon right now?" } ], tools=[ { "type": "function", "function": { "name": "get_weather", "description": "Current weather for a city", "parameters": { "type": "object", "properties": { "city": { "type": "string" } }, "required": [ "city" ] } } } ], ) print(response.choices[0].message.content) ``` **Node.js** (`index.mjs`): ```javascript import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.zurelay.com/v1", apiKey: process.env.ZURELAY_API_KEY, }); const response = await client.chat.completions.create({ "model": "gpt-6-sol", "messages": [ { "role": "user", "content": "What's the weather in Lisbon right now?" } ], "tools": [ { "type": "function", "function": { "name": "get_weather", "description": "Current weather for a city", "parameters": { "type": "object", "properties": { "city": { "type": "string" } }, "required": [ "city" ] } } } ] }); console.log(response.choices[0].message.content); ``` **cURL**: ```bash curl https://api.zurelay.com/v1/chat/completions \ -H "Authorization: Bearer $ZURELAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-6-sol", "messages": [ { "role": "user", "content": "What'\''s the weather in Lisbon right now?" } ], "tools": [ { "type": "function", "function": { "name": "get_weather", "description": "Current weather for a city", "parameters": { "type": "object", "properties": { "city": { "type": "string" } }, "required": [ "city" ] } } } ] }' ``` ## Structured output Ask for JSON with `response_format`. A `json_schema` holds models that support it to your schema; `json_object` asks for any valid JSON (say so in your prompt too). **Python** (`main.py`): ```python import os from openai import OpenAI client = OpenAI( base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"], ) response = client.chat.completions.create( model="gpt-6.1-sol", messages=[ { "role": "user", "content": "Extract the people: 'Ana met Ben and Chloe in Porto.'" } ], response_format={ "type": "json_schema", "json_schema": { "name": "people", "schema": { "type": "object", "properties": { "names": { "type": "array", "items": { "type": "string" } } }, "required": [ "names" ] } } }, ) print(response.choices[0].message.content) ``` **Node.js** (`index.mjs`): ```javascript import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.zurelay.com/v1", apiKey: process.env.ZURELAY_API_KEY, }); const response = await client.chat.completions.create({ "model": "gpt-6.1-sol", "messages": [ { "role": "user", "content": "Extract the people: 'Ana met Ben and Chloe in Porto.'" } ], "response_format": { "type": "json_schema", "json_schema": { "name": "people", "schema": { "type": "object", "properties": { "names": { "type": "array", "items": { "type": "string" } } }, "required": [ "names" ] } } } }); console.log(response.choices[0].message.content); ``` **cURL**: ```bash curl https://api.zurelay.com/v1/chat/completions \ -H "Authorization: Bearer $ZURELAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-6.1-sol", "messages": [ { "role": "user", "content": "Extract the people: '\''Ana met Ben and Chloe in Porto.'\''" } ], "response_format": { "type": "json_schema", "json_schema": { "name": "people", "schema": { "type": "object", "properties": { "names": { "type": "array", "items": { "type": "string" } } }, "required": [ "names" ] } } } }' ``` ## Reasoning Reasoning models think before they answer. Set `reasoning_effort` to trade speed and cost for depth. Thinking counts as output tokens (`completion_tokens_details.reasoning_tokens`) and can take minutes at high effort: give your client a generous timeout (see [Reliability and timeouts](https://zurelay.com/docs/reliability)). ### Claude’s extended thinking Through this endpoint, Claude models take `reasoning_effort` too. For Anthropic’s own `thinking` parameter and thinking blocks, use [Anthropic Messages](https://zurelay.com/docs/anthropic). --- # Images, PDFs, audio and video input > Send pictures, documents, recordings and clips to the models that read them, as data or links, in OpenAI or Anthropic format. Source: https://zurelay.com/docs/multimodal Put pictures, PDFs, audio and video in a message for the models that read them. Send the file itself (base64) or a link: we fetch links ourselves, so any public `https` link to a file up to 20 MB works. ## At a glance | To send | /v1/chat/completions (every model) | /v1/messages (Claude) | Formats | | --- | --- | --- | --- | | A picture | `image_url` part | `image` block | PNG, JPEG, WebP, GIF | | A PDF | `file` part | `document` block | PDF | | A recording | `input_audio` part | Claude doesn’t hear audio | WAV, MP3 | | A video | `video_url` part | Claude doesn’t watch video | MP4, MOV, WebM | | Text, code, CSV, Word, Excel | The text itself, in the message | A `text` block, or a plain-text `document` | Text | Every part takes a public `https` link or the file itself as a base64 data URL, such as `data:application/pdf;base64,JVBERi0...`. Several files can go in one message, in any mix the model reads. ## Which models read what Every chat model reads text. This table is live: it shows what each model reads today. A model given something it doesn’t read refuses it straight away with `unsupported_content`, so it never answers about a file it didn’t see. | Model | ID | Images | PDFs | Audio | Video | | --- | --- | --- | --- | --- | --- | | Claude Opus 5.5 | `claude-opus-5-5` | Yes | Yes | No | No | | GPT-6 Astra | `gpt-6-astra` | Yes | Yes | No | No | | Claude Sonnet 5.5 | `claude-sonnet-5-5` | Yes | Yes | No | No | | GPT-6.1 Sol | `gpt-6.1-sol` | Yes | Yes | No | No | | GPT-6 Sol | `gpt-6-sol` | Yes | Yes | No | No | | GPT-6 Luna | `gpt-6-luna` | Yes | Yes | No | No | | Grok 4.7 | `grok-4.7` | Yes | No | No | No | | MiMo V2.6 Pro | `mimo-v2.6-pro` | Yes | No | Yes | No | | MiMo V2.6 Flash | `mimo-v2.6-flash` | Yes | No | Yes | No | | DeepSeek V4.1 Flash | `deepseek-v4.1-flash` | Yes | No | No | No | | DeepSeek V4 Flash | `deepseek-v4-flash` | Yes | No | No | No | | DeepSeek V4 Flash 0731 | `deepseek-v4-flash-0731` | No | No | No | No | | DeepSeek V4 Pro 0813 | `deepseek-v4-pro-0813` | No | No | No | No | | Claude Fable 5.1 | `claude-fable-5-1` | Yes | Yes | No | No | | Claude Fable 5 | `claude-fable-5` | Yes | No | No | No | | Claude Opus 5 | `claude-opus-5` | Yes | Yes | No | No | | Claude Opus 4.8 | `claude-opus-4-8` | Yes | Yes | No | No | | Claude Opus 4.7 | `claude-opus-4-7` | Yes | No | No | No | | Claude Opus 4.6 | `claude-opus-4-6` | Yes | No | No | No | | Claude Haiku 4.5 | `claude-haiku-4-5` | Yes | Yes | No | No | | GPT-5.6 Sol | `gpt-5.6-sol` | Yes | Yes | No | No | | GPT-5.6 Terra | `gpt-5.6-terra` | Yes | Yes | No | No | | GPT-5.6 Luna | `gpt-5.6-luna` | Yes | Yes | No | No | | Gemini 3.7 Flash | `gemini-3.7-flash` | Yes | Yes | Yes | Yes | | Gemini 3.5 Flash-Lite | `gemini-3.5-flash-lite` | Yes | Yes | Yes | Yes | | Gemini 3.1 Flash-Lite | `gemini-3.1-flash-lite` | Yes | Yes | Yes | Yes | | GLM-5.3 | `glm-5.3` | No | No | No | No | | GLM-5.2 | `glm-5.2` | No | No | No | No | | Kimi K3 | `kimi-k3` | Yes | No | No | No | | Qwen 3.8 Max | `qwen3.8-max` | Yes | Yes | No | Yes | | Qwen 3.8 Flash | `qwen3.8-flash` | Yes | Yes | No | Yes | | Qwen 3.8 Omni Flash | `qwen3.8-omni-flash` | Yes | No | No | Yes | | HY4 Preview | `hy4-preview` | Yes | No | No | No | | MiniMax M2.7 | `minimax-m2.7` | No | No | No | No | ## Images Add an `image_url` part with a link or a data URL. PNG, JPEG, WebP and GIF. Several images in one message work on every model that reads images. **Python** (`main.py`): ```python import os from openai import OpenAI client = OpenAI( base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"], ) response = client.chat.completions.create( model="gemini-3.7-flash", messages=[ { "role": "user", "content": [ { "type": "text", "text": "What's in this picture? One sentence." }, { "type": "image_url", "image_url": { "url": "https://zurelay.com/examples/gpt-image-2.5-flare/watch-800.webp" } } ] } ], ) print(response.choices[0].message.content) ``` **Node.js** (`index.mjs`): ```javascript import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.zurelay.com/v1", apiKey: process.env.ZURELAY_API_KEY, }); const response = await client.chat.completions.create({ "model": "gemini-3.7-flash", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "What's in this picture? One sentence." }, { "type": "image_url", "image_url": { "url": "https://zurelay.com/examples/gpt-image-2.5-flare/watch-800.webp" } } ] } ] }); console.log(response.choices[0].message.content); ``` **cURL**: ```bash curl https://api.zurelay.com/v1/chat/completions \ -H "Authorization: Bearer $ZURELAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "gemini-3.7-flash", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "What'\''s in this picture? One sentence." }, { "type": "image_url", "image_url": { "url": "https://zurelay.com/examples/gpt-image-2.5-flare/watch-800.webp" } } ] } ] }' ``` To send a file from disk, base64-encode it into a data URL: **Python** (`main.py`): ```python import base64, os from openai import OpenAI client = OpenAI(base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"]) with open("receipt.jpg", "rb") as f: data_url = "data:image/jpeg;base64," + base64.b64encode(f.read()).decode() response = client.chat.completions.create( model="claude-sonnet-5-5", messages=[{ "role": "user", "content": [ {"type": "text", "text": "What's the total on this receipt?"}, {"type": "image_url", "image_url": {"url": data_url}}, ], }], ) print(response.choices[0].message.content) ``` **Node.js** (`index.mjs`): ```javascript import { readFile } from "node:fs/promises"; import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.zurelay.com/v1", apiKey: process.env.ZURELAY_API_KEY }); const dataUrl = "data:image/jpeg;base64," + (await readFile("receipt.jpg")).toString("base64"); const response = await client.chat.completions.create({ model: "claude-sonnet-5-5", messages: [{ role: "user", content: [ { type: "text", text: "What's the total on this receipt?" }, { type: "image_url", image_url: { url: dataUrl } }, ], }], }); console.log(response.choices[0].message.content); ``` **cURL**: ```bash # The request goes in a file: a base64 image is too long for a command-line argument. cat > request.json < request.json <\n{notes}\n\n\n\n{sales}\n", }], ) print(response.choices[0].message.content) ``` ## Audio For models that hear (see the table): an `input_audio` part with the base64 audio and its format, `wav` or `mp3`. `input_audio` takes the audio itself; for a link, use a `file` part with the link in `file_data`. Recordings made in a browser are usually WebM, which is read as video: convert them to MP3 or WAV for models that only hear. **Python** (`main.py`): ```python import base64, os from openai import OpenAI client = OpenAI(base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"]) with open("call.mp3", "rb") as f: audio = base64.b64encode(f.read()).decode() response = client.chat.completions.create( model="gemini-3.7-flash", messages=[{ "role": "user", "content": [ {"type": "text", "text": "Transcribe this call, then list the action items."}, {"type": "input_audio", "input_audio": {"data": audio, "format": "mp3"}}, ], }], ) print(response.choices[0].message.content) ``` **Node.js** (`index.mjs`): ```javascript import { readFile } from "node:fs/promises"; import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.zurelay.com/v1", apiKey: process.env.ZURELAY_API_KEY }); const audio = (await readFile("call.mp3")).toString("base64"); const response = await client.chat.completions.create({ model: "gemini-3.7-flash", messages: [{ role: "user", content: [ { type: "text", text: "Transcribe this call, then list the action items." }, { type: "input_audio", input_audio: { data: audio, format: "mp3" } }, ], }], }); console.log(response.choices[0].message.content); ``` ## Video For models that watch video: a `video_url` part with a link or a data URL (a `file` part works too). MP4, MOV and WebM. Models with sound read the soundtrack as well, so you can ask what’s said. **Python** (`main.py`): ```python import os from openai import OpenAI client = OpenAI( base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"], ) response = client.chat.completions.create( model="gemini-3.7-flash", messages=[ { "role": "user", "content": [ { "type": "text", "text": "Describe what happens, and transcribe anything that's said." }, { "type": "video_url", "video_url": { "url": "https://example.com/clip.mp4" } } ] } ], ) print(response.choices[0].message.content) ``` **Node.js** (`index.mjs`): ```javascript import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.zurelay.com/v1", apiKey: process.env.ZURELAY_API_KEY, }); const response = await client.chat.completions.create({ "model": "gemini-3.7-flash", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe what happens, and transcribe anything that's said." }, { "type": "video_url", "video_url": { "url": "https://example.com/clip.mp4" } } ] } ] }); console.log(response.choices[0].message.content); ``` **cURL**: ```bash curl https://api.zurelay.com/v1/chat/completions \ -H "Authorization: Bearer $ZURELAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "gemini-3.7-flash", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe what happens, and transcribe anything that'\''s said." }, { "type": "video_url", "video_url": { "url": "https://example.com/clip.mp4" } } ] } ] }' ``` A clip from disk goes the same way, as a data URL: Python: ```python with open("clip.mp4", "rb") as f: video = "data:video/mp4;base64," + base64.b64encode(f.read()).decode() content = [ {"type": "text", "text": "What happens in this clip?"}, {"type": "video_url", "video_url": {"url": video}}, ] ``` > **One shape in, the right shape out:** Labs take media in different shapes (a video can be a `video_url`, an `image_url` or a `file` part, depending on whose docs you read). Send any of them: we hand each model its media in the shape it reads. ## Several files in one message Add a part for each file, in the order you want them read, and refer to them in your text by order or by name (“the second image”, “contract.pdf”). Pictures, PDFs, audio and video can be mixed as long as the model reads each kind. Text parts can sit between files to label them. ## Links - Links must be public `https` URLs. We follow up to three redirects and wait up to 20 seconds. - A link that doesn’t download is refused straight away with `invalid_url` and the status it answered, before any model is called. - Signed links (S3, Cloud Storage, our own image links) work while they’re valid. ## Anthropic format On [/v1/messages](https://zurelay.com/docs/anthropic), Claude models take Anthropic’s blocks: `image` and `document`, with a `base64` or `url` source. A few Claude models read PDFs only through `/v1/chat/completions`; send one to `/v1/messages` and the error says where to send it instead. Content blocks: ```json [ { "type": "image", "source": { "type": "base64", "media_type": "image/png", "data": "iVBORw0KGgo..." } }, { "type": "document", "source": { "type": "url", "url": "https://example.com/report.pdf" } }, { "type": "document", "source": { "type": "base64", "media_type": "application/pdf", "data": "JVBERi0..." } }, { "type": "text", "text": "Compare the chart with the report's conclusion." } ] ``` ## Limits | Limit | Value | | --- | --- | | One file | 20 MB | | All media in one request | 30 MB | | Request body (base64 adds about a third) | 25 MB | | Media parts per request | 100 | | Links per request | 20 | For bigger files, send a link: it doesn’t count toward the request body. Files uploaded to a lab’s own Files API (`file_id`) can’t be used; send the file or a link instead. ## Errors | Status | Error | When | | --- | --- | --- | | 400 | `unsupported_content` | The model doesn’t read that kind of file. Pick one that does from the table above. | | 400 | `invalid_url` | A link didn’t download. The message says why (its status, a timeout, not a media file). | | 400 | `invalid_request_error` | A part isn’t a data URL or an https link, is empty, or is a type models don’t read (Word, plain text). | | 413 | `invalid_request_error` | A file over 20 MB, more than 30 MB of media, or a request body over 25 MB. | The first two are the error’s `code`; the others are its `type`, with a message that says what to change. All of them come back before any model is called, so they cost nothing. See [Errors](https://zurelay.com/docs/errors) for the rest. ## How media are billed Models turn media into input tokens, billed at the model’s input price. Each lab counts its own way; these are typical figures we measured: | | Claude | GPT | Gemini | | --- | --- | --- | --- | | A 1024 × 1024 picture | about 1,400 tokens | about 1,250 to 1,650 | about 1,100 | | A large photo (12 MP) | up to about 4,800 | scales with pixels | about 1,100 to 1,200 | | One PDF page | about 1,600 to 2,500 | the page’s text | about 560 | | One second of audio | not read | not read | about 25 | | One second of video | not read | not read | about 100 to 300 | Your request log shows exactly what each request counted. Shrinking pictures to the size you need is the easiest saving: most models read a 1024 px picture as well as a 4000 px one. > **Tip:** Sending the same picture or document in many requests? Put it early in the conversation and keep that prefix the same: models with prompt caching bill repeated input at their cached price. --- # Anthropic Messages and Claude Code > POST /v1/messages for Claude models with the Anthropic SDKs, Claude Code and anything else that speaks Anthropic's API. Source: https://zurelay.com/docs/anthropic Claude models also speak Anthropic’s Messages API, so Anthropic’s SDKs, Claude Code and anything built for them work with a new base URL and key. `POST /v1/messages` Base URL `https://api.zurelay.com` (the SDKs add `/v1/messages`). Send the key as `x-api-key` or `Authorization: Bearer`, with `anthropic-version: 2023-06-01`. Only Claude models are served here; other models use [/v1/chat/completions](https://zurelay.com/docs/chat). ## With Anthropic’s SDKs **Python** (`main.py`): ```python import os import anthropic client = anthropic.Anthropic( base_url="https://api.zurelay.com", api_key=os.environ["ZURELAY_API_KEY"], ) message = client.messages.create( model="claude-opus-5-5", max_tokens=1024, system="You are a careful code reviewer.", messages=[{"role": "user", "content": "Review this function: def add(a, b): return a - b"}], ) print(message.content[0].text) ``` **Node.js** (`index.mjs`): ```javascript import Anthropic from "@anthropic-ai/sdk"; const client = new Anthropic({ baseURL: "https://api.zurelay.com", apiKey: process.env.ZURELAY_API_KEY, }); const message = await client.messages.create({ model: "claude-opus-5-5", max_tokens: 1024, system: "You are a careful code reviewer.", messages: [{ role: "user", content: "Review this function: def add(a, b): return a - b" }], }); console.log(message.content[0].text); ``` **cURL**: ```bash curl https://api.zurelay.com/v1/messages \ -H "x-api-key: $ZURELAY_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "content-type: application/json" \ -d '{ "model": "claude-opus-5-5", "max_tokens": 1024, "messages": [{ "role": "user", "content": "Hello, Claude" }] }' ``` Everything in Anthropic’s Messages API passes through (except `metadata`, which stays with us): `system` as a string or blocks (with `cache_control` for prompt caching), `tools`, `thinking`, `stop_sequences`, `temperature`, streaming with the same events, and `anthropic-beta` headers. Images and PDFs go in as `image` and `document` blocks: see [Images, PDFs, audio and video](https://zurelay.com/docs/multimodal). ## Claude Code Point Claude Code at zurelay with five environment variables (the base URL, your key and the three model aliases), then run it as usual: Shell: ```bash export ANTHROPIC_BASE_URL="https://api.zurelay.com" export ANTHROPIC_AUTH_TOKEN="$ZURELAY_API_KEY" export ANTHROPIC_DEFAULT_OPUS_MODEL="claude-opus-5-5" export ANTHROPIC_DEFAULT_SONNET_MODEL="claude-sonnet-5-5" export ANTHROPIC_DEFAULT_HAIKU_MODEL="claude-sonnet-5-5" claude ``` To keep it, put the same values in the `env` block of `~/.claude/settings.json`. The full guide, and guides for Cursor, Cline and others, are under [Integrations](https://zurelay.com/docs/integrations). ## Counting tokens `POST /v1/messages/count_tokens` Takes the same body as `/v1/messages` and returns `{ "input_tokens": 1234 }`, an estimate of what the request would count. Free. ## Errors Errors come back in Anthropic’s shape (`{ "type": "error", "error": { "type", "message" } }`), with the same meanings as on the [errors page](https://zurelay.com/docs/errors). They carry no `code`: use the HTTP status and `error.type`. > **Note:** A few Claude models read PDFs only through `/v1/chat/completions`. Send one to `/v1/messages` and the error says which endpoint to use. The table on [Images, PDFs, audio and video](https://zurelay.com/docs/multimodal) shows what each model reads. - **Tip:** Anthropic’s token counter and ours can differ slightly; usage in the response is what’s billed. --- # Image generation > POST /v1/images/generations: prompts, resolutions, aspect ratios, pixel sizes, response formats and prices. Source: https://zurelay.com/docs/images Make pictures from a prompt with OpenAI’s Images API: Nano Banana and GPT Image, one call each. `POST /v1/images/generations` | Model | ID | Sizes and prices | Reference images | | --- | --- | --- | --- | | [Nano Banana Pro](https://zurelay.com/models/nano-banana-pro) | `nano-banana-pro` | 1K and 2K $0.035 · 4K $0.065 | up to 14 | | [Nano Banana 2](https://zurelay.com/models/nano-banana-2) | `nano-banana-2` | 1K $0.025 · 2K $0.034 · 4K $0.058 | up to 14 | | [GPT Image 2.5 Flare](https://zurelay.com/models/gpt-image-2.5-flare) | `gpt-image-2.5-flare` | $0.02 an image | up to 16 | | [GPT Image 2.5 Sunburst](https://zurelay.com/models/gpt-image-2.5-sunburst) | `gpt-image-2.5-sunburst` | $0.02 an image | up to 16 | | [GPT Image 2](https://zurelay.com/models/gpt-image-2) | `gpt-image-2` | $0.02 an image | up to 16 | ## Make an image **Python** (`main.py`): ```python import os from openai import OpenAI client = OpenAI(base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"]) result = client.images.generate( model="nano-banana-pro", prompt="A red fox curled up asleep in fresh snow at dawn, soft golden light, 85mm photo", size="2K", extra_body={ "aspect_ratio": "3:2" }, ) print(result.data[0].url) ``` **Node.js** (`index.mjs`): ```javascript import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.zurelay.com/v1", apiKey: process.env.ZURELAY_API_KEY }); const result = await client.images.generate({ "model": "nano-banana-pro", "prompt": "A red fox curled up asleep in fresh snow at dawn, soft golden light, 85mm photo", "size": "2K", "aspect_ratio": "3:2" }); console.log(result.data[0].url); ``` **cURL**: ```bash curl https://api.zurelay.com/v1/images/generations \ -H "Authorization: Bearer $ZURELAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "nano-banana-pro", "prompt": "A red fox curled up asleep in fresh snow at dawn, soft golden light, 85mm photo", "size": "2K", "aspect_ratio": "3:2" }' ``` Response: ```json { "created": 1790870400, "data": [ { "url": "https://...signed link, valid for 24 hours..." } ] } ``` ## Parameters - `model` (string, required): An image model ID. - `prompt` (string, required): What to make, up to 32,000 characters. - `n` (integer): How many images, 1 to 4. They’re made at once; you pay for the ones delivered. - `size` (string): Nano Banana: `1K`, `2K` or `4K` (or pixels like `2048x1152`, read as the nearest resolution and shape). GPT Image: pixels, such as `1024x1024` or `3840x2160`. Leave it out for the default. - `aspect_ratio` (string): Nano Banana: `1:1`, `2:3`, `3:2`, `3:4`, `4:3`, `4:5`, `5:4`, `9:16`, `16:9` or `21:9`. - `response_format` (string): `url` (default): a link valid for 24 hours. `b64_json`: the image itself, base64-encoded. - `images` (array): Reference images to edit or work from. See [Image editing](https://zurelay.com/docs/image-editing). - `quality, background, output_format` (string): GPT Image only: see the [GPT Image guide](https://zurelay.com/docs/gpt-image). ## Sizes Nano Banana is priced by resolution and makes exactly the size asked for: an image that comes back smaller is retried, never delivered. At 1K: | Aspect ratio | 1K | 2K | 4K | | --- | --- | --- | --- | | 1:1 | 1024 × 1024 | 2048 × 2048 | 4096 × 4096 | | 2:3 | 848 × 1264 | 1696 × 2528 | 3392 × 5056 | | 3:2 | 1264 × 848 | 2528 × 1696 | 5056 × 3392 | | 3:4 | 896 × 1200 | 1792 × 2400 | 3584 × 4800 | | 4:3 | 1200 × 896 | 2400 × 1792 | 4800 × 3584 | | 4:5 | 928 × 1152 | 1856 × 2304 | 3712 × 4608 | | 5:4 | 1152 × 928 | 2304 × 1856 | 4608 × 3712 | | 9:16 | 768 × 1376 | 1536 × 2752 | 3072 × 5504 | | 16:9 | 1376 × 768 | 2752 × 1536 | 5504 × 3072 | | 21:9 | 1584 × 672 | 3168 × 1344 | 6336 × 2688 | GPT Image costs the same at every size: `1024x1024`, `1536x1024`, `1024x1536`, `2048x2048`, `3840x2160`, `2160x3840`. ## Links and history - Links in `url` are signed and work for 24 hours. Download what you want to keep. - Images returned as `url` are also kept for 30 days in your dashboard’s [Library](https://zurelay.com/app/library), where you can download them, delete them, or use them as reference images. Images returned as `b64_json` aren’t stored. - Prompts sent through the API aren’t stored with the images. ## Timing Most images arrive in 10 to 60 seconds, longer at 4K. If an attempt stalls, the request is retried on another path automatically, so give your client a timeout of 10 minutes. ## When an image isn’t made If the model declines a prompt (say, for its content policy), you get `400 image_not_generated` with the model’s reason, and you’re not charged. With `n` above 1, you get the images that were made and pay only for those. - [Edit and combine images](https://zurelay.com/docs/image-editing): Reference images, up to 14 or 16 a request. - [Nano Banana guide](https://zurelay.com/docs/nano-banana): Prompting, text in images, references. --- # Image editing and reference images > Edit a picture or combine several: POST /v1/images/edits (OpenAI's multipart form or JSON) and reference images on generations. Source: https://zurelay.com/docs/image-editing Give an image model pictures to work from: edit a photo, put a product in a new scene, keep a character the same across shots, or blend several images into one. ## Two ways to send images | Endpoint | Takes | Good for | | --- | --- | --- | | `POST /v1/images/edits` | OpenAI’s multipart form (image files), or JSON | OpenAI’s SDKs: client.images.edit(...) | | `POST /v1/images/generations` | JSON with an `images` list | Any HTTP client; images as links or data URLs | Both take the same models, sizes and prices as [generation](https://zurelay.com/docs/images). An edit costs the same as a new image of that size. ## How many images each model takes | Model | ID | Sizes and prices | Reference images | | --- | --- | --- | --- | | [Nano Banana Pro](https://zurelay.com/models/nano-banana-pro) | `nano-banana-pro` | 1K and 2K $0.035 · 4K $0.065 | up to 14 | | [Nano Banana 2](https://zurelay.com/models/nano-banana-2) | `nano-banana-2` | 1K $0.025 · 2K $0.034 · 4K $0.058 | up to 14 | | [GPT Image 2.5 Flare](https://zurelay.com/models/gpt-image-2.5-flare) | `gpt-image-2.5-flare` | $0.02 an image | up to 16 | | [GPT Image 2.5 Sunburst](https://zurelay.com/models/gpt-image-2.5-sunburst) | `gpt-image-2.5-sunburst` | $0.02 an image | up to 16 | | [GPT Image 2](https://zurelay.com/models/gpt-image-2) | `gpt-image-2` | $0.02 an image | up to 16 | ## With OpenAI’s SDKs **Python** (`main.py`): ```python import os from openai import OpenAI client = OpenAI(base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"]) result = client.images.edit( model="gpt-image-2", image=[open("product.png", "rb"), open("background.jpg", "rb")], prompt="Put the bottle from the first image on the marble counter in the second. Keep its label exactly.", ) print(result.data[0].url) ``` **Node.js** (`index.mjs`): ```javascript import fs from "node:fs"; import OpenAI, { toFile } from "openai"; const client = new OpenAI({ baseURL: "https://api.zurelay.com/v1", apiKey: process.env.ZURELAY_API_KEY }); const result = await client.images.edit({ model: "gpt-image-2", image: [ await toFile(fs.createReadStream("product.png"), null, { type: "image/png" }), await toFile(fs.createReadStream("background.jpg"), null, { type: "image/jpeg" }), ], prompt: "Put the bottle from the first image on the marble counter in the second. Keep its label exactly.", }); console.log(result.data[0].url); ``` **cURL**: ```bash curl https://api.zurelay.com/v1/images/edits \ -H "Authorization: Bearer $ZURELAY_API_KEY" \ -F model=gpt-image-2 \ -F "image[]=@product.png" \ -F "image[]=@background.jpg" \ -F prompt="Put the bottle from the first image on the marble counter in the second." ``` ## As JSON On either endpoint, list the pictures in `images` as links or data URLs (`image` with one picture works too): **Python** (`main.py`): ```python import os import requests response = requests.post( "https://api.zurelay.com/v1/images/generations", headers={"Authorization": f"Bearer {os.environ['ZURELAY_API_KEY']}"}, json={ "model": "nano-banana-pro", "prompt": "Make the jacket in the first image deep green, and give the model the pose from the second.", "images": [ "https://example.com/model.jpg", "data:image/png;base64,iVBORw0KGgo..." ], "size": "2K", "aspect_ratio": "4:5" }, ) response.raise_for_status() print(response.json()) ``` **Node.js** (`index.mjs`): ```javascript const response = await fetch("https://api.zurelay.com/v1/images/generations", { method: "POST", headers: { Authorization: `Bearer ${process.env.ZURELAY_API_KEY}`, "Content-Type": "application/json", }, body: JSON.stringify({ "model": "nano-banana-pro", "prompt": "Make the jacket in the first image deep green, and give the model the pose from the second.", "images": [ "https://example.com/model.jpg", "data:image/png;base64,iVBORw0KGgo..." ], "size": "2K", "aspect_ratio": "4:5" }), }); console.log(await response.json()); ``` **cURL**: ```bash curl https://api.zurelay.com/v1/images/generations \ -H "Authorization: Bearer $ZURELAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "nano-banana-pro", "prompt": "Make the jacket in the first image deep green, and give the model the pose from the second.", "images": [ "https://example.com/model.jpg", "data:image/png;base64,iVBORw0KGgo..." ], "size": "2K", "aspect_ratio": "4:5" }' ``` ## Rules for input images - PNG, JPEG or WebP, up to 20 MB each. - Links must be public `https` URLs; we download them ourselves. One that doesn’t download is refused with `invalid_image` before anything is made. - Refer to images by their order in the prompt: “the first image”, “the second image”. - Masks aren’t supported. Describe the change instead (“replace only the sky”); these models edit just what you name. - Over a model’s limit, the request is refused with a clear error naming the limit. ## What to try | Goal | Prompt pattern | | --- | --- | | Edit one thing | “Change the car to matte black. Keep everything else the same.” | | New background | “Place the person from the image on a beach at sunset, matching the light.” | | Product in a scene | “Put the product from the first image on the shelf in the second.” | | Same character, new shot | “The same woman as in the references, now riding a bike in Paris.” | | Style transfer | “Redraw the first image in the style of the second.” | | Combine many | “A group photo of the six people in the images, standing on stairs.” | --- # Nano Banana > Everything about Nano Banana Pro and Nano Banana 2 on zurelay: sizes, prices, up to 14 reference images, editing, text in images and prompting. Source: https://zurelay.com/docs/nano-banana Google’s Nano Banana models make and edit pictures from text and up to 14 reference images, with crisp text in images and faithful edits. ## The models | | Nano Banana Pro | Nano Banana 2 | | --- | --- | --- | | ID | `nano-banana-pro` | `nano-banana-2` | | Best for | The highest quality: detail, text, faithful edits | Speed and price, with very good quality | | Prices | 1K and 2K $0.035, 4K $0.065 | 1K $0.025, 2K $0.034, 4K $0.058 | | Resolutions | 1K, 2K, 4K | 1K, 2K, 4K | | Aspect ratios | 10, from 21:9 to 9:16 | 10, from 21:9 to 9:16 | | Reference images | up to 14 | up to 14 | ## Make an image **Python** (`main.py`): ```python import os from openai import OpenAI client = OpenAI(base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"]) result = client.images.generate( model="nano-banana-pro", prompt="A minimalist poster for a jazz night: a brass saxophone in silhouette, the words 'Blue Hour' in tall serif letters, deep navy and gold", size="2K", extra_body={ "aspect_ratio": "2:3" }, ) print(result.data[0].url) ``` **Node.js** (`index.mjs`): ```javascript import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.zurelay.com/v1", apiKey: process.env.ZURELAY_API_KEY }); const result = await client.images.generate({ "model": "nano-banana-pro", "prompt": "A minimalist poster for a jazz night: a brass saxophone in silhouette, the words 'Blue Hour' in tall serif letters, deep navy and gold", "size": "2K", "aspect_ratio": "2:3" }); console.log(result.data[0].url); ``` **cURL**: ```bash curl https://api.zurelay.com/v1/images/generations \ -H "Authorization: Bearer $ZURELAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "nano-banana-pro", "prompt": "A minimalist poster for a jazz night: a brass saxophone in silhouette, the words '\''Blue Hour'\'' in tall serif letters, deep navy and gold", "size": "2K", "aspect_ratio": "2:3" }' ``` Sizes: `1K`, `2K` or `4K`, with any of ten aspect ratios (the full table is on [Image generation](https://zurelay.com/docs/images)). Without `size`, you get a 1K image. ## Edit and combine with references Pass up to 14 images in `images` (or as files on `/v1/images/edits`) and say what to do with them. Nano Banana keeps faces, products and logos recognizable across edits. **Python** (`main.py`): ```python import os import requests response = requests.post( "https://api.zurelay.com/v1/images/generations", headers={"Authorization": f"Bearer {os.environ['ZURELAY_API_KEY']}"}, json={ "model": "nano-banana-pro", "prompt": "Studio product shot: the sneaker from the first image on the concrete block from the second, lit like the third image. Keep the sneaker's logo and colors exactly.", "images": [ "https://example.com/sneaker.png", "https://example.com/concrete.jpg", "https://example.com/lighting-ref.jpg" ], "size": "2K", "aspect_ratio": "1:1" }, ) response.raise_for_status() print(response.json()) ``` **Node.js** (`index.mjs`): ```javascript const response = await fetch("https://api.zurelay.com/v1/images/generations", { method: "POST", headers: { Authorization: `Bearer ${process.env.ZURELAY_API_KEY}`, "Content-Type": "application/json", }, body: JSON.stringify({ "model": "nano-banana-pro", "prompt": "Studio product shot: the sneaker from the first image on the concrete block from the second, lit like the third image. Keep the sneaker's logo and colors exactly.", "images": [ "https://example.com/sneaker.png", "https://example.com/concrete.jpg", "https://example.com/lighting-ref.jpg" ], "size": "2K", "aspect_ratio": "1:1" }), }); console.log(await response.json()); ``` **cURL**: ```bash curl https://api.zurelay.com/v1/images/generations \ -H "Authorization: Bearer $ZURELAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "nano-banana-pro", "prompt": "Studio product shot: the sneaker from the first image on the concrete block from the second, lit like the third image. Keep the sneaker'\''s logo and colors exactly.", "images": [ "https://example.com/sneaker.png", "https://example.com/concrete.jpg", "https://example.com/lighting-ref.jpg" ], "size": "2K", "aspect_ratio": "1:1" }' ``` ## Prompting ### Describe a scene, not keywords Write the way you’d brief a photographer: subject, setting, composition, light, lens or medium, mood. “A weathered fisherman mending a net on a pier at dusk, low warm light, 50mm, shallow depth of field” beats “fisherman, pier, dusk, 4k, masterpiece”. ### Text in images Put the exact words in quotes and say where they go and how they look: “the headline ‘Grand Opening’ in bold white sans-serif across the top”. Keep each piece of text short. ### Edits - Name the change and what must stay: “Change only the wall color to sage green. Keep the furniture, light and camera angle.” - Edit in steps for big changes: one call per change, feeding each result into the next. - Refer to references by order and role: “the person from the first image, the outfit from the second”. ### Shape and size Pick the aspect ratio for where the image goes (9:16 for stories, 16:9 for headers, 4:5 for feeds). Draft at 1K, then make the keeper at 2K or 4K: the price is per image, by resolution. ## Good to know - Images take 10 to 60 seconds, longer at 4K. Set a client timeout of 10 minutes, which leaves room for a retry. - Prompts the model declines return `image_not_generated` with its reason, free of charge. - Every image carries the lab’s invisible SynthID watermark, as it does from Google. - Links last 24 hours; the Library in your dashboard keeps images returned as links for 30 days. > **Try it without code:** The dashboard [Playground](https://zurelay.com/app/playground) has an image studio: drop in reference images, pick a size, and see the exact code for the request. Model details and examples: [Nano Banana Pro](https://zurelay.com/models/nano-banana-pro), [Nano Banana 2](https://zurelay.com/models/nano-banana-2). - **Pricing:** 1K and 2K $0.035, 4K $0.065 for Pro; 1K $0.025, 2K $0.034, 4K $0.058 for Nano Banana 2. --- # GPT Image > GPT Image 2, 2.5 Flare and 2.5 Sunburst on zurelay: sizes, quality, editing with up to 16 images. Source: https://zurelay.com/docs/gpt-image OpenAI’s GPT Image models: sharp detail, strong text rendering, and editing with up to 16 images, at one price for every size. ## The models | Model | ID | Price per image | Reference images | | --- | --- | --- | --- | | GPT Image 2.5 Flare | `gpt-image-2.5-flare` | $0.02 | up to 16 | | GPT Image 2.5 Sunburst | `gpt-image-2.5-sunburst` | $0.02 | up to 16 | | GPT Image 2 | `gpt-image-2` | $0.02 | up to 16 | GPT Image 2.5 Flare is the quick one; 2.5 Sunburst puts quality first; GPT Image 2 stays for workflows tuned to it. ## Make an image **Python** (`main.py`): ```python import os from openai import OpenAI client = OpenAI(base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"]) result = client.images.generate( model="gpt-image-2.5-sunburst", prompt="An isometric illustration of a tiny bakery at night, warm window light, a cat on the awning", size="1536x1024", extra_body={ "quality": "high" }, ) print(result.data[0].url) ``` **Node.js** (`index.mjs`): ```javascript import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.zurelay.com/v1", apiKey: process.env.ZURELAY_API_KEY }); const result = await client.images.generate({ "model": "gpt-image-2.5-sunburst", "prompt": "An isometric illustration of a tiny bakery at night, warm window light, a cat on the awning", "size": "1536x1024", "quality": "high" }); console.log(result.data[0].url); ``` **cURL**: ```bash curl https://api.zurelay.com/v1/images/generations \ -H "Authorization: Bearer $ZURELAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-image-2.5-sunburst", "prompt": "An isometric illustration of a tiny bakery at night, warm window light, a cat on the awning", "size": "1536x1024", "quality": "high" }' ``` ## Parameters - `size` (string): `1024x1024, 1536x1024, 1024x1536, 2048x2048, 3840x2160, 2160x3840`, or `auto`. Same price at every size. - `quality` (string): `low`, `medium`, `high` or `auto`. - `background` (string): `transparent`, `opaque` or `auto`. Transparent needs PNG or WebP output. - `output_format` (string): `png`, `jpeg` or `webp`. - `n` (integer): 1 to 4 images. ## Edit with up to 16 images Use `client.images.edit` with one or more files, or JSON with `images`. Masks aren’t supported; describe the area to change. Python: ```python result = client.images.edit( model="gpt-image-2.5-flare", image=[open(f"tile-{i}.png", "rb") for i in range(1, 9)], prompt="Arrange the eight product photos into a clean 4 x 2 catalog grid on white.", ) ``` More on sending images: [Image editing and reference images](https://zurelay.com/docs/image-editing). --- # Video generation > POST /v1/videos and polling: prompts, length, resolution, aspect ratio, audio, start and end frames, reference images and prices. Source: https://zurelay.com/docs/video Make videos with Seedance from a prompt, from a first frame (and optionally a last one), or from up to 9 reference images. Videos take a few minutes, so you start one and check back. `POST /v1/videos` | Model | ID | Price per second | Length | | --- | --- | --- | --- | | [Seedance 2.5](https://zurelay.com/models/seedance-2.5) | `seedance-2.5` | 480p $0.14 · 720p $0.305 · 1080p $0.75 | 4 to 30 s | | [Seedance 2.0](https://zurelay.com/models/seedance-2.0) | `seedance-2.0` | 480p $0.097 · 720p $0.207 | 4 to 15 s | | [Seedance 2.0 Fast](https://zurelay.com/models/seedance-2.0-fast) | `seedance-2.0-fast` | 480p $0.074 · 720p $0.159 | 4 to 15 s | | [Seedance 2.0 Mini](https://zurelay.com/models/seedance-2.0-mini) | `seedance-2.0-mini` | 480p $0.047 · 720p $0.10 | 4 to 15 s | ## Start, poll, download **Python** (`video.py`): ```python import os, time, requests API = "https://api.zurelay.com/v1" HEADERS = {"Authorization": f"Bearer {os.environ['ZURELAY_API_KEY']}"} # 1. Start the video response = requests.post(f"{API}/videos", headers=HEADERS, json={ "model": "seedance-2.0", "prompt": "A paper boat drifting down a rain-soaked street at night, neon reflections, slow dolly shot", "seconds": 5, "resolution": "720p", "aspect_ratio": "16:9", }) response.raise_for_status() # a 4xx says what to fix job = response.json() # 2. Poll until it's done (usually 2 to 6 minutes) while job["status"] in ("queued", "in_progress"): time.sleep(10) job = requests.get(f"{API}/videos/{job['id']}", headers=HEADERS).json() if job["status"] == "failed": raise SystemExit(job["error"]["message"]) # 3. Download the MP4 video = requests.get(f"{API}/videos/{job['id']}/content", headers=HEADERS) open("boat.mp4", "wb").write(video.content) ``` **Node.js** (`video.mjs`): ```javascript import { writeFile } from "node:fs/promises"; const API = "https://api.zurelay.com/v1"; const headers = { Authorization: `Bearer ${process.env.ZURELAY_API_KEY}`, "Content-Type": "application/json", }; // 1. Start the video let job = await (await fetch(`${API}/videos`, { method: "POST", headers, body: JSON.stringify({ model: "seedance-2.0", prompt: "A paper boat drifting down a rain-soaked street at night, neon reflections, slow dolly shot", seconds: 5, resolution: "720p", aspect_ratio: "16:9", }), })).json(); // 2. Poll until it's done (usually 2 to 6 minutes) while (job.status === "queued" || job.status === "in_progress") { await new Promise((resolve) => setTimeout(resolve, 10_000)); job = await (await fetch(`${API}/videos/${job.id}`, { headers })).json(); } if (job.status === "failed") throw new Error(job.error.message); // 3. Download the MP4 const video = await fetch(`${API}/videos/${job.id}/content`, { headers }); await writeFile("boat.mp4", Buffer.from(await video.arrayBuffer())); ``` **cURL**: ```bash # 1. Start the video curl https://api.zurelay.com/v1/videos \ -H "Authorization: Bearer $ZURELAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "seedance-2.0", "prompt": "A paper boat drifting down a rain-soaked street at night", "seconds": 5, "resolution": "720p" }' # 2. Check on it (repeat until "status" is "completed") curl https://api.zurelay.com/v1/videos/video_abc123 -H "Authorization: Bearer $ZURELAY_API_KEY" # 3. Download it curl -L https://api.zurelay.com/v1/videos/video_abc123/content \ -H "Authorization: Bearer $ZURELAY_API_KEY" -o boat.mp4 ``` ## Parameters - `model` (string, required): A video model ID. - `prompt` (string): What happens, up to 8,000 characters. Required unless you give a first frame or reference images. - `seconds` (integer): Length, within the model’s range (see the table). Default 5. - `resolution` (string): `480p`, `720p` or `1080p`, as the model offers. Default 720p. - `aspect_ratio` (string): `16:9`, `9:16`, `1:1`, `4:3`, `3:4`, `21:9`, or `adaptive` (the first frame’s shape; the default when you give one). Otherwise the default is `16:9`. - `size` (string): Instead of resolution and ratio: pixels like `1280x720`. - `generate_audio` (boolean): Sound and speech made with the video. Default `true`. (`audio` works too.) - `seed` (integer): The same seed and settings give similar results. - `first_frame` (string): An image the video starts on: a link or a data URL. - `last_frame` (string): An image it ends on. Needs a first frame. - `reference_images` (array): Up to 9 images of characters, products, places or a style to use. Not with frames. ## Starting from images ### First and last frame Animate a still: the video opens exactly on `first_frame`. Add `last_frame` and it lands on that image, so you control both ends of the shot. Request body: ```json { "model": "seedance-2.0", "first_frame": "https://example.com/storefront-day.jpg", "last_frame": "https://example.com/storefront-night.jpg", "prompt": "Time passes from afternoon to night; lights come on inside the shop.", "seconds": 6, "resolution": "720p" } ``` ### Reference images Give up to 9 images and mention them in the prompt by order (`@Image1`, `@Image2`…): the people, products and places in them appear in the video. Request body: ```json { "model": "seedance-2.5", "reference_images": [ "https://example.com/chef.jpg", "https://example.com/kitchen.jpg", "https://example.com/dish.jpg" ], "prompt": "The chef from @Image1 plates the dish from @Image3 in the kitchen from @Image2, close-up, warm light.", "seconds": 8, "resolution": "720p", "aspect_ratio": "9:16" } ``` ### Rules for images - PNG, JPEG or WebP, as a public `https` link or a data URL, up to 10 MB. - 300 to 6000 pixels on each side, and no wider or taller than 5:2. Data URLs outside these are refused up front with a message saying why. Links are read when the video starts: one that isn’t a usable image fails the video with `invalid_image`, free of charge. - Frames or reference images, not both in one request. Reference videos and audio aren’t supported. - Requests are JSON: images go in as links or data URLs, not as multipart uploads. ## The video object GET /v1/videos/video_abc123: ```json { "id": "video_abc123", "object": "video", "model": "seedance-2.0", "status": "completed", "progress": 100, "created_at": 1790870400, "completed_at": 1790870562, "expires_at": 1793462400, "seconds": 5, "resolution": "720p", "aspect_ratio": "16:9", "size": "1280x720", "generate_audio": true, "url": "https://...signed link...", "url_expires_at": 1790956962, "cost": 1.035, "error": null } ``` | Field | Means | | --- | --- | | `status` | `queued`, `in_progress`, `completed` or `failed`. | | `progress` | An estimate from 0 to 100, from how long the model usually takes. | | `url` | The MP4, once completed. The link works for 24 hours; fetch the job again for a fresh one. | | `expires_at` | When the video is deleted: 30 days after it was made, or sooner if someone deletes it in the Library. | | `cost` | What it cost, once completed. | | `error` | For failed videos: `code` (`content_policy`, `invalid_image`, `generation_failed`, or `generation_timeout` after 90 minutes) and a message. | ## Other endpoints | Endpoint | Returns | | --- | --- | | `GET /v1/videos/{id}` | One video, as above. | | `GET /v1/videos/{id}/content` | The MP4 itself. | | `GET /v1/videos?limit=20&after=video_abc` | Your workspace’s videos, newest first, with `first_id` and `last_id` for paging. | ## Paying for video A video costs its price per second (by resolution) times its length. When you start one, its price is set aside from your balance; it’s charged when the video is done, and released if it fails. A failed video is free. > **Tip:** Draft at 480p and a short length, then make the final cut at 720p or 1080p once the prompt is right. Most of the wait is the queue, so drafts aren’t much faster, but they cost a fraction. --- # Seedance > Seedance 2.5, 2.0, 2.0 Fast and 2.0 Mini on zurelay: which to pick, image-to-video, up to 9 reference images and prompting. Source: https://zurelay.com/docs/seedance ByteDance’s Seedance models: cinematic motion, native sound, and control from frames or reference images. ## Which one to use | Model | Pick it for | Prices per second | Length | | --- | --- | --- | --- | | Seedance 2.5 | The best quality and longest clips, up to 1080p | 480p $0.14, 720p $0.305, 1080p $0.75 | 4 to 30 s | | Seedance 2.0 | High quality at a lower price | 480p $0.097, 720p $0.207 | 4 to 15 s | | Seedance 2.0 Fast | Quicker turnarounds | 480p $0.074, 720p $0.159 | 4 to 15 s | | Seedance 2.0 Mini | The lowest price, for drafts and volume | 480p $0.047, 720p $0.10 | 4 to 15 s | All four take a prompt, a first and last frame, or up to 9 reference images, and make sound unless you set `generate_audio` to `false`. ## Prompting - Describe the motion, not just the scene: who moves, how, and what the camera does (“slow push in”, “handheld tracking shot”, “static wide shot”). - Name the light and mood: “overcast morning”, “neon at night”, “golden hour backlight”. - Keep one main action per clip. Long sequences come out better as several clips. - For speech, put the line in quotes and say who says it: `the barista says "One flat white, coming up"`. - From a first frame, describe what happens next rather than what’s already in the picture. ## Image-to-video A first frame fixes the opening shot exactly; the video takes its shape unless you set `aspect_ratio`. Add a last frame to land on a specific image, for transitions and before-and-after shots. ## Reference images Up to 9 images of characters, products, places or a look. Refer to them in the prompt as `@Image1`, `@Image2` and so on, in the order you sent them. The same references across several clips keep a character or product consistent. ## Good to know - Most of the wait (2 to 6 minutes) is queueing; a 480p clip takes about as long as a 720p one. - Prompts that may show real people or protected characters can be declined (`content_policy`); declined videos are free. - Videos are kept for 30 days; download the ones you want to keep. Full API details: [Video generation](https://zurelay.com/docs/video). Try it in the dashboard’s [Playground](https://zurelay.com/app/playground), which shows the code for each video. --- # Smart routing > When a model is down, its closest alternative answers instead of an error. Automatic or your own rules, per key or per request. Source: https://zurelay.com/docs/smart-routing When the model a request names is down, smart routing answers with a close alternative instead of an error. Your app keeps working through an outage without a line of retry code. ## How it works 1. Every request first goes to the model you asked for. Most problems end here: if one path to the model stalls or errors, the request is retried on another before you see anything. 2. Only when the model itself can’t answer (every path failed, or none is available) does smart routing step in. 3. It tries the alternatives in order, each with the same self-healing, until one answers. 4. The answer tells you which model replied (`model` in the body, `x-zurelay-model` and `x-zurelay-routed-from` headers), and it’s billed at that model’s price. When it’s on, the model you asked for gets about half of the usual time before the alternatives take over, so an outage costs seconds, not minutes. A model known to be down is skipped straight away. ## Turning it on | Where | How | Applies to | | --- | --- | --- | | Workspace default | Settings → Smart routing | Every key that doesn’t set its own, and the Playground | | Per key | API keys → Edit → Smart routing | Requests with that key | | Per request | a `models` list, or the `x-zurelay-fallback` header | That request | A key’s setting is one of: - `Workspace` (default): Follow the workspace setting (off unless you turned it on). - `Off` (mode): If the model is down, the request fails with model_unavailable, which you can retry. - `Automatic` (mode): Use our list of close alternatives for each model (below): the same family and a similar price first. - `Custom` (mode): Your own rules: for each model, up to three alternatives in order. A rule for “any other model” (`*`) covers the rest. ## Per request ### A list of alternatives Pass `models` next to `model`: up to three alternatives, tried in order if `model` is down, whatever the key’s setting or the header. It’s never sent to the model. The OpenAI Python SDK sends it through `extra_body`. **Python** (`main.py`): ```python import os from openai import OpenAI client = OpenAI( base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"], ) response = client.chat.completions.create( model="gpt-6-sol", messages=[ { "role": "user", "content": "Summarize this ticket in one line: ..." } ], extra_body={ "models": [ "gpt-6.1-sol", "gpt-6-astra" ] }, ) print(response.choices[0].message.content) ``` **Node.js** (`index.mjs`): ```javascript import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.zurelay.com/v1", apiKey: process.env.ZURELAY_API_KEY, }); const response = await client.chat.completions.create({ "model": "gpt-6-sol", "models": [ "gpt-6.1-sol", "gpt-6-astra" ], "messages": [ { "role": "user", "content": "Summarize this ticket in one line: ..." } ] }); console.log(response.choices[0].message.content); ``` **cURL**: ```bash curl https://api.zurelay.com/v1/chat/completions \ -H "Authorization: Bearer $ZURELAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-6-sol", "models": [ "gpt-6.1-sol", "gpt-6-astra" ], "messages": [ { "role": "user", "content": "Summarize this ticket in one line: ..." } ] }' ``` ### The header Headers: ``` x-zurelay-fallback: off # never reroute this request x-zurelay-fallback: auto # use the automatic alternatives for this request ``` What applies, first to last: the request’s `models`, then the header, then the key’s setting, then the workspace default. ## What alternatives can answer - **They read what you sent.** A request with a PDF or a video only goes to alternatives that read it. If none does, you get the error rather than an answer that ignores your file. - **The key may use them.** Alternatives outside a key’s allowed models are skipped. - **They speak the endpoint.** On `/v1/messages`, alternatives are Claude models. - **The request itself is fine.** If the model refused the request (a bad parameter, too long a prompt), it isn’t rerouted: the alternatives would refuse it too. ## Automatic alternatives Live from the catalog: the alternatives each model falls back to in automatic mode. | If this model is down | Automatic smart routing tries, in order | | --- | --- | | Claude Opus 5.5 (`claude-opus-5-5`) | Claude Opus 5 → Claude Opus 4.8 → Claude Sonnet 5.5 | | GPT-6 Astra (`gpt-6-astra`) | GPT-5.6 Sol → GPT-6.1 Sol → Claude Opus 5.5 | | Claude Sonnet 5.5 (`claude-sonnet-5-5`) | Claude Opus 5.5 → Claude Opus 4.8 → Claude Haiku 4.5 | | GPT-6.1 Sol (`gpt-6.1-sol`) | GPT-6 Sol → GPT-5.6 Terra → GPT-5.6 Sol | | GPT-6 Sol (`gpt-6-sol`) | GPT-6.1 Sol → GPT-5.6 Terra → GPT-5.6 Sol | | GPT-6 Luna (`gpt-6-luna`) | GPT-5.6 Luna → Gemini 3.1 Flash-Lite → DeepSeek V4.1 Flash | | Grok 4.7 (`grok-4.7`) | GPT-6.1 Sol → Gemini 3.7 Flash → GLM-5.3 | | MiMo V2.6 Pro (`mimo-v2.6-pro`) | MiMo V2.6 Flash → DeepSeek V4 Pro 0813 → GLM-5.3 | | MiMo V2.6 Flash (`mimo-v2.6-flash`) | MiMo V2.6 Pro → DeepSeek V4.1 Flash → Qwen 3.8 Flash | | DeepSeek V4.1 Flash (`deepseek-v4.1-flash`) | DeepSeek V4 Flash 0731 → DeepSeek V4 Pro 0813 → Qwen 3.8 Flash | | DeepSeek V4 Flash (`deepseek-v4-flash`) | DeepSeek V4 Flash 0731 → DeepSeek V4 Pro 0813 → Qwen 3.8 Flash | | DeepSeek V4 Flash 0731 (`deepseek-v4-flash-0731`) | DeepSeek V4.1 Flash → MiMo V2.6 Flash → Qwen 3.8 Flash | | DeepSeek V4 Pro 0813 (`deepseek-v4-pro-0813`) | DeepSeek V4.1 Flash → GLM-5.3 → MiMo V2.6 Pro | | Claude Fable 5.1 (`claude-fable-5-1`) | Claude Fable 5 → Claude Opus 5.5 → Claude Opus 5 | | Claude Fable 5 (`claude-fable-5`) | Claude Fable 5.1 → Claude Opus 5.5 → Claude Opus 5 | | Claude Opus 5 (`claude-opus-5`) | Claude Opus 4.8 → Claude Opus 5.5 → Claude Opus 4.7 | | Claude Opus 4.8 (`claude-opus-4-8`) | Claude Opus 5 → Claude Opus 4.7 → Claude Opus 5.5 | | Claude Opus 4.7 (`claude-opus-4-7`) | Claude Opus 4.8 → Claude Opus 4.6 → Claude Opus 5 | | Claude Opus 4.6 (`claude-opus-4-6`) | Claude Opus 4.7 → Claude Opus 4.8 → Claude Opus 5 | | Claude Haiku 4.5 (`claude-haiku-4-5`) | Claude Sonnet 5.5 | | GPT-5.6 Sol (`gpt-5.6-sol`) | GPT-6.1 Sol → GPT-5.6 Terra → GPT-6 Sol | | GPT-5.6 Terra (`gpt-5.6-terra`) | GPT-6.1 Sol → GPT-6 Sol → GPT-5.6 Sol | | GPT-5.6 Luna (`gpt-5.6-luna`) | GPT-6 Luna → Gemini 3.5 Flash-Lite → DeepSeek V4.1 Flash | | Gemini 3.7 Flash (`gemini-3.7-flash`) | GPT-6.1 Sol → Gemini 3.5 Flash-Lite → DeepSeek V4 Flash | | Gemini 3.5 Flash-Lite (`gemini-3.5-flash-lite`) | Gemini 3.1 Flash-Lite → Gemini 3.7 Flash → GPT-5.6 Luna | | Gemini 3.1 Flash-Lite (`gemini-3.1-flash-lite`) | Gemini 3.5 Flash-Lite → GPT-6 Luna → DeepSeek V4.1 Flash | | GLM-5.3 (`glm-5.3`) | GLM-5.2 → DeepSeek V4 Pro 0813 → MiniMax M2.7 | | GLM-5.2 (`glm-5.2`) | GLM-5.3 → DeepSeek V4 Pro 0813 → MiniMax M2.7 | | Kimi K3 (`kimi-k3`) | Qwen 3.8 Max → GLM-5.3 → DeepSeek V4 Pro 0813 | | Qwen 3.8 Max (`qwen3.8-max`) | Kimi K3 → GLM-5.3 → DeepSeek V4 Pro 0813 | | Qwen 3.8 Flash (`qwen3.8-flash`) | Qwen 3.8 Omni Flash → DeepSeek V4.1 Flash → MiMo V2.6 Flash | | Qwen 3.8 Omni Flash (`qwen3.8-omni-flash`) | Qwen 3.8 Flash → Gemini 3.1 Flash-Lite → MiMo V2.6 Flash | | HY4 Preview (`hy4-preview`) | GLM-5.3 → DeepSeek V4 Pro 0813 → Qwen 3.8 Flash | | MiniMax M2.7 (`minimax-m2.7`) | GLM-5.2 → DeepSeek V4 Pro 0813 → MiMo V2.6 Pro | ## Things to know - An alternative may cost more or less than the model you asked for. Your request log shows the model that answered and its cost. - Streaming works the same: nothing is sent until a model is answering, and the chunks name that model. - Images and videos aren’t rerouted: a different image or video model makes a different picture. > **Tip:** Check `x-zurelay-routed-from` in your logs or monitoring: it’s set only when an alternative answered, so you can see how often a model you depend on was down. --- # Key limits and rules > Monthly and daily budgets, rate limits, expiry, allowed models and allowed IP addresses for each API key. Source: https://zurelay.com/docs/key-rules Hold each API key to what it’s for: how much it spends, how fast, which models, from where, and for how long. Set them when you create a key, or any time under API keys → Edit. ## The rules | Rule | What happens past it | Error | | --- | --- | --- | | Monthly budget (USD) | Requests are refused until the next month (UTC) | `402 key_budget_exceeded` | | Daily budget (USD) | Refused until midnight UTC | `402 key_daily_budget_exceeded` | | Rate limit (requests a minute) | Refused for a moment, with `retry-after` | `429 rate_limit_exceeded` | | Expiry | The key stops working | `401 expired_api_key` | | Allowed models | Other models are refused | `403 model_not_allowed` | | Allowed IP addresses | Requests from elsewhere are refused | `403 ip_not_allowed` | Refused requests are never charged. Budgets are checked before each request, so the request that crosses one still completes. Changes apply within 15 seconds, everywhere. Your workspace balance applies to every key: when it’s empty, requests get `402 insufficient_quota`. ## Allowed models Pick any models, text, image or video. A key limited to, say, `gpt-6-sol` and `gemini-3.7-flash` can’t call anything else, and smart routing only uses alternatives from that list. ## Allowed IP addresses Single addresses and CIDR ranges, IPv4 and IPv6, up to 100 per key. The address checked is the one your request comes from on the internet (behind a proxy or NAT, that’s the proxy’s). Examples: ``` 203.0.113.7 one address 198.51.100.0/24 a range of 256 addresses 2001:db8::/32 an IPv6 range ``` > **Warning:** Serverless platforms and some clouds send requests from changing addresses. Use a range your provider publishes for egress, or leave IPs open on those keys and lean on the other rules. ## Budgets in practice - Give every app or environment its own key and budget, so one runaway loop can’t drain the balance. - A daily budget catches a bug the same day; a monthly one caps the bill. - The API keys page shows each key’s spend today and this month next to its budgets. --- # Errors > Every error zurelay returns, what it means and what to do about it. Source: https://zurelay.com/docs/errors Errors come back as JSON in the format of the endpoint you called, with a message written for people and, on the OpenAI-style endpoints, a stable code for programs. ## The shape OpenAI-style endpoints: ```json { "error": { "message": "This API key can't use `gpt-6-astra`. It's limited to gpt-6-sol. Change that at https://zurelay.com/app/keys.", "type": "permission_error", "param": "model", "code": "model_not_allowed" } } ``` On `/v1/messages` the same errors come in Anthropic’s shape, which has no code: use the HTTP status and `error.type`: `{ "type": "error", "error": { "type": "...", "message": "..." } }`. Every chat, image and video response, error or not, has an `x-request-id` header. ## Codes | Status | Code | What to do | | --- | --- | --- | | 400 | `invalid_request` | Fix the request: the message says what’s wrong and param says where. | | 400 | `unsupported_content` | The model doesn’t read that kind of file. Remove it or pick a model that does. | | 400 | `invalid_url` | A link in the request didn’t download. The message has the status it answered. | | 400 | `invalid_image` | An input image couldn’t be used (type, size or link). | | 400 | `image_not_generated` | The model declined the prompt. Not charged. | | 400 | `context_length_exceeded` | The model says the conversation is too long for it. Shorten it or pick a model with a longer context. | | 401 | `missing_api_key` | Send the key in Authorization or x-api-key. | | 401 | `invalid_api_key` | The key is wrong or revoked. | | 401 | `expired_api_key` | Create a new key. | | 402 | `insufficient_quota` | Add credit under Billing. | | 402 | `key_budget_exceeded` | Raise the key’s monthly budget, or wait for next month. | | 402 | `key_daily_budget_exceeded` | Raise the key’s daily budget, or wait for midnight UTC. | | 403 | `model_not_allowed` | The key isn’t allowed to use this model. | | 403 | `ip_not_allowed` | The key only works from other addresses. | | 404 | `model_not_found` | Check the model ID on the models page. | | 404 | `unknown_url` | No such endpoint. Check the path and the base URL. | | 404 | `video_not_found` | No video with that ID in this workspace. | | 409 | `video_not_ready` | The video isn’t finished: poll it until its status is completed. | | 409 | `video_failed` | The video failed, so there’s nothing to download. Its error says why. | | 410 | `video_expired` | The video was deleted, from the Library or after its 30 days. | | 413 | `too_large` | The request or a file is over the limit. Send big files as links. | | 429 | `rate_limit_exceeded` | Wait for retry-after seconds, then retry. | | 500 | `server_error` | Something went wrong on our side. Retry. | | 502 | `stream_interrupted` | Sent as the last event of a stream that broke off. Not charged; retry. | | 503 | `model_unavailable` | The model is down right now. Retry after retry-after seconds, or turn on smart routing. | | 503 | `storage_unavailable` | An image was made but couldn’t be delivered. Not charged; retry. | | 503 | `service_unavailable` | A brief problem on our side. Retry. | ## Retrying - Retry `429` and `503`, waiting the `retry-after` header’s seconds (or backing off exponentially). - Don’t retry other `4xx` errors unchanged: they’ll fail the same way. - Failed requests are free, so retrying costs nothing until one succeeds. - The official OpenAI and Anthropic SDKs already retry the right errors for you. ## Errors mid-stream Errors a model returns about your request itself (a parameter it doesn’t take, a prompt too long for it) come through with the model’s own message and code, as a 400 or 422. ## Errors mid-stream Once a stream has started, a failure can’t become an HTTP status. You get a final error event instead (in the endpoint’s format) and the stream ends, and the request isn’t charged. This is rare: most problems are caught and retried before the first token. --- # Reliability, retries and timeouts > How zurelay keeps requests answered, how long it waits, and when your client should retry. Source: https://zurelay.com/docs/reliability Every request is watched from start to first token. If a path to the model stalls or fails, the request moves to another before anything reaches you. ## Self-healing - Most models are served over several paths. A request starts on the healthiest one. - A path that errors, or is slower to start than it usually is, is dropped for that request and retried elsewhere. - Paths that keep failing are moved to the back of the line for everyone, and tried again once they recover. - Nothing reaches you until a real token has arrived, so all of this is invisible: you see a slightly slower answer, not an error. The models page shows each model’s live status (operational, degraded or down), from real requests and our own checks: every chat model is tested at least every 10 minutes. ## Timeouts | | Limit | | --- | --- | | Waiting for the first token | About 100 seconds, plus 1 second per 5,000 prompt tokens (up to 30), plus 2 minutes at high reasoning effort | | A quiet stream (no data) | 90 seconds; once a stream has started, keep-alive comments arrive every 15 seconds | | A whole stream | No limit while data keeps coming | | A whole non-streaming request | 10 minutes | | An image | 3 minutes per attempt, retried on another path if it stalls | Before the first token nothing is sent, not even headers, and a non-streaming request sends nothing until it’s done. Set your client’s timeout to at least 10 minutes for long reasoning, long outputs and images. The OpenAI SDKs default to 10 minutes; some HTTP clients default to 30 or 60 seconds. ## Rate limits There’s no global rate limit on your account: keys are limited only if you set a rate limit on them. A model that’s overloaded is retried on its other paths; if none answers, you get `503 model_unavailable` with `retry-after`. Back off and retry, or turn on [smart routing](https://zurelay.com/docs/smart-routing). ## When a model is down Rarely, every path to a model fails. Without smart routing, the request ends with `503 model_unavailable` and `retry-after`, within about 100 seconds rather than hanging. With smart routing, a close alternative answers instead. --- # Pricing and billing > Prepaid credit, how tokens, images, media and video seconds are counted, failed requests, and the request log. Source: https://zurelay.com/docs/billing Prepaid credit, pay as you go. Each request costs what the model’s price says, and the request log shows exactly what it counted. ## Credit - Add credit by card under [Billing](https://zurelay.com/app/billing), from $10. It’s shared by every key in the workspace. - Credit you buy doesn’t expire while your account is open. - Auto-recharge, under [Billing](https://zurelay.com/app/billing) → Payment method: save a card (it goes to Stripe, never to us) and we add a set amount whenever the balance falls below your threshold, at most once every two minutes and ten times a day, with a receipt by email. A declined card pauses it, and the owners and admins are told why. - When the balance runs out, requests are refused with `402 insufficient_quota` until you add more. The balance is checked before each request, so the last one can take it slightly below zero; the next top-up covers that. ## What a request costs | Kind | Billed by | | --- | --- | | Chat | Input tokens and output tokens, at the model’s price per million. Cached input at the cached price where the model has one. | | Images, PDFs, audio, video in a message | The input tokens the model counts for them (see Images, PDFs, audio and video) | | Reasoning | Thinking tokens count as output tokens | | Images made | Per image, by resolution | | Videos | Per second, by resolution | Prices are on the [models page](https://zurelay.com/docs/models). Each model’s own page shows the lab’s list price beside ours. ## What’s free - Failed requests, and requests refused by a key’s rules or the balance. - Images a model declined, and failed videos. - Listing models and counting tokens. - The instructions we add to keep models answering as themselves: you’re only billed for what you send and receive. A stream you cancel is billed for its input and the tokens generated up to that point. With smart routing, a request answered by an alternative is billed at the alternative’s price. ## Where to see it - The [request log](https://zurelay.com/app/requests): every request with its model, tokens, cost and speed. - Each API key’s spend today and this month on the API keys page. - Your credit history and receipts under Billing. > **Note:** Questions about a charge? Open a request under [Support](https://zurelay.com/app/support) in the dashboard with the request’s `x-request-id`. --- # Integrations > Set up Claude Code, Cursor, Cline, the OpenAI and Anthropic SDKs, LangChain, n8n and more with zurelay. Source: https://zurelay.com/docs/integrations zurelay speaks the OpenAI and Anthropic APIs, so most tools work by changing a base URL and a key. Step-by-step guides for the popular ones: ## Coding agents - [Claude Code](https://zurelay.com/docs/integrations/claude-code): Anthropic’s coding agent for the terminal and IDE. - [OpenCode](https://zurelay.com/docs/integrations/opencode): Open-source coding agent for the terminal. - [Cursor](https://zurelay.com/docs/integrations/cursor): The AI code editor. Chat and agent on your key. - [Cline](https://zurelay.com/docs/integrations/cline): Autonomous coding agent for VS Code and JetBrains. - [Kilo Code](https://zurelay.com/docs/integrations/kilo-code): Open-source coding agent for VS Code and CLI. - [Continue](https://zurelay.com/docs/integrations/continue): Open-source assistant for VS Code and JetBrains. - [Aider](https://zurelay.com/docs/integrations/aider): AI pair programming in your terminal. - [Zed](https://zurelay.com/docs/integrations/zed): The fast editor with a built-in agent panel. ## SDKs & frameworks - [OpenAI Node SDK](https://zurelay.com/docs/integrations/openai-node): Official OpenAI library for TypeScript and JS. - [OpenAI Python SDK](https://zurelay.com/docs/integrations/openai-python): Official OpenAI library for Python. - [Anthropic SDK](https://zurelay.com/docs/integrations/anthropic-sdk): Official Claude libraries for Python and TS. - [Vercel AI SDK](https://zurelay.com/docs/integrations/vercel-ai-sdk): Streaming, tools and UI hooks for TypeScript. - [LangChain](https://zurelay.com/docs/integrations/langchain): Agents and LangGraph, in Python or JavaScript. - [LlamaIndex](https://zurelay.com/docs/integrations/llamaindex): RAG and document agents over your own data. - [cURL](https://zurelay.com/docs/integrations/curl): Plain HTTP, for any language or a quick test. ## Chat apps & assistants - [Hermes Agent](https://zurelay.com/docs/integrations/hermes-agent): Nous Research’s agent for terminal and chat apps. - [OpenClaw](https://zurelay.com/docs/integrations/openclaw): Open-source personal assistant on your devices. - [Open WebUI](https://zurelay.com/docs/integrations/open-webui): Self-hosted, ChatGPT-style chat for your team. - [LibreChat](https://zurelay.com/docs/integrations/librechat): Open-source, multi-user chat with agents. - [LobeHub](https://zurelay.com/docs/integrations/lobehub): Open-source AI workspace, formerly LobeChat. - [Cherry Studio](https://zurelay.com/docs/integrations/cherry-studio): A desktop AI client for Windows, macOS and Linux. - [SillyTavern](https://zurelay.com/docs/integrations/sillytavern): Local-first frontend for character chat. ## Automation - [n8n](https://zurelay.com/docs/integrations/n8n): Workflow automation with AI agent nodes. - [Dify](https://zurelay.com/docs/integrations/dify): AI workflows, agents and RAG on a canvas. ## Anything else If a tool lets you set an OpenAI base URL (sometimes called “OpenAI-compatible” or “custom provider”), use `https://api.zurelay.com/v1` and your zurelay key. If it takes an Anthropic base URL, use `https://api.zurelay.com`. Then pick a model ID from the [models page](https://zurelay.com/docs/models). --- # Claude Code with zurelay > Anthropic’s coding agent for the terminal and IDE. Made by Anthropic. Source: https://zurelay.com/docs/integrations/claude-code The examples use `claude-opus-5-5`. Any Claude model ID from [the models page](https://zurelay.com/docs/models) works. ## Set it up 1. Set `ANTHROPIC_BASE_URL` to `https://api.zurelay.com`, with no `/v1`. Claude Code adds `/v1/messages` itself. 2. Set `ANTHROPIC_AUTH_TOKEN` to your zurelay key. It’s sent as a Bearer token, with no approval prompt. 3. Point the Opus, Sonnet and Haiku aliases at zurelay models, then run `claude`. 4. To keep the setup, put the same values in the `env` block of `~/.claude/settings.json`. ## Configuration **Shell** (`terminal`): ```bash export ANTHROPIC_BASE_URL="https://api.zurelay.com" export ANTHROPIC_AUTH_TOKEN="$ZURELAY_API_KEY" export ANTHROPIC_DEFAULT_OPUS_MODEL="claude-opus-5-5" export ANTHROPIC_DEFAULT_SONNET_MODEL="claude-sonnet-5-5" export ANTHROPIC_DEFAULT_HAIKU_MODEL="claude-sonnet-5-5" claude --model claude-opus-5-5 ``` **settings.json** (`~/.claude/settings.json`): ```json { "env": { "ANTHROPIC_BASE_URL": "https://api.zurelay.com", "ANTHROPIC_AUTH_TOKEN": "", "ANTHROPIC_DEFAULT_OPUS_MODEL": "claude-opus-5-5", "ANTHROPIC_DEFAULT_SONNET_MODEL": "claude-sonnet-5-5", "ANTHROPIC_DEFAULT_HAIKU_MODEL": "claude-sonnet-5-5" }, "model": "claude-opus-5-5" } ``` > **Note:** Claude Code speaks the Anthropic API, so it runs on Claude models. Map the Haiku alias too: Claude Code uses it for background work. In the VS Code extension, add the same variables under `claudeCode.environmentVariables`. [Claude Code’s own docs](https://code.claude.com/docs/en/llm-gateway-connect) --- # OpenCode with zurelay > Open-source coding agent for the terminal. Made by SST. Source: https://zurelay.com/docs/integrations/opencode The examples use `claude-opus-5-5`. Any model ID from [the models page](https://zurelay.com/docs/models) works. ## Set it up 1. Export `ZURELAY_API_KEY` in your shell profile. 2. Add zurelay as a provider in `~/.config/opencode/opencode.json`, or in `opencode.json` at your project root. 3. Run `opencode` and switch models any time with `/models`. ## Configuration opencode.json (`~/.config/opencode/opencode.json`): ```json { "$schema": "https://opencode.ai/config.json", "model": "zurelay/claude-opus-5-5", "small_model": "zurelay/deepseek-v4.1-flash", "provider": { "zurelay": { "npm": "@ai-sdk/openai-compatible", "name": "zurelay", "options": { "baseURL": "https://api.zurelay.com/v1", "apiKey": "{env:ZURELAY_API_KEY}" }, "models": { "claude-opus-5-5": { "name": "Claude Opus 5.5" }, "gpt-6-astra": { "name": "GPT-6 Astra" }, "claude-sonnet-5-5": { "name": "Claude Sonnet 5.5" }, "gpt-6.1-sol": { "name": "GPT-6.1 Sol" }, "gpt-6-sol": { "name": "GPT-6 Sol" }, "gpt-6-luna": { "name": "GPT-6 Luna" }, "grok-4.7": { "name": "Grok 4.7" }, "mimo-v2.6-pro": { "name": "MiMo V2.6 Pro" }, "mimo-v2.6-flash": { "name": "MiMo V2.6 Flash" }, "deepseek-v4.1-flash": { "name": "DeepSeek V4.1 Flash" }, "deepseek-v4-flash": { "name": "DeepSeek V4 Flash" }, "deepseek-v4-flash-0731": { "name": "DeepSeek V4 Flash 0731" }, "deepseek-v4-pro-0813": { "name": "DeepSeek V4 Pro 0813" }, "claude-fable-5-1": { "name": "Claude Fable 5.1" }, "claude-fable-5": { "name": "Claude Fable 5" }, "claude-opus-5": { "name": "Claude Opus 5" }, "claude-opus-4-8": { "name": "Claude Opus 4.8" }, "claude-opus-4-7": { "name": "Claude Opus 4.7" }, "claude-opus-4-6": { "name": "Claude Opus 4.6" }, "claude-haiku-4-5": { "name": "Claude Haiku 4.5" }, "gpt-5.6-sol": { "name": "GPT-5.6 Sol" }, "gpt-5.6-terra": { "name": "GPT-5.6 Terra" }, "gpt-5.6-luna": { "name": "GPT-5.6 Luna" }, "gemini-3.7-flash": { "name": "Gemini 3.7 Flash" }, "gemini-3.5-flash-lite": { "name": "Gemini 3.5 Flash-Lite" }, "gemini-3.1-flash-lite": { "name": "Gemini 3.1 Flash-Lite" }, "glm-5.3": { "name": "GLM-5.3" }, "glm-5.2": { "name": "GLM-5.2" }, "kimi-k3": { "name": "Kimi K3" }, "qwen3.8-max": { "name": "Qwen 3.8 Max" }, "qwen3.8-flash": { "name": "Qwen 3.8 Flash" }, "qwen3.8-omni-flash": { "name": "Qwen 3.8 Omni Flash" }, "hy4-preview": { "name": "HY4 Preview" }, "minimax-m2.7": { "name": "MiniMax M2.7" } } } } } ``` > **Note:** Keep the `@ai-sdk/openai-compatible` package: it calls Chat Completions, which every zurelay model supports. `small_model` handles titles and other quick jobs. [OpenCode’s own docs](https://opencode.ai/docs/providers) --- # Cursor with zurelay > The AI code editor. Chat and agent on your key. Made by Anysphere. Source: https://zurelay.com/docs/integrations/cursor The examples use `claude-opus-5-5`. Any model ID from [the models page](https://zurelay.com/docs/models) works. ## Set it up 1. Open Cursor Settings and go to Models. 2. Under API Keys, paste your zurelay key into OpenAI API Key. Turn on Override OpenAI Base URL and enter the base URL. 3. Add the model ID as a custom model, make sure it’s switched on, and pick it in the chat’s model menu. - **OpenAI API Key**: Your zurelay API key - **Base URL**: `https://api.zurelay.com/v1` - **Custom model**: `claude-opus-5-5` > **Note:** Cursor sends these requests from its own servers. Chat and agent use your key; Tab completion keeps using Cursor’s built-in models. [Cursor’s own docs](https://cursor.com/help/models-and-usage/api-keys) --- # Cline with zurelay > Autonomous coding agent for VS Code and JetBrains. Source: https://zurelay.com/docs/integrations/cline The examples use `claude-opus-5-5`. Any model ID from [the models page](https://zurelay.com/docs/models) works. ## Set it up 1. Open Cline and click the settings gear. 2. Set API Provider to OpenAI Compatible and fill in the fields below. 3. Optionally set the context window and image support in the model settings, then start a task. - **API Provider**: OpenAI Compatible - **Base URL**: `https://api.zurelay.com/v1` - **API Key**: Your zurelay API key - **Model ID**: `claude-opus-5-5` [Cline’s own docs](https://docs.cline.bot/provider-config/openai-compatible) --- # Kilo Code with zurelay > Open-source coding agent for VS Code and CLI. Source: https://zurelay.com/docs/integrations/kilo-code The examples use `claude-opus-5-5`. Any model ID from [the models page](https://zurelay.com/docs/models) works. ## Set it up 1. Open Kilo Code’s settings and go to the Providers tab. 2. Add a custom provider with the ID `zurelay`, and set Provider API to OpenAI Compatible. 3. Enter the base URL and your key. Kilo loads the model list from zurelay, so you can pick any model. - **Provider ID**: `zurelay` - **Provider API**: OpenAI Compatible - **Base URL**: `https://api.zurelay.com/v1` - **API key**: Your zurelay API key - **Model**: `claude-opus-5-5` > **Note:** Choose OpenAI Compatible even for GPT and Grok models. Kilo’s OpenAI Responses option calls an API zurelay doesn’t serve. [Kilo Code’s own docs](https://kilo.ai/docs/ai-providers/openai-compatible) --- # Continue with zurelay > Open-source assistant for VS Code and JetBrains. Source: https://zurelay.com/docs/integrations/continue The examples use `claude-opus-5-5`. Any model ID from [the models page](https://zurelay.com/docs/models) works. ## Set it up 1. Add `ZURELAY_API_KEY=…` to `~/.continue/.env`. The IDE extensions read secrets from there, not from your shell. 2. Add a zurelay model to `~/.continue/config.yaml`. 3. Reload Continue and pick the model in the chat panel. ## Configuration config.yaml (`~/.continue/config.yaml`): ```yaml name: zurelay version: 1.0.0 schema: v1 models: - name: "Claude Opus 5.5" provider: openai model: "claude-opus-5-5" apiBase: https://api.zurelay.com/v1 apiKey: ${{ secrets.ZURELAY_API_KEY }} useResponsesApi: false roles: - chat - edit - apply ``` > **Note:** `useResponsesApi: false` keeps GPT models on Chat Completions, which zurelay serves. [Continue’s own docs](https://docs.continue.dev/customize/model-providers/top-level/openai) --- # Aider with zurelay > AI pair programming in your terminal. Source: https://zurelay.com/docs/integrations/aider The examples use `claude-opus-5-5`. Any model ID from [the models page](https://zurelay.com/docs/models) works. ## Set it up 1. Point Aider’s OpenAI settings at zurelay with `OPENAI_API_BASE` and `OPENAI_API_KEY`. 2. Prefix model IDs with `openai/` so Aider sends them to that endpoint. 3. To keep the setup, put the same settings in `.aider.conf.yml` in your home directory or repo. ## Configuration **Shell** (`terminal`): ```bash export OPENAI_API_BASE="https://api.zurelay.com/v1" export OPENAI_API_KEY="$ZURELAY_API_KEY" aider --model openai/claude-opus-5-5 --weak-model openai/deepseek-v4.1-flash ``` **.aider.conf.yml** (`~/.aider.conf.yml`): ```yaml model: "openai/claude-opus-5-5" weak-model: "openai/deepseek-v4.1-flash" openai-api-base: https://api.zurelay.com/v1 openai-api-key: ``` > **Note:** Aider may warn that it doesn’t know a model’s context size. It still works; add the model to `.aider.model.metadata.json` to silence the warning. [Aider’s own docs](https://aider.chat/docs/llms/openai-compat.html) --- # Zed with zurelay > The fast editor with a built-in agent panel. Made by Zed Industries. Source: https://zurelay.com/docs/integrations/zed The examples use `claude-opus-5-5`. Any model ID from [the models page](https://zurelay.com/docs/models) works. ## Set it up 1. Export `ZURELAY_API_KEY`. Zed reads it for a provider named `zurelay`, or you can paste the key in Agent Settings. 2. Add zurelay under `language_models.openai_compatible` in your settings, or use Add Provider under LLM Providers in Agent Settings. 3. Pick the model from the Agent Panel’s model menu. ## Configuration settings.json (`~/.config/zed/settings.json`): ```json { "language_models": { "openai_compatible": { "zurelay": { "api_url": "https://api.zurelay.com/v1", "available_models": [ { "name": "claude-opus-5-5", "display_name": "Claude Opus 5.5", "max_tokens": 1000000 } ] } } } } ``` > **Note:** Keep the key out of `settings.json`. Zed reads `ZURELAY_API_KEY` from your environment, or stores a key you paste in the system keychain. [Zed’s own docs](https://zed.dev/docs/ai/use-api-access) --- # OpenAI Node SDK with zurelay > Official OpenAI library for TypeScript and JS. Made by OpenAI. Source: https://zurelay.com/docs/integrations/openai-node The examples use `claude-opus-5-5`. Any model ID from [the models page](https://zurelay.com/docs/models) works. ## Set it up 1. Install the SDK: `npm install openai`. 2. Create the client with zurelay’s base URL and your key. 3. Call any chat model in the catalog by its ID. Streaming, tools and JSON mode work as usual. ## Configuration **Chat** (`app.ts`): ```typescript import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.zurelay.com/v1", apiKey: process.env.ZURELAY_API_KEY, }); const completion = await client.chat.completions.create({ model: "claude-opus-5-5", messages: [{ role: "user", content: "Hello" }], }); console.log(completion.choices[0]?.message.content); ``` **Images** (`image.ts`): ```typescript import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.zurelay.com/v1", apiKey: process.env.ZURELAY_API_KEY, }); const result = await client.images.generate({ model: "nano-banana-pro", prompt: "A glass prism splitting white light into a rainbow", n: 1, }); console.log(result.data?.[0]?.url); ``` > **Note:** The SDK also reads `OPENAI_BASE_URL` and `OPENAI_API_KEY`, so existing code can switch with two environment variables. [OpenAI Node SDK’s own docs](https://github.com/openai/openai-node) --- # OpenAI Python SDK with zurelay > Official OpenAI library for Python. Made by OpenAI. Source: https://zurelay.com/docs/integrations/openai-python The examples use `claude-opus-5-5`. Any model ID from [the models page](https://zurelay.com/docs/models) works. ## Set it up 1. Install the SDK: `pip install openai`. 2. Create the client with zurelay’s base URL and your key. 3. Call any chat model in the catalog by its ID. `AsyncOpenAI` takes the same arguments. ## Configuration **Chat** (`app.py`): ```python import os from openai import OpenAI client = OpenAI( base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"], ) completion = client.chat.completions.create( model="claude-opus-5-5", messages=[{"role": "user", "content": "Hello"}], ) print(completion.choices[0].message.content) ``` **Images** (`image.py`): ```python import os from openai import OpenAI client = OpenAI( base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"], ) result = client.images.generate( model="nano-banana-pro", prompt="A glass prism splitting white light into a rainbow", n=1, ) print(result.data[0].url) ``` > **Note:** The SDK also reads `OPENAI_BASE_URL` and `OPENAI_API_KEY`, so existing code can switch with two environment variables. [OpenAI Python SDK’s own docs](https://github.com/openai/openai-python) --- # Anthropic SDK with zurelay > Official Claude libraries for Python and TS. Made by Anthropic. Source: https://zurelay.com/docs/integrations/anthropic-sdk The examples use `claude-opus-5-5`. Any Claude model ID from [the models page](https://zurelay.com/docs/models) works. ## Set it up 1. Install the SDK: `pip install anthropic` or `npm install @anthropic-ai/sdk`. 2. Set the base URL to `https://api.zurelay.com`, with no `/v1`. The SDK adds `/v1/messages` itself. 3. Use a Claude model ID. Other models answer on the OpenAI-compatible URL. ## Configuration **Python** (`app.py`): ```python import os from anthropic import Anthropic client = Anthropic( base_url="https://api.zurelay.com", api_key=os.environ["ZURELAY_API_KEY"], ) message = client.messages.create( model="claude-opus-5-5", max_tokens=1024, messages=[{"role": "user", "content": "Hello"}], ) print(message.content[0].text) ``` **TypeScript** (`app.ts`): ```typescript import Anthropic from "@anthropic-ai/sdk"; const client = new Anthropic({ baseURL: "https://api.zurelay.com", apiKey: process.env.ZURELAY_API_KEY, }); const message = await client.messages.create({ model: "claude-opus-5-5", max_tokens: 1024, messages: [{ role: "user", content: "Hello" }], }); console.log(message.content); ``` > **Note:** Both SDKs also read `ANTHROPIC_BASE_URL` and `ANTHROPIC_API_KEY` from the environment. [Anthropic SDK’s own docs](https://platform.claude.com/docs/en/cli-sdks-libraries/overview) --- # Vercel AI SDK with zurelay > Streaming, tools and UI hooks for TypeScript. Made by Vercel. Source: https://zurelay.com/docs/integrations/vercel-ai-sdk The examples use `claude-opus-5-5`. Any model ID from [the models page](https://zurelay.com/docs/models) works. ## Set it up 1. Install `ai` and the OpenAI-compatible provider: `npm install ai @ai-sdk/openai-compatible`. 2. Create a zurelay provider with the base URL and your key. 3. Pass `zurelay(modelId)` to `streamText`, `generateText` or any other AI SDK call. ## Configuration TypeScript (`lib/ai.ts`): ```typescript import { createOpenAICompatible } from "@ai-sdk/openai-compatible"; import { streamText } from "ai"; const zurelay = createOpenAICompatible({ name: "zurelay", baseURL: "https://api.zurelay.com/v1", apiKey: process.env.ZURELAY_API_KEY, }); const result = streamText({ model: zurelay("claude-opus-5-5"), prompt: "Write a haiku about routers.", }); for await (const text of result.textStream) process.stdout.write(text); ``` [Vercel AI SDK’s own docs](https://ai-sdk.dev/providers/openai-compatible-providers) --- # LangChain with zurelay > Agents and LangGraph, in Python or JavaScript. Source: https://zurelay.com/docs/integrations/langchain The examples use `claude-opus-5-5`. Any model ID from [the models page](https://zurelay.com/docs/models) works. ## Set it up 1. Install the OpenAI integration: `pip install langchain-openai` or `npm install @langchain/openai`. 2. Create `ChatOpenAI` with zurelay’s base URL, your key and a catalog model. 3. Use it anywhere LangChain or LangGraph takes a chat model. ## Configuration **Python** (`agent.py`): ```python import os from langchain_openai import ChatOpenAI llm = ChatOpenAI( model="claude-opus-5-5", base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"], ) print(llm.invoke("Hello").content) ``` **TypeScript** (`agent.ts`): ```typescript import { ChatOpenAI } from "@langchain/openai"; const llm = new ChatOpenAI({ model: "claude-opus-5-5", apiKey: process.env.ZURELAY_API_KEY, configuration: { baseURL: "https://api.zurelay.com/v1" }, }); const reply = await llm.invoke("Hello"); console.log(reply.content); ``` [LangChain’s own docs](https://docs.langchain.com/oss/python/integrations/chat/openai) --- # LlamaIndex with zurelay > RAG and document agents over your own data. Source: https://zurelay.com/docs/integrations/llamaindex The examples use `claude-opus-5-5`. Any model ID from [the models page](https://zurelay.com/docs/models) works. ## Set it up 1. Install the OpenAI-compatible LLM: `pip install llama-index-llms-openai-like`. 2. Create `OpenAILike` with zurelay’s base URL, your key and a catalog model. 3. Set it as `Settings.llm` so every index and query engine uses it. ## Configuration Python (`app.py`): ```python import os from llama_index.core import Settings from llama_index.llms.openai_like import OpenAILike Settings.llm = OpenAILike( model="claude-opus-5-5", api_base="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"], is_chat_model=True, is_function_calling_model=True, context_window=1000000, ) print(Settings.llm.complete("Hello")) ``` [LlamaIndex’s own docs](https://developers.llamaindex.ai/python/framework-api-reference/llms/openai_like/) --- # cURL with zurelay > Plain HTTP, for any language or a quick test. Made by curl. Source: https://zurelay.com/docs/integrations/curl The examples use `claude-opus-5-5`. Any model ID from [the models page](https://zurelay.com/docs/models) works. ## Set it up 1. Send your key as a Bearer token in the `Authorization` header. 2. Use Chat Completions for any model, or the Messages API for Claude models. 3. Image models take the same key at `/v1/images/generations`. ## Configuration **Chat Completions** (`terminal`): ```bash curl https://api.zurelay.com/v1/chat/completions \ -H "Authorization: Bearer $ZURELAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "claude-opus-5-5", "messages": [{ "role": "user", "content": "Hello" }] }' ``` **Messages** (`terminal`): ```bash curl https://api.zurelay.com/v1/messages \ -H "Authorization: Bearer $ZURELAY_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "Content-Type: application/json" \ -d '{ "model": "claude-opus-5-5", "max_tokens": 1024, "messages": [{ "role": "user", "content": "Hello" }] }' ``` **Images** (`terminal`): ```bash curl https://api.zurelay.com/v1/images/generations \ -H "Authorization: Bearer $ZURELAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "nano-banana-pro", "prompt": "A glass prism splitting white light into a rainbow", "n": 1 }' ``` [cURL’s own docs](https://curl.se/docs/manpage.html) --- # Hermes Agent with zurelay > Nous Research’s agent for terminal and chat apps. Made by Nous Research. Source: https://zurelay.com/docs/integrations/hermes-agent The examples use `claude-opus-5-5`. Any model ID from [the models page](https://zurelay.com/docs/models) works. ## Set it up 1. Run `hermes model` and choose Custom endpoint. 2. Enter the base URL, your key and a model ID. If it asks for an API mode, choose `chat_completions`. 3. Or set it up by hand: the endpoint in `~/.hermes/config.yaml`, the key in `~/.hermes/.env`. 4. Start `hermes`. Switch models mid-chat with `/model custom:`. - **Base URL**: `https://api.zurelay.com/v1` - **API key**: Your zurelay API key - **Model**: `claude-opus-5-5` ## Configuration **config.yaml** (`~/.hermes/config.yaml`): ```yaml model: default: "claude-opus-5-5" provider: custom base_url: https://api.zurelay.com/v1 key_env: ZURELAY_API_KEY ``` **.env** (`~/.hermes/.env`): ``` ZURELAY_API_KEY= ``` [Hermes Agent’s own docs](https://hermes-agent.nousresearch.com/docs/integrations/providers) --- # OpenClaw with zurelay > Open-source personal assistant on your devices. Source: https://zurelay.com/docs/integrations/openclaw The examples use `claude-opus-5-5`. Any model ID from [the models page](https://zurelay.com/docs/models) works. ## Set it up 1. Put `ZURELAY_API_KEY=…` in `~/.openclaw/.env`. 2. Add zurelay as a custom provider in `~/.openclaw/openclaw.json` and make it the default model. 3. Restart with `openclaw gateway restart`. Switch later with `openclaw models set zurelay/`. ## Configuration **openclaw.json** (`~/.openclaw/openclaw.json`): ```json { agents: { defaults: { model: { primary: "zurelay/claude-opus-5-5" } }, }, models: { mode: "merge", providers: { zurelay: { baseUrl: "https://api.zurelay.com/v1", apiKey: "${ZURELAY_API_KEY}", api: "openai-completions", models: [ { id: "claude-opus-5-5", name: "Claude Opus 5.5" }, { id: "gpt-6-astra", name: "GPT-6 Astra" }, { id: "claude-sonnet-5-5", name: "Claude Sonnet 5.5" }, { id: "gpt-6.1-sol", name: "GPT-6.1 Sol" }, { id: "gpt-6-sol", name: "GPT-6 Sol" }, { id: "gpt-6-luna", name: "GPT-6 Luna" }, { id: "grok-4.7", name: "Grok 4.7" }, { id: "mimo-v2.6-pro", name: "MiMo V2.6 Pro" }, { id: "mimo-v2.6-flash", name: "MiMo V2.6 Flash" }, { id: "deepseek-v4.1-flash", name: "DeepSeek V4.1 Flash" }, { id: "deepseek-v4-flash", name: "DeepSeek V4 Flash" }, { id: "deepseek-v4-flash-0731", name: "DeepSeek V4 Flash 0731" }, { id: "deepseek-v4-pro-0813", name: "DeepSeek V4 Pro 0813" }, { id: "claude-fable-5-1", name: "Claude Fable 5.1" }, { id: "claude-fable-5", name: "Claude Fable 5" }, { id: "claude-opus-5", name: "Claude Opus 5" }, { id: "claude-opus-4-8", name: "Claude Opus 4.8" }, { id: "claude-opus-4-7", name: "Claude Opus 4.7" }, { id: "claude-opus-4-6", name: "Claude Opus 4.6" }, { id: "claude-haiku-4-5", name: "Claude Haiku 4.5" }, { id: "gpt-5.6-sol", name: "GPT-5.6 Sol" }, { id: "gpt-5.6-terra", name: "GPT-5.6 Terra" }, { id: "gpt-5.6-luna", name: "GPT-5.6 Luna" }, { id: "gemini-3.7-flash", name: "Gemini 3.7 Flash" }, { id: "gemini-3.5-flash-lite", name: "Gemini 3.5 Flash-Lite" }, { id: "gemini-3.1-flash-lite", name: "Gemini 3.1 Flash-Lite" }, { id: "glm-5.3", name: "GLM-5.3" }, { id: "glm-5.2", name: "GLM-5.2" }, { id: "kimi-k3", name: "Kimi K3" }, { id: "qwen3.8-max", name: "Qwen 3.8 Max" }, { id: "qwen3.8-flash", name: "Qwen 3.8 Flash" }, { id: "qwen3.8-omni-flash", name: "Qwen 3.8 Omni Flash" }, { id: "hy4-preview", name: "HY4 Preview" }, { id: "minimax-m2.7", name: "MiniMax M2.7" }, ], }, }, }, } ``` **Onboarding** (`terminal`): ```bash openclaw onboard --auth-choice custom-api-key \ --custom-base-url "https://api.zurelay.com/v1" \ --custom-model-id "claude-opus-5-5" \ --custom-api-key "$ZURELAY_API_KEY" \ --custom-compatibility openai ``` > **Note:** Use the `openai-completions` API, not `openai-responses`: zurelay serves Chat Completions. [OpenClaw’s own docs](https://docs.openclaw.ai/concepts/model-providers/custom-providers) --- # Open WebUI with zurelay > Self-hosted, ChatGPT-style chat for your team. Source: https://zurelay.com/docs/integrations/open-webui ## Set it up 1. Open Admin Panel, then Settings, then Connections. 2. Under OpenAI API, add a connection with the URL and key below. Leave API Type on Chat Completions. 3. Save. Every zurelay model shows up in the model picker; list Model IDs to show only some. - **URL**: `https://api.zurelay.com/v1` - **API Key**: Your zurelay API key - **API Type**: Chat Completions ## Configuration Docker (`terminal`): ```bash docker run -d -p 3000:8080 \ -e OPENAI_API_BASE_URL="https://api.zurelay.com/v1" \ -e OPENAI_API_KEY="$ZURELAY_API_KEY" \ -v open-webui:/app/backend/data \ --name open-webui ghcr.io/open-webui/open-webui:main ``` > **Note:** Environment variables only seed the first start. After that, Open WebUI keeps connections in its database, so change them in the Admin Panel. [Open WebUI’s own docs](https://docs.openwebui.com/getting-started/quick-start/connect-a-provider/starting-with-openai-compatible/) --- # LibreChat with zurelay > Open-source, multi-user chat with agents. Source: https://zurelay.com/docs/integrations/librechat The examples use `claude-opus-5-5`. Any model ID from [the models page](https://zurelay.com/docs/models) works. ## Set it up 1. Add `ZURELAY_API_KEY=…` to LibreChat’s `.env`. 2. Add zurelay as a custom endpoint in `librechat.yaml`. With `fetch: true`, the full model list loads from zurelay. 3. With Docker, mount the file in `docker-compose.override.yml`, then run `docker compose down && docker compose up -d`. ## Configuration **librechat.yaml** (`librechat.yaml`): ```yaml version: 1.3.17 endpoints: custom: - name: "zurelay" apiKey: "${ZURELAY_API_KEY}" baseURL: "https://api.zurelay.com/v1" models: default: ["claude-opus-5-5"] fetch: true titleConvo: true titleModel: "deepseek-v4.1-flash" modelDisplayLabel: "zurelay" ``` **docker-compose.override.yml** (`docker-compose.override.yml`): ```yaml services: api: volumes: - type: bind source: ./librechat.yaml target: /app/librechat.yaml ``` [LibreChat’s own docs](https://www.librechat.ai/docs/configuration/librechat_yaml/object_structure/custom_endpoint) --- # LobeHub with zurelay > Open-source AI workspace, formerly LobeChat. Source: https://zurelay.com/docs/integrations/lobehub The examples use `claude-opus-5-5`. Any model ID from [the models page](https://zurelay.com/docs/models) works. ## Set it up 1. Open Settings, then Provider (AI Service Provider in older versions), and choose Add Custom Provider. 2. Set Request Format to OpenAI, and enter the proxy URL and your key below. 3. Fetch the model list, or add model IDs by hand. Leave Use Responses API Specification off. - **Provider ID**: `zurelay` - **Request Format**: OpenAI - **Proxy URL**: `https://api.zurelay.com/v1` - **API Key**: Your zurelay API key - **Model ID**: `claude-opus-5-5` ## Configuration Self-hosted .env (`.env`): ``` OPENAI_API_KEY= OPENAI_PROXY_URL=https://api.zurelay.com/v1 OPENAI_MODEL_LIST=-all,+claude-opus-5-5,+gpt-6-astra,+claude-sonnet-5-5,+gpt-6.1-sol,+gpt-6-sol,+gpt-6-luna,+grok-4.7,+mimo-v2.6-pro,+mimo-v2.6-flash,+deepseek-v4.1-flash,+deepseek-v4-flash,+deepseek-v4-flash-0731,+deepseek-v4-pro-0813,+claude-fable-5-1,+claude-fable-5,+claude-opus-5,+claude-opus-4-8,+claude-opus-4-7,+claude-opus-4-6,+claude-haiku-4-5,+gpt-5.6-sol,+gpt-5.6-terra,+gpt-5.6-luna,+gemini-3.7-flash,+gemini-3.5-flash-lite,+gemini-3.1-flash-lite,+glm-5.3,+glm-5.2,+kimi-k3,+qwen3.8-max,+qwen3.8-flash,+qwen3.8-omni-flash,+hy4-preview,+minimax-m2.7 ``` [LobeHub’s own docs](https://lobehub.com/docs/usage/providers) --- # Cherry Studio with zurelay > A desktop AI client for Windows, macOS and Linux. Made by CherryHQ. Source: https://zurelay.com/docs/integrations/cherry-studio The examples use `claude-opus-5-5`. Any model ID from [the models page](https://zurelay.com/docs/models) works. ## Set it up 1. Open Settings, then Model Provider, and add a custom provider named zurelay. 2. Paste your key and use `https://api.zurelay.com` as the API host. Cherry Studio adds `/v1` itself. 3. Load the model list, add the models you want, and run the connection check. - **Provider type**: OpenAI - **API Key**: Your zurelay API key - **API Host**: `https://api.zurelay.com` - **Model ID**: `claude-opus-5-5` [Cherry Studio’s own docs](https://cherryai.com/docs/en/pre-basic/providers/zi-ding-yi-fu-wu-shang/) --- # SillyTavern with zurelay > Local-first frontend for character chat. Source: https://zurelay.com/docs/integrations/sillytavern The examples use `claude-opus-5-5`. Any model ID from [the models page](https://zurelay.com/docs/models) works. ## Set it up 1. Open API Connections and set API to Chat Completion. 2. Choose Custom (OpenAI-compatible) as the source and fill in the fields below. 3. Click Connect, then pick a model from Available Models. - **Custom Endpoint**: `https://api.zurelay.com/v1` - **Custom API Key**: Your zurelay API key - **Model ID**: `claude-opus-5-5` [SillyTavern’s own docs](https://docs.sillytavern.app/usage/api-connections/openai/) --- # n8n with zurelay > Workflow automation with AI agent nodes. Source: https://zurelay.com/docs/integrations/n8n The examples use `claude-opus-5-5`. Any model ID from [the models page](https://zurelay.com/docs/models) works. ## Set it up 1. Create an OpenAI credential with your zurelay key as API Key and the base URL below as Base URL. 2. In an AI Agent, add an OpenAI Chat Model and select that credential. 3. Pick the model From List, or By ID. Turn off Use Responses API so the node calls Chat Completions. - **API Key**: Your zurelay API key - **Base URL**: `https://api.zurelay.com/v1` - **Model**: `claude-opus-5-5` > **Note:** For Claude over the Messages API, the Anthropic credential has a Base URL too. Set it to `https://api.zurelay.com` and use the Anthropic Chat Model node. [n8n’s own docs](https://docs.n8n.io/integrations/builtin/credentials/openai) --- # Dify with zurelay > AI workflows, agents and RAG on a canvas. Made by LangGenius. Source: https://zurelay.com/docs/integrations/dify The examples use `claude-opus-5-5`. Any model ID from [the models page](https://zurelay.com/docs/models) works. ## Set it up 1. Go to Integrations, then Model Provider (under Settings in older versions), and install OpenAI-API-compatible from the Marketplace. 2. Click Add Model on its card, set Model Type to LLM, and fill in the fields below. 3. Raise Model context size from its 4,096 default, and set Function Call Type to Tool Call for agents. - **Model Name**: `claude-opus-5-5` - **API Key**: Your zurelay API key - **API Base URL**: `https://api.zurelay.com/v1` - **Completion mode**: Chat [Dify’s own docs](https://docs.dify.ai/en/use-dify/workspace/model-providers)