# Chat completions > POST /v1/chat/completions: messages, streaming, tool calling, structured output, reasoning effort and usage. Source: https://zurelay.com/docs/chat OpenAI’s Chat Completions API, for every text model: the same request and response, the same streaming, the same SDKs. `POST /v1/chat/completions` ## Request - `model` (string, required): A model ID from the [models page](https://zurelay.com/docs/models), such as `claude-opus-5-5` or `gpt-6-sol`. - `messages` (array, required): The conversation: `system` (or `developer`), `user`, `assistant` and `tool` messages. `content` is a string, or an array of parts to add [images, PDFs, audio or video](https://zurelay.com/docs/multimodal). - `stream` (boolean): Send the answer as server-sent events while it’s written. See Streaming below. - `stream_options` (object): `{ "include_usage": true }` adds a final chunk with token usage. - `max_completion_tokens` (integer): The most tokens to generate, reasoning included. `max_tokens` works too. - `temperature` (number): 0 to 2. Lower is more focused, higher more varied. Some reasoning models ignore it. - `top_p` (number): Nucleus sampling, 0 to 1. Change this or temperature, not both. - `stop` (string | array): Up to 4 sequences where generation stops. - `tools` (array): Functions the model may call. See Tool calling below. - `tool_choice` (string | object): `auto`, `none`, `required`, or a specific function. - `response_format` (object): JSON output: `{ "type": "json_object" }`, or `json_schema` with a schema. See Structured output below. - `reasoning_effort` (string): For reasoning models: `low`, `medium` or `high` (some take `minimal` or `xhigh`). Higher thinks longer and costs more output tokens. - `seed` (integer): Best-effort reproducibility, for models that support it. - `models` (array): Smart routing for this request: alternatives to try, in order, if the model is down. See [Smart routing](https://zurelay.com/docs/smart-routing). > **Note:** Other OpenAI parameters pass through to the model unchanged, except `user` and `metadata`, which stay with us. Which ones a model honors (tools, JSON schemas, reasoning) is up to the model, as it is with its own lab. With the OpenAI Python SDK, send `models` as `extra_body={"models": [...]}`. ## Response Response: ```json { "id": "chatcmpl-AbC123...", "object": "chat.completion", "created": 1790870400, "model": "claude-opus-5-5", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Here's the summary..." }, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 1204, "completion_tokens": 312, "total_tokens": 1516, "prompt_tokens_details": { "cached_tokens": 1024 }, "completion_tokens_details": { "reasoning_tokens": 128 } } } ``` - `model` is the model that answered. It’s the one you asked for, unless smart routing answered with an alternative. - Reasoning models may add `reasoning_content` to the message: the model’s thinking, when it shares it. - `usage` is what you’re billed for. `cached_tokens` are input tokens read from the model’s prompt cache, billed at the cached price where a model has one. Every response also has these headers: | Header | Value | | --- | --- | | `x-request-id` | The request’s ID, as it appears in your request log. Quote it to support. | | `x-zurelay-model` | The model that answered. | | `x-zurelay-routed-from` | Only after smart routing: the model you asked for. | ## Streaming With `stream: true`, the answer arrives as server-sent events in OpenAI’s chunk format, ending with `data: [DONE]`. Every SDK reads these for you. **Python** (`main.py`): ```python import os from openai import OpenAI client = OpenAI( base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"], ) stream = client.chat.completions.create( model="claude-sonnet-5-5", messages=[ { "role": "user", "content": "Write a limerick about databases." } ], stream=True, stream_options={ "include_usage": True }, ) for chunk in stream: if chunk.choices and chunk.choices[0].delta.content: print(chunk.choices[0].delta.content, end="", flush=True) ``` **Node.js** (`index.mjs`): ```javascript import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.zurelay.com/v1", apiKey: process.env.ZURELAY_API_KEY, }); const stream = await client.chat.completions.create({ "model": "claude-sonnet-5-5", "messages": [ { "role": "user", "content": "Write a limerick about databases." } ], "stream": true, "stream_options": { "include_usage": true } }); for await (const chunk of stream) { process.stdout.write(chunk.choices[0]?.delta?.content ?? ""); } ``` **cURL**: ```bash curl https://api.zurelay.com/v1/chat/completions \ -H "Authorization: Bearer $ZURELAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "claude-sonnet-5-5", "messages": [ { "role": "user", "content": "Write a limerick about databases." } ], "stream": true, "stream_options": { "include_usage": true } }' ``` - Nothing is sent until the model has produced its first token, so a retry or a switch to another path stays invisible to you. - Once the stream has started, quiet stretches (a model thinking between steps) are filled with SSE comment lines (`: keep-alive`) every 15 seconds. Clients ignore them; they keep proxies from closing the connection. Before the first token nothing is sent, not even headers: allow up to about 100 seconds, or 4 minutes at high reasoning effort. - With `include_usage`, the last chunk before `[DONE]` has an empty `choices` list and the `usage`. ## Tool calling Describe functions in `tools`; when the model wants one, it answers with `tool_calls` instead of text. Run the function, send the result back as a `tool` message with the call’s ID, and the model continues. **Python** (`main.py`): ```python import os from openai import OpenAI client = OpenAI( base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"], ) response = client.chat.completions.create( model="gpt-6-sol", messages=[ { "role": "user", "content": "What's the weather in Lisbon right now?" } ], tools=[ { "type": "function", "function": { "name": "get_weather", "description": "Current weather for a city", "parameters": { "type": "object", "properties": { "city": { "type": "string" } }, "required": [ "city" ] } } } ], ) print(response.choices[0].message.content) ``` **Node.js** (`index.mjs`): ```javascript import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.zurelay.com/v1", apiKey: process.env.ZURELAY_API_KEY, }); const response = await client.chat.completions.create({ "model": "gpt-6-sol", "messages": [ { "role": "user", "content": "What's the weather in Lisbon right now?" } ], "tools": [ { "type": "function", "function": { "name": "get_weather", "description": "Current weather for a city", "parameters": { "type": "object", "properties": { "city": { "type": "string" } }, "required": [ "city" ] } } } ] }); console.log(response.choices[0].message.content); ``` **cURL**: ```bash curl https://api.zurelay.com/v1/chat/completions \ -H "Authorization: Bearer $ZURELAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-6-sol", "messages": [ { "role": "user", "content": "What'\''s the weather in Lisbon right now?" } ], "tools": [ { "type": "function", "function": { "name": "get_weather", "description": "Current weather for a city", "parameters": { "type": "object", "properties": { "city": { "type": "string" } }, "required": [ "city" ] } } } ] }' ``` ## Structured output Ask for JSON with `response_format`. A `json_schema` holds models that support it to your schema; `json_object` asks for any valid JSON (say so in your prompt too). **Python** (`main.py`): ```python import os from openai import OpenAI client = OpenAI( base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"], ) response = client.chat.completions.create( model="gpt-6.1-sol", messages=[ { "role": "user", "content": "Extract the people: 'Ana met Ben and Chloe in Porto.'" } ], response_format={ "type": "json_schema", "json_schema": { "name": "people", "schema": { "type": "object", "properties": { "names": { "type": "array", "items": { "type": "string" } } }, "required": [ "names" ] } } }, ) print(response.choices[0].message.content) ``` **Node.js** (`index.mjs`): ```javascript import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.zurelay.com/v1", apiKey: process.env.ZURELAY_API_KEY, }); const response = await client.chat.completions.create({ "model": "gpt-6.1-sol", "messages": [ { "role": "user", "content": "Extract the people: 'Ana met Ben and Chloe in Porto.'" } ], "response_format": { "type": "json_schema", "json_schema": { "name": "people", "schema": { "type": "object", "properties": { "names": { "type": "array", "items": { "type": "string" } } }, "required": [ "names" ] } } } }); console.log(response.choices[0].message.content); ``` **cURL**: ```bash curl https://api.zurelay.com/v1/chat/completions \ -H "Authorization: Bearer $ZURELAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-6.1-sol", "messages": [ { "role": "user", "content": "Extract the people: '\''Ana met Ben and Chloe in Porto.'\''" } ], "response_format": { "type": "json_schema", "json_schema": { "name": "people", "schema": { "type": "object", "properties": { "names": { "type": "array", "items": { "type": "string" } } }, "required": [ "names" ] } } } }' ``` ## Reasoning Reasoning models think before they answer. Set `reasoning_effort` to trade speed and cost for depth. Thinking counts as output tokens (`completion_tokens_details.reasoning_tokens`) and can take minutes at high effort: give your client a generous timeout (see [Reliability and timeouts](https://zurelay.com/docs/reliability)). ### Claude’s extended thinking Through this endpoint, Claude models take `reasoning_effort` too. For Anthropic’s own `thinking` parameter and thinking blocks, use [Anthropic Messages](https://zurelay.com/docs/anthropic).