Chat
Chat completions
OpenAI’s Chat Completions API, for every text model: the same request and response, the same streaming, the same SDKs.
/v1/chat/completionsRequest
modelstringrequired- A model ID from the models page, such as
claude-opus-5-5orgpt-6-sol. messagesarrayrequired- The conversation:
system(ordeveloper),user,assistantandtoolmessages.contentis a string, or an array of parts to add images, PDFs, audio or video. streamboolean- Send the answer as server-sent events while it’s written. See Streaming below.
stream_optionsobject{ "include_usage": true }adds a final chunk with token usage.max_completion_tokensinteger- The most tokens to generate, reasoning included.
max_tokensworks too. temperaturenumber- 0 to 2. Lower is more focused, higher more varied. Some reasoning models ignore it.
top_pnumber- Nucleus sampling, 0 to 1. Change this or temperature, not both.
stopstring | array- Up to 4 sequences where generation stops.
toolsarray- Functions the model may call. See Tool calling below.
tool_choicestring | objectauto,none,required, or a specific function.response_formatobject- JSON output:
{ "type": "json_object" }, orjson_schemawith a schema. See Structured output below. reasoning_effortstring- For reasoning models:
low,mediumorhigh(some takeminimalorxhigh). Higher thinks longer and costs more output tokens. seedinteger- Best-effort reproducibility, for models that support it.
modelsarray- Smart routing for this request: alternatives to try, in order, if the model is down. See Smart routing.
user and metadata, which stay with us. Which ones a model honors (tools, JSON schemas, reasoning) is up to the model, as it is with its own lab. With the OpenAI Python SDK, send models as extra_body={"models": [...]}.Response
{ "id": "chatcmpl-AbC123...", "object": "chat.completion", "created": 1790870400, "model": "claude-opus-5-5", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Here's the summary..." }, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 1204, "completion_tokens": 312, "total_tokens": 1516, "prompt_tokens_details": { "cached_tokens": 1024 }, "completion_tokens_details": { "reasoning_tokens": 128 } }}modelis the model that answered. It’s the one you asked for, unless smart routing answered with an alternative.- Reasoning models may add
reasoning_contentto the message: the model’s thinking, when it shares it. usageis what you’re billed for.cached_tokensare input tokens read from the model’s prompt cache, billed at the cached price where a model has one.
Every response also has these headers:
Streaming
With stream: true, the answer arrives as server-sent events in OpenAI’s chunk format, ending with data: [DONE]. Every SDK reads these for you.
import osfrom openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"],)stream = client.chat.completions.create( model="claude-sonnet-5-5", messages=[ { "role": "user", "content": "Write a limerick about databases." } ], stream=True, stream_options={ "include_usage": True },)for chunk in stream: if chunk.choices and chunk.choices[0].delta.content: print(chunk.choices[0].delta.content, end="", flush=True)- Nothing is sent until the model has produced its first token, so a retry or a switch to another path stays invisible to you.
- Once the stream has started, quiet stretches (a model thinking between steps) are filled with SSE comment lines (
: keep-alive) every 15 seconds. Clients ignore them; they keep proxies from closing the connection. Before the first token nothing is sent, not even headers: allow up to about 100 seconds, or 4 minutes at high reasoning effort. - With
include_usage, the last chunk before[DONE]has an emptychoiceslist and theusage.
Tool calling
Describe functions in tools; when the model wants one, it answers with tool_calls instead of text. Run the function, send the result back as a tool message with the call’s ID, and the model continues.
import osfrom openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"],)response = client.chat.completions.create( model="gpt-6-sol", messages=[ { "role": "user", "content": "What's the weather in Lisbon right now?" } ], tools=[ { "type": "function", "function": { "name": "get_weather", "description": "Current weather for a city", "parameters": { "type": "object", "properties": { "city": { "type": "string" } }, "required": [ "city" ] } } } ],)print(response.choices[0].message.content)Structured output
Ask for JSON with response_format. A json_schema holds models that support it to your schema; json_object asks for any valid JSON (say so in your prompt too).
import osfrom openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"],)response = client.chat.completions.create( model="gpt-6.1-sol", messages=[ { "role": "user", "content": "Extract the people: 'Ana met Ben and Chloe in Porto.'" } ], response_format={ "type": "json_schema", "json_schema": { "name": "people", "schema": { "type": "object", "properties": { "names": { "type": "array", "items": { "type": "string" } } }, "required": [ "names" ] } } },)print(response.choices[0].message.content)Reasoning
Reasoning models think before they answer. Set reasoning_effort to trade speed and cost for depth. Thinking counts as output tokens (completion_tokens_details.reasoning_tokens) and can take minutes at high effort: give your client a generous timeout (see Reliability and timeouts).
Claude’s extended thinking
Through this endpoint, Claude models take reasoning_effort too. For Anthropic’s own thinking parameter and thinking blocks, use Anthropic Messages.
Questions, or something missing? Ask support in your dashboard or email support@zurelay.com.