Chat

Chat completions

OpenAI’s Chat Completions API, for every text model: the same request and response, the same streaming, the same SDKs.

POST/v1/chat/completions

Request

modelstringrequired
A model ID from the models page, such as claude-opus-5-5 or gpt-6-sol.
messagesarrayrequired
The conversation: system (or developer), user, assistant and tool messages. content is a string, or an array of parts to add images, PDFs, audio or video.
streamboolean
Send the answer as server-sent events while it’s written. See Streaming below.
stream_optionsobject
{ "include_usage": true } adds a final chunk with token usage.
max_completion_tokensinteger
The most tokens to generate, reasoning included. max_tokens works too.
temperaturenumber
0 to 2. Lower is more focused, higher more varied. Some reasoning models ignore it.
top_pnumber
Nucleus sampling, 0 to 1. Change this or temperature, not both.
stopstring | array
Up to 4 sequences where generation stops.
toolsarray
Functions the model may call. See Tool calling below.
tool_choicestring | object
auto, none, required, or a specific function.
response_formatobject
JSON output: { "type": "json_object" }, or json_schema with a schema. See Structured output below.
reasoning_effortstring
For reasoning models: low, medium or high (some take minimal or xhigh). Higher thinks longer and costs more output tokens.
seedinteger
Best-effort reproducibility, for models that support it.
modelsarray
Smart routing for this request: alternatives to try, in order, if the model is down. See Smart routing.
Other OpenAI parameters pass through to the model unchanged, except user and metadata, which stay with us. Which ones a model honors (tools, JSON schemas, reasoning) is up to the model, as it is with its own lab. With the OpenAI Python SDK, send models as extra_body={"models": [...]}.

Response

Response
{
"id": "chatcmpl-AbC123...",
"object": "chat.completion",
"created": 1790870400,
"model": "claude-opus-5-5",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Here's the summary..."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 1204,
"completion_tokens": 312,
"total_tokens": 1516,
"prompt_tokens_details": { "cached_tokens": 1024 },
"completion_tokens_details": { "reasoning_tokens": 128 }
}
}
  • model is the model that answered. It’s the one you asked for, unless smart routing answered with an alternative.
  • Reasoning models may add reasoning_content to the message: the model’s thinking, when it shares it.
  • usage is what you’re billed for. cached_tokens are input tokens read from the model’s prompt cache, billed at the cached price where a model has one.

Every response also has these headers:

HeaderValue
x-request-idThe request’s ID, as it appears in your request log. Quote it to support.
x-zurelay-modelThe model that answered.
x-zurelay-routed-fromOnly after smart routing: the model you asked for.

Streaming

With stream: true, the answer arrives as server-sent events in OpenAI’s chunk format, ending with data: [DONE]. Every SDK reads these for you.

import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key=os.environ["ZURELAY_API_KEY"],
)
stream = client.chat.completions.create(
model="claude-sonnet-5-5",
messages=[
{
"role": "user",
"content": "Write a limerick about databases."
}
],
stream=True,
stream_options={
"include_usage": True
},
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
  • Nothing is sent until the model has produced its first token, so a retry or a switch to another path stays invisible to you.
  • Once the stream has started, quiet stretches (a model thinking between steps) are filled with SSE comment lines (: keep-alive) every 15 seconds. Clients ignore them; they keep proxies from closing the connection. Before the first token nothing is sent, not even headers: allow up to about 100 seconds, or 4 minutes at high reasoning effort.
  • With include_usage, the last chunk before [DONE] has an empty choices list and the usage.

Tool calling

Describe functions in tools; when the model wants one, it answers with tool_calls instead of text. Run the function, send the result back as a tool message with the call’s ID, and the model continues.

import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key=os.environ["ZURELAY_API_KEY"],
)
response = client.chat.completions.create(
model="gpt-6-sol",
messages=[
{
"role": "user",
"content": "What's the weather in Lisbon right now?"
}
],
tools=[
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather for a city",
"parameters": {
"type": "object",
"properties": {
"city": {
"type": "string"
}
},
"required": [
"city"
]
}
}
}
],
)
print(response.choices[0].message.content)

Structured output

Ask for JSON with response_format. A json_schema holds models that support it to your schema; json_object asks for any valid JSON (say so in your prompt too).

import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key=os.environ["ZURELAY_API_KEY"],
)
response = client.chat.completions.create(
model="gpt-6.1-sol",
messages=[
{
"role": "user",
"content": "Extract the people: 'Ana met Ben and Chloe in Porto.'"
}
],
response_format={
"type": "json_schema",
"json_schema": {
"name": "people",
"schema": {
"type": "object",
"properties": {
"names": {
"type": "array",
"items": {
"type": "string"
}
}
},
"required": [
"names"
]
}
}
},
)
print(response.choices[0].message.content)

Reasoning

Reasoning models think before they answer. Set reasoning_effort to trade speed and cost for depth. Thinking counts as output tokens (completion_tokens_details.reasoning_tokens) and can take minutes at high effort: give your client a generous timeout (see Reliability and timeouts).

Claude’s extended thinking

Through this endpoint, Claude models take reasoning_effort too. For Anthropic’s own thinking parameter and thinking blocks, use Anthropic Messages.

Questions, or something missing? Ask support in your dashboard or email support@zurelay.com.