Chat

Images, PDFs, audio and video input

Put pictures, PDFs, audio and video in a message for the models that read them. Send the file itself (base64) or a link: we fetch links ourselves, so any public https link to a file up to 20 MB works.

At a glance

To send/v1/chat/completions (every model)/v1/messages (Claude)Formats
A pictureimage_url partimage blockPNG, JPEG, WebP, GIF
A PDFfile partdocument blockPDF
A recordinginput_audio partClaude doesn’t hear audioWAV, MP3
A videovideo_url partClaude doesn’t watch videoMP4, MOV, WebM
Text, code, CSV, Word, ExcelThe text itself, in the messageA text block, or a plain-text documentText

Every part takes a public https link or the file itself as a base64 data URL, such as data:application/pdf;base64,JVBERi0.... Several files can go in one message, in any mix the model reads.

Which models read what

Every chat model reads text. This table is live: it shows what each model reads today. A model given something it doesn’t read refuses it straight away with unsupported_content, so it never answers about a file it didn’t see.

Images

Add an image_url part with a link or a data URL. PNG, JPEG, WebP and GIF. Several images in one message work on every model that reads images.

import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key=os.environ["ZURELAY_API_KEY"],
)
response = client.chat.completions.create(
model="gemini-3.7-flash",
messages=[
{
"role": "user",
"content": [
{
"type": "text",
"text": "What's in this picture? One sentence."
},
{
"type": "image_url",
"image_url": {
"url": "https://zurelay.com/examples/gpt-image-2.5-flare/watch-800.webp"
}
}
]
}
],
)
print(response.choices[0].message.content)

To send a file from disk, base64-encode it into a data URL:

import base64, os
from openai import OpenAI
client = OpenAI(base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"])
with open("receipt.jpg", "rb") as f:
data_url = "data:image/jpeg;base64," + base64.b64encode(f.read()).decode()
response = client.chat.completions.create(
model="claude-sonnet-5-5",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "What's the total on this receipt?"},
{"type": "image_url", "image_url": {"url": data_url}},
],
}],
)
print(response.choices[0].message.content)

detail (low, high, auto) passes through for models that use it.

PDFs

Add a file part with the PDF as a data URL in file_data, and its file name. Models read the text and, where they can, the pages as pictures (charts, scans, tables). A public link works in file_data too.

import base64, os
from openai import OpenAI
client = OpenAI(base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"])
with open("contract.pdf", "rb") as f:
pdf = "data:application/pdf;base64," + base64.b64encode(f.read()).decode()
response = client.chat.completions.create(
model="gemini-3.7-flash",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "Who are the parties, and how can the contract be ended?"},
{"type": "file", "file": {"filename": "contract.pdf", "file_data": pdf}},
],
}],
)
print(response.choices[0].message.content)

Long PDFs cost input tokens per page (see billing below), so send the pages you need. Ask about several documents at once by adding a file part for each, with names that tell them apart.

Word, Excel, text and code

Models read PDFs, not Word, Excel or PowerPoint files. Put text, CSV, JSON, Markdown or code straight into the message as text: it’s billed as ordinary input tokens, and it’s what the model reads best. When the layout matters (tables, charts, scans), export the file to PDF and send that. A data URL of any other type, such as text/plain, is refused with an error that says so.

import os
from pathlib import Path
from openai import OpenAI
client = OpenAI(base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"])
notes = Path("meeting-notes.md").read_text()
sales = Path("sales.csv").read_text()
response = client.chat.completions.create(
model="deepseek-v4.1-flash",
messages=[{
"role": "user",
"content": f"Summarize the notes, then total the sales by region.\n\n"
f"<notes>\n{notes}\n</notes>\n\n<sales>\n{sales}\n</sales>",
}],
)
print(response.choices[0].message.content)

Audio

For models that hear (see the table): an input_audio part with the base64 audio and its format, wav or mp3. input_audio takes the audio itself; for a link, use a file part with the link in file_data. Recordings made in a browser are usually WebM, which is read as video: convert them to MP3 or WAV for models that only hear.

import base64, os
from openai import OpenAI
client = OpenAI(base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"])
with open("call.mp3", "rb") as f:
audio = base64.b64encode(f.read()).decode()
response = client.chat.completions.create(
model="gemini-3.7-flash",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "Transcribe this call, then list the action items."},
{"type": "input_audio", "input_audio": {"data": audio, "format": "mp3"}},
],
}],
)
print(response.choices[0].message.content)

Video

For models that watch video: a video_url part with a link or a data URL (a file part works too). MP4, MOV and WebM. Models with sound read the soundtrack as well, so you can ask what’s said.

import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key=os.environ["ZURELAY_API_KEY"],
)
response = client.chat.completions.create(
model="gemini-3.7-flash",
messages=[
{
"role": "user",
"content": [
{
"type": "text",
"text": "Describe what happens, and transcribe anything that's said."
},
{
"type": "video_url",
"video_url": {
"url": "https://example.com/clip.mp4"
}
}
]
}
],
)
print(response.choices[0].message.content)

A clip from disk goes the same way, as a data URL:

Python
with open("clip.mp4", "rb") as f:
video = "data:video/mp4;base64," + base64.b64encode(f.read()).decode()
content = [
{"type": "text", "text": "What happens in this clip?"},
{"type": "video_url", "video_url": {"url": video}},
]

One shape in, the right shape out

Labs take media in different shapes (a video can be a video_url, an image_url or a file part, depending on whose docs you read). Send any of them: we hand each model its media in the shape it reads.

Several files in one message

Add a part for each file, in the order you want them read, and refer to them in your text by order or by name (“the second image”, “contract.pdf”). Pictures, PDFs, audio and video can be mixed as long as the model reads each kind. Text parts can sit between files to label them.

  • Links must be public https URLs. We follow up to three redirects and wait up to 20 seconds.
  • A link that doesn’t download is refused straight away with invalid_url and the status it answered, before any model is called.
  • Signed links (S3, Cloud Storage, our own image links) work while they’re valid.

Anthropic format

On /v1/messages, Claude models take Anthropic’s blocks: image and document, with a base64 or url source. A few Claude models read PDFs only through /v1/chat/completions; send one to /v1/messages and the error says where to send it instead.

Content blocks
[
{ "type": "image", "source": { "type": "base64", "media_type": "image/png", "data": "iVBORw0KGgo..." } },
{ "type": "document", "source": { "type": "url", "url": "https://example.com/report.pdf" } },
{ "type": "document", "source": { "type": "base64", "media_type": "application/pdf", "data": "JVBERi0..." } },
{ "type": "text", "text": "Compare the chart with the report's conclusion." }
]

Limits

LimitValue
One file20 MB
All media in one request30 MB
Request body (base64 adds about a third)25 MB
Media parts per request100
Links per request20

For bigger files, send a link: it doesn’t count toward the request body. Files uploaded to a lab’s own Files API (file_id) can’t be used; send the file or a link instead.

Errors

StatusErrorWhen
400unsupported_contentThe model doesn’t read that kind of file. Pick one that does from the table above.
400invalid_urlA link didn’t download. The message says why (its status, a timeout, not a media file).
400invalid_request_errorA part isn’t a data URL or an https link, is empty, or is a type models don’t read (Word, plain text).
413invalid_request_errorA file over 20 MB, more than 30 MB of media, or a request body over 25 MB.

The first two are the error’s code; the others are its type, with a message that says what to change. All of them come back before any model is called, so they cost nothing. See Errors for the rest.

How media are billed

Models turn media into input tokens, billed at the model’s input price. Each lab counts its own way; these are typical figures we measured:

ClaudeGPTGemini
A 1024 × 1024 pictureabout 1,400 tokensabout 1,250 to 1,650about 1,100
A large photo (12 MP)up to about 4,800scales with pixelsabout 1,100 to 1,200
One PDF pageabout 1,600 to 2,500the page’s textabout 560
One second of audionot readnot readabout 25
One second of videonot readnot readabout 100 to 300

Your request log shows exactly what each request counted. Shrinking pictures to the size you need is the easiest saving: most models read a 1024 px picture as well as a 4000 px one.

Sending the same picture or document in many requests? Put it early in the conversation and keep that prefix the same: models with prompt caching bill repeated input at their cached price.

Questions, or something missing? Ask support in your dashboard or email support@zurelay.com.