Chat
Images, PDFs, audio and video input
Put pictures, PDFs, audio and video in a message for the models that read them. Send the file itself (base64) or a link: we fetch links ourselves, so any public https link to a file up to 20 MB works.
At a glance
Every part takes a public https link or the file itself as a base64 data URL, such as data:application/pdf;base64,JVBERi0.... Several files can go in one message, in any mix the model reads.
Which models read what
Every chat model reads text. This table is live: it shows what each model reads today. A model given something it doesn’t read refuses it straight away with unsupported_content, so it never answers about a file it didn’t see.
Images
Add an image_url part with a link or a data URL. PNG, JPEG, WebP and GIF. Several images in one message work on every model that reads images.
import osfrom openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"],)response = client.chat.completions.create( model="gemini-3.7-flash", messages=[ { "role": "user", "content": [ { "type": "text", "text": "What's in this picture? One sentence." }, { "type": "image_url", "image_url": { "url": "https://zurelay.com/examples/gpt-image-2.5-flare/watch-800.webp" } } ] } ],)print(response.choices[0].message.content)To send a file from disk, base64-encode it into a data URL:
import base64, osfrom openai import OpenAIclient = OpenAI(base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"])with open("receipt.jpg", "rb") as f: data_url = "data:image/jpeg;base64," + base64.b64encode(f.read()).decode()response = client.chat.completions.create( model="claude-sonnet-5-5", messages=[{ "role": "user", "content": [ {"type": "text", "text": "What's the total on this receipt?"}, {"type": "image_url", "image_url": {"url": data_url}}, ], }],)print(response.choices[0].message.content)detail (low, high, auto) passes through for models that use it.
PDFs
Add a file part with the PDF as a data URL in file_data, and its file name. Models read the text and, where they can, the pages as pictures (charts, scans, tables). A public link works in file_data too.
import base64, osfrom openai import OpenAIclient = OpenAI(base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"])with open("contract.pdf", "rb") as f: pdf = "data:application/pdf;base64," + base64.b64encode(f.read()).decode()response = client.chat.completions.create( model="gemini-3.7-flash", messages=[{ "role": "user", "content": [ {"type": "text", "text": "Who are the parties, and how can the contract be ended?"}, {"type": "file", "file": {"filename": "contract.pdf", "file_data": pdf}}, ], }],)print(response.choices[0].message.content)Long PDFs cost input tokens per page (see billing below), so send the pages you need. Ask about several documents at once by adding a file part for each, with names that tell them apart.
Word, Excel, text and code
Models read PDFs, not Word, Excel or PowerPoint files. Put text, CSV, JSON, Markdown or code straight into the message as text: it’s billed as ordinary input tokens, and it’s what the model reads best. When the layout matters (tables, charts, scans), export the file to PDF and send that. A data URL of any other type, such as text/plain, is refused with an error that says so.
import osfrom pathlib import Pathfrom openai import OpenAIclient = OpenAI(base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"])notes = Path("meeting-notes.md").read_text()sales = Path("sales.csv").read_text()response = client.chat.completions.create( model="deepseek-v4.1-flash", messages=[{ "role": "user", "content": f"Summarize the notes, then total the sales by region.\n\n" f"<notes>\n{notes}\n</notes>\n\n<sales>\n{sales}\n</sales>", }],)print(response.choices[0].message.content)Audio
For models that hear (see the table): an input_audio part with the base64 audio and its format, wav or mp3. input_audio takes the audio itself; for a link, use a file part with the link in file_data. Recordings made in a browser are usually WebM, which is read as video: convert them to MP3 or WAV for models that only hear.
import base64, osfrom openai import OpenAIclient = OpenAI(base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"])with open("call.mp3", "rb") as f: audio = base64.b64encode(f.read()).decode()response = client.chat.completions.create( model="gemini-3.7-flash", messages=[{ "role": "user", "content": [ {"type": "text", "text": "Transcribe this call, then list the action items."}, {"type": "input_audio", "input_audio": {"data": audio, "format": "mp3"}}, ], }],)print(response.choices[0].message.content)Video
For models that watch video: a video_url part with a link or a data URL (a file part works too). MP4, MOV and WebM. Models with sound read the soundtrack as well, so you can ask what’s said.
import osfrom openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"],)response = client.chat.completions.create( model="gemini-3.7-flash", messages=[ { "role": "user", "content": [ { "type": "text", "text": "Describe what happens, and transcribe anything that's said." }, { "type": "video_url", "video_url": { "url": "https://example.com/clip.mp4" } } ] } ],)print(response.choices[0].message.content)A clip from disk goes the same way, as a data URL:
with open("clip.mp4", "rb") as f: video = "data:video/mp4;base64," + base64.b64encode(f.read()).decode()content = [ {"type": "text", "text": "What happens in this clip?"}, {"type": "video_url", "video_url": {"url": video}},]One shape in, the right shape out
video_url, an image_url or a file part, depending on whose docs you read). Send any of them: we hand each model its media in the shape it reads.Several files in one message
Add a part for each file, in the order you want them read, and refer to them in your text by order or by name (“the second image”, “contract.pdf”). Pictures, PDFs, audio and video can be mixed as long as the model reads each kind. Text parts can sit between files to label them.
Links
- Links must be public
httpsURLs. We follow up to three redirects and wait up to 20 seconds. - A link that doesn’t download is refused straight away with
invalid_urland the status it answered, before any model is called. - Signed links (S3, Cloud Storage, our own image links) work while they’re valid.
Anthropic format
On /v1/messages, Claude models take Anthropic’s blocks: image and document, with a base64 or url source. A few Claude models read PDFs only through /v1/chat/completions; send one to /v1/messages and the error says where to send it instead.
[ { "type": "image", "source": { "type": "base64", "media_type": "image/png", "data": "iVBORw0KGgo..." } }, { "type": "document", "source": { "type": "url", "url": "https://example.com/report.pdf" } }, { "type": "document", "source": { "type": "base64", "media_type": "application/pdf", "data": "JVBERi0..." } }, { "type": "text", "text": "Compare the chart with the report's conclusion." }]Limits
For bigger files, send a link: it doesn’t count toward the request body. Files uploaded to a lab’s own Files API (file_id) can’t be used; send the file or a link instead.
Errors
The first two are the error’s code; the others are its type, with a message that says what to change. All of them come back before any model is called, so they cost nothing. See Errors for the rest.
How media are billed
Models turn media into input tokens, billed at the model’s input price. Each lab counts its own way; these are typical figures we measured:
Your request log shows exactly what each request counted. Shrinking pictures to the size you need is the easiest saving: most models read a 1024 px picture as well as a 4000 px one.
Questions, or something missing? Ask support in your dashboard or email support@zurelay.com.