# Images, PDFs, audio and video input > Send pictures, documents, recordings and clips to the models that read them, as data or links, in OpenAI or Anthropic format. Source: https://zurelay.com/docs/multimodal Put pictures, PDFs, audio and video in a message for the models that read them. Send the file itself (base64) or a link: we fetch links ourselves, so any public `https` link to a file up to 20 MB works. ## At a glance | To send | /v1/chat/completions (every model) | /v1/messages (Claude) | Formats | | --- | --- | --- | --- | | A picture | `image_url` part | `image` block | PNG, JPEG, WebP, GIF | | A PDF | `file` part | `document` block | PDF | | A recording | `input_audio` part | Claude doesn’t hear audio | WAV, MP3 | | A video | `video_url` part | Claude doesn’t watch video | MP4, MOV, WebM | | Text, code, CSV, Word, Excel | The text itself, in the message | A `text` block, or a plain-text `document` | Text | Every part takes a public `https` link or the file itself as a base64 data URL, such as `data:application/pdf;base64,JVBERi0...`. Several files can go in one message, in any mix the model reads. ## Which models read what Every chat model reads text. This table is live: it shows what each model reads today. A model given something it doesn’t read refuses it straight away with `unsupported_content`, so it never answers about a file it didn’t see. | Model | ID | Images | PDFs | Audio | Video | | --- | --- | --- | --- | --- | --- | | Claude Opus 5.5 | `claude-opus-5-5` | Yes | Yes | No | No | | GPT-6 Astra | `gpt-6-astra` | Yes | Yes | No | No | | Claude Sonnet 5.5 | `claude-sonnet-5-5` | Yes | Yes | No | No | | GPT-6.1 Sol | `gpt-6.1-sol` | Yes | Yes | No | No | | GPT-6 Sol | `gpt-6-sol` | Yes | Yes | No | No | | GPT-6 Luna | `gpt-6-luna` | Yes | Yes | No | No | | Grok 4.7 | `grok-4.7` | Yes | No | No | No | | MiMo V2.6 Pro | `mimo-v2.6-pro` | Yes | No | Yes | No | | MiMo V2.6 Flash | `mimo-v2.6-flash` | Yes | No | Yes | No | | DeepSeek V4.1 Flash | `deepseek-v4.1-flash` | Yes | No | No | No | | DeepSeek V4 Flash | `deepseek-v4-flash` | Yes | No | No | No | | DeepSeek V4 Flash 0731 | `deepseek-v4-flash-0731` | No | No | No | No | | DeepSeek V4 Pro 0813 | `deepseek-v4-pro-0813` | No | No | No | No | | Claude Fable 5.1 | `claude-fable-5-1` | Yes | Yes | No | No | | Claude Fable 5 | `claude-fable-5` | Yes | No | No | No | | Claude Opus 5 | `claude-opus-5` | Yes | Yes | No | No | | Claude Opus 4.8 | `claude-opus-4-8` | Yes | Yes | No | No | | Claude Opus 4.7 | `claude-opus-4-7` | Yes | No | No | No | | Claude Opus 4.6 | `claude-opus-4-6` | Yes | No | No | No | | Claude Haiku 4.5 | `claude-haiku-4-5` | Yes | Yes | No | No | | GPT-5.6 Sol | `gpt-5.6-sol` | Yes | Yes | No | No | | GPT-5.6 Terra | `gpt-5.6-terra` | Yes | Yes | No | No | | GPT-5.6 Luna | `gpt-5.6-luna` | Yes | Yes | No | No | | Gemini 3.7 Flash | `gemini-3.7-flash` | Yes | Yes | Yes | Yes | | Gemini 3.5 Flash-Lite | `gemini-3.5-flash-lite` | Yes | Yes | Yes | Yes | | Gemini 3.1 Flash-Lite | `gemini-3.1-flash-lite` | Yes | Yes | Yes | Yes | | GLM-5.3 | `glm-5.3` | No | No | No | No | | GLM-5.2 | `glm-5.2` | No | No | No | No | | Kimi K3 | `kimi-k3` | Yes | No | No | No | | Qwen 3.8 Max | `qwen3.8-max` | Yes | Yes | No | Yes | | Qwen 3.8 Flash | `qwen3.8-flash` | Yes | Yes | No | Yes | | Qwen 3.8 Omni Flash | `qwen3.8-omni-flash` | Yes | No | No | Yes | | HY4 Preview | `hy4-preview` | Yes | No | No | No | | MiniMax M2.7 | `minimax-m2.7` | No | No | No | No | ## Images Add an `image_url` part with a link or a data URL. PNG, JPEG, WebP and GIF. Several images in one message work on every model that reads images. **Python** (`main.py`): ```python import os from openai import OpenAI client = OpenAI( base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"], ) response = client.chat.completions.create( model="gemini-3.7-flash", messages=[ { "role": "user", "content": [ { "type": "text", "text": "What's in this picture? One sentence." }, { "type": "image_url", "image_url": { "url": "https://zurelay.com/examples/gpt-image-2.5-flare/watch-800.webp" } } ] } ], ) print(response.choices[0].message.content) ``` **Node.js** (`index.mjs`): ```javascript import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.zurelay.com/v1", apiKey: process.env.ZURELAY_API_KEY, }); const response = await client.chat.completions.create({ "model": "gemini-3.7-flash", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "What's in this picture? One sentence." }, { "type": "image_url", "image_url": { "url": "https://zurelay.com/examples/gpt-image-2.5-flare/watch-800.webp" } } ] } ] }); console.log(response.choices[0].message.content); ``` **cURL**: ```bash curl https://api.zurelay.com/v1/chat/completions \ -H "Authorization: Bearer $ZURELAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "gemini-3.7-flash", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "What'\''s in this picture? One sentence." }, { "type": "image_url", "image_url": { "url": "https://zurelay.com/examples/gpt-image-2.5-flare/watch-800.webp" } } ] } ] }' ``` To send a file from disk, base64-encode it into a data URL: **Python** (`main.py`): ```python import base64, os from openai import OpenAI client = OpenAI(base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"]) with open("receipt.jpg", "rb") as f: data_url = "data:image/jpeg;base64," + base64.b64encode(f.read()).decode() response = client.chat.completions.create( model="claude-sonnet-5-5", messages=[{ "role": "user", "content": [ {"type": "text", "text": "What's the total on this receipt?"}, {"type": "image_url", "image_url": {"url": data_url}}, ], }], ) print(response.choices[0].message.content) ``` **Node.js** (`index.mjs`): ```javascript import { readFile } from "node:fs/promises"; import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.zurelay.com/v1", apiKey: process.env.ZURELAY_API_KEY }); const dataUrl = "data:image/jpeg;base64," + (await readFile("receipt.jpg")).toString("base64"); const response = await client.chat.completions.create({ model: "claude-sonnet-5-5", messages: [{ role: "user", content: [ { type: "text", text: "What's the total on this receipt?" }, { type: "image_url", image_url: { url: dataUrl } }, ], }], }); console.log(response.choices[0].message.content); ``` **cURL**: ```bash # The request goes in a file: a base64 image is too long for a command-line argument. cat > request.json < request.json <\n{notes}\n\n\n\n{sales}\n", }], ) print(response.choices[0].message.content) ``` ## Audio For models that hear (see the table): an `input_audio` part with the base64 audio and its format, `wav` or `mp3`. `input_audio` takes the audio itself; for a link, use a `file` part with the link in `file_data`. Recordings made in a browser are usually WebM, which is read as video: convert them to MP3 or WAV for models that only hear. **Python** (`main.py`): ```python import base64, os from openai import OpenAI client = OpenAI(base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"]) with open("call.mp3", "rb") as f: audio = base64.b64encode(f.read()).decode() response = client.chat.completions.create( model="gemini-3.7-flash", messages=[{ "role": "user", "content": [ {"type": "text", "text": "Transcribe this call, then list the action items."}, {"type": "input_audio", "input_audio": {"data": audio, "format": "mp3"}}, ], }], ) print(response.choices[0].message.content) ``` **Node.js** (`index.mjs`): ```javascript import { readFile } from "node:fs/promises"; import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.zurelay.com/v1", apiKey: process.env.ZURELAY_API_KEY }); const audio = (await readFile("call.mp3")).toString("base64"); const response = await client.chat.completions.create({ model: "gemini-3.7-flash", messages: [{ role: "user", content: [ { type: "text", text: "Transcribe this call, then list the action items." }, { type: "input_audio", input_audio: { data: audio, format: "mp3" } }, ], }], }); console.log(response.choices[0].message.content); ``` ## Video For models that watch video: a `video_url` part with a link or a data URL (a `file` part works too). MP4, MOV and WebM. Models with sound read the soundtrack as well, so you can ask what’s said. **Python** (`main.py`): ```python import os from openai import OpenAI client = OpenAI( base_url="https://api.zurelay.com/v1", api_key=os.environ["ZURELAY_API_KEY"], ) response = client.chat.completions.create( model="gemini-3.7-flash", messages=[ { "role": "user", "content": [ { "type": "text", "text": "Describe what happens, and transcribe anything that's said." }, { "type": "video_url", "video_url": { "url": "https://example.com/clip.mp4" } } ] } ], ) print(response.choices[0].message.content) ``` **Node.js** (`index.mjs`): ```javascript import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.zurelay.com/v1", apiKey: process.env.ZURELAY_API_KEY, }); const response = await client.chat.completions.create({ "model": "gemini-3.7-flash", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe what happens, and transcribe anything that's said." }, { "type": "video_url", "video_url": { "url": "https://example.com/clip.mp4" } } ] } ] }); console.log(response.choices[0].message.content); ``` **cURL**: ```bash curl https://api.zurelay.com/v1/chat/completions \ -H "Authorization: Bearer $ZURELAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "gemini-3.7-flash", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe what happens, and transcribe anything that'\''s said." }, { "type": "video_url", "video_url": { "url": "https://example.com/clip.mp4" } } ] } ] }' ``` A clip from disk goes the same way, as a data URL: Python: ```python with open("clip.mp4", "rb") as f: video = "data:video/mp4;base64," + base64.b64encode(f.read()).decode() content = [ {"type": "text", "text": "What happens in this clip?"}, {"type": "video_url", "video_url": {"url": video}}, ] ``` > **One shape in, the right shape out:** Labs take media in different shapes (a video can be a `video_url`, an `image_url` or a `file` part, depending on whose docs you read). Send any of them: we hand each model its media in the shape it reads. ## Several files in one message Add a part for each file, in the order you want them read, and refer to them in your text by order or by name (“the second image”, “contract.pdf”). Pictures, PDFs, audio and video can be mixed as long as the model reads each kind. Text parts can sit between files to label them. ## Links - Links must be public `https` URLs. We follow up to three redirects and wait up to 20 seconds. - A link that doesn’t download is refused straight away with `invalid_url` and the status it answered, before any model is called. - Signed links (S3, Cloud Storage, our own image links) work while they’re valid. ## Anthropic format On [/v1/messages](https://zurelay.com/docs/anthropic), Claude models take Anthropic’s blocks: `image` and `document`, with a `base64` or `url` source. A few Claude models read PDFs only through `/v1/chat/completions`; send one to `/v1/messages` and the error says where to send it instead. Content blocks: ```json [ { "type": "image", "source": { "type": "base64", "media_type": "image/png", "data": "iVBORw0KGgo..." } }, { "type": "document", "source": { "type": "url", "url": "https://example.com/report.pdf" } }, { "type": "document", "source": { "type": "base64", "media_type": "application/pdf", "data": "JVBERi0..." } }, { "type": "text", "text": "Compare the chart with the report's conclusion." } ] ``` ## Limits | Limit | Value | | --- | --- | | One file | 20 MB | | All media in one request | 30 MB | | Request body (base64 adds about a third) | 25 MB | | Media parts per request | 100 | | Links per request | 20 | For bigger files, send a link: it doesn’t count toward the request body. Files uploaded to a lab’s own Files API (`file_id`) can’t be used; send the file or a link instead. ## Errors | Status | Error | When | | --- | --- | --- | | 400 | `unsupported_content` | The model doesn’t read that kind of file. Pick one that does from the table above. | | 400 | `invalid_url` | A link didn’t download. The message says why (its status, a timeout, not a media file). | | 400 | `invalid_request_error` | A part isn’t a data URL or an https link, is empty, or is a type models don’t read (Word, plain text). | | 413 | `invalid_request_error` | A file over 20 MB, more than 30 MB of media, or a request body over 25 MB. | The first two are the error’s `code`; the others are its `type`, with a message that says what to change. All of them come back before any model is called, so they cost nothing. See [Errors](https://zurelay.com/docs/errors) for the rest. ## How media are billed Models turn media into input tokens, billed at the model’s input price. Each lab counts its own way; these are typical figures we measured: | | Claude | GPT | Gemini | | --- | --- | --- | --- | | A 1024 × 1024 picture | about 1,400 tokens | about 1,250 to 1,650 | about 1,100 | | A large photo (12 MP) | up to about 4,800 | scales with pixels | about 1,100 to 1,200 | | One PDF page | about 1,600 to 2,500 | the page’s text | about 560 | | One second of audio | not read | not read | about 25 | | One second of video | not read | not read | about 100 to 300 | Your request log shows exactly what each request counted. Shrinking pictures to the size you need is the easiest saving: most models read a 1024 px picture as well as a 4000 px one. > **Tip:** Sending the same picture or document in many requests? Put it early in the conversation and keep that prefix the same: models with prompt caching bill repeated input at their cached price.