# Text to speech > POST /v1/audio/speech: Gemini text-to-speech through OpenAI's speech API, with 30 voices, styles in plain words, MP3, WAV or PCM, and prices. Source: https://zurelay.com/docs/speech Turn text into natural speech with Gemini’s text-to-speech models, through OpenAI’s speech API: 30 voices, a style you describe in plain words, and MP3, WAV or PCM back in a few seconds. `POST /v1/audio/speech` | Model | ID | Text in, per 1M tokens | Audio out, per 1M tokens | Audio tokens a second | About a minute | | --- | --- | --- | --- | --- | --- | | [Gemini 3.1 Flash TTS](https://zurelay.com/models/gemini-3.1-flash-tts) | `gemini-3.1-flash-tts` | $0.60 | $12.00 | 32 | $0.023 | | [Gemini 2.5 Pro TTS](https://zurelay.com/models/gemini-2.5-pro-tts) | `gemini-2.5-pro-tts` | $0.60 | $12.00 | 25 | $0.018 | | [Gemini 2.5 Flash TTS](https://zurelay.com/models/gemini-2.5-flash-tts) | `gemini-2.5-flash-tts` | $0.30 | $6.00 | 25 | $0.009 | ## Make speech **Python** (`speak.py`): ```python from openai import OpenAI client = OpenAI(base_url="https://api.zurelay.com/v1", api_key="YOUR_ZURELAY_KEY") speech = client.audio.speech.create( model="gemini-3.1-flash-tts", voice="Kore", input="Your order has shipped and arrives on Friday.", instructions="Say it warmly, at a relaxed pace", ) speech.write_to_file("shipped.mp3") ``` **Node.js** (`speak.mjs`): ```javascript import OpenAI from "openai"; import { writeFile } from "node:fs/promises"; const client = new OpenAI({ baseURL: "https://api.zurelay.com/v1", apiKey: process.env.ZURELAY_API_KEY }); const speech = await client.audio.speech.create({ model: "gemini-3.1-flash-tts", voice: "Kore", input: "Your order has shipped and arrives on Friday.", instructions: "Say it warmly, at a relaxed pace", }); await writeFile("shipped.mp3", Buffer.from(await speech.arrayBuffer())); ``` **cURL**: ```bash curl https://api.zurelay.com/v1/audio/speech \ -H "Authorization: Bearer $ZURELAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "gemini-3.1-flash-tts", "voice": "Kore", "input": "Your order has shipped and arrives on Friday.", "response_format": "wav" }' -o shipped.wav ``` Any OpenAI SDK works: point it at `https://api.zurelay.com/v1` and call its speech method. The reply is the audio file itself. ## Parameters - `model` (string, required): A speech model ID from the table. - `input` (string, required): The text to speak, up to 5,000 characters. Split longer text into several requests. - `voice` (string): One of the 30 voices below, or an OpenAI voice name such as `alloy` or `nova`. Default `Kore`. - `instructions` (string): How to say it, in plain words: `Say cheerfully`, `Whisper, like a secret`, `Read slowly, like a bedtime story`. Up to 2,000 characters. - `response_format` (string): `mp3` (the default), `wav`, or `pcm`: raw 24 kHz, 16-bit, mono samples. - `speed` (number): Not supported: ask for the pace in `instructions` instead. Only `1` is accepted. - `models` (array): Smart routing for this request: other speech models to try, in order, if the model is down. See [Smart routing](https://zurelay.com/docs/smart-routing). > **Tip:** The audio comes back when all of it is made. Measured at full length (5,000 characters): the Flash models take about a minute, and Gemini 2.5 Pro TTS about five and a half. Give your HTTP client a timeout of 10 minutes (the OpenAI SDKs’ default), or send long text in parts. ## Voices | Voice | Sounds | Voice | Sounds | | --- | --- | --- | --- | | `Zephyr` | Bright | `Puck` | Upbeat | | `Charon` | Informative | `Kore` | Firm | | `Fenrir` | Excitable | `Leda` | Youthful | | `Orus` | Firm | `Aoede` | Breezy | | `Callirrhoe` | Easy-going | `Autonoe` | Bright | | `Enceladus` | Breathy | `Iapetus` | Clear | | `Umbriel` | Easy-going | `Algieba` | Smooth | | `Despina` | Smooth | `Erinome` | Clear | | `Algenib` | Gravelly | `Rasalgethi` | Informative | | `Laomedeia` | Upbeat | `Achernar` | Soft | | `Alnilam` | Firm | `Schedar` | Even | | `Gacrux` | Mature | `Pulcherrima` | Forward | | `Achird` | Friendly | `Zubenelgenubi` | Casual | | `Vindemiatrix` | Gentle | `Sadachbia` | Lively | | `Sadaltager` | Knowledgeable | `Sulafat` | Warm | Code written for OpenAI keeps working: its voice names map to a Gemini voice of a similar character. | OpenAI voice | Speaks as | | --- | --- | | `alloy` | Kore | | `ash` | Charon | | `ballad` | Algieba | | `coral` | Aoede | | `echo` | Puck | | `fable` | Fenrir | | `nova` | Leda | | `onyx` | Orus | | `sage` | Zephyr | | `shimmer` | Callirrhoe | | `verse` | Umbriel | | `marin` | Despina | | `cedar` | Iapetus | ## Style and pacing - Describe the delivery the way you’d brief a voice actor: mood, pace, volume, accent. - Keep the direction in `instructions` and only the words to say in `input`, so the direction isn’t read out. - The language is picked up from the text, so a Spanish sentence is spoken in Spanish. ## Paying for speech Speech is billed on tokens: the text you send at the input price, and the audio you get back at the output price. Each second of audio is a set number of tokens, which the table lists for each model, so its last column is what a minute costs. Each reply says what it was billed on in its `x-zurelay-input-tokens` and `x-zurelay-output-tokens` headers, and every request is in your request log. Requests that fail cost nothing. With [smart routing](https://zurelay.com/docs/smart-routing) on, another speech model may answer when the one you asked for is down: `x-zurelay-model` names it, and it’s billed at its price. > **Tip:** Speech isn’t stored: save the file you get back. The same text and voice give a slightly different reading each time. Try voices and styles in the dashboard’s [Playground](https://zurelay.com/app/playground).