Speech

Text to speech

Turn text into natural speech with Gemini’s text-to-speech models, through OpenAI’s speech API: 30 voices, a style you describe in plain words, and MP3, WAV or PCM back in a few seconds.

POST/v1/audio/speech
ModelIDText in, per 1M tokensAudio out, per 1M tokensAudio tokens a secondAbout a minute
Gemini 3.1 Flash TTSgemini-3.1-flash-tts$0.60$12.0032$0.023
Gemini 2.5 Pro TTSgemini-2.5-pro-tts$0.60$12.0025$0.018
Gemini 2.5 Flash TTSgemini-2.5-flash-tts$0.30$6.0025$0.009

Make speech

from openai import OpenAI
client = OpenAI(base_url="https://api.zurelay.com/v1", api_key="YOUR_ZURELAY_KEY")
speech = client.audio.speech.create(
model="gemini-3.1-flash-tts",
voice="Kore",
input="Your order has shipped and arrives on Friday.",
instructions="Say it warmly, at a relaxed pace",
)
speech.write_to_file("shipped.mp3")

Any OpenAI SDK works: point it at https://api.zurelay.com/v1 and call its speech method. The reply is the audio file itself.

Parameters

modelstringrequired
A speech model ID from the table.
inputstringrequired
The text to speak, up to 5,000 characters. Split longer text into several requests.
voicestring
One of the 30 voices below, or an OpenAI voice name such as alloy or nova. Default Kore.
instructionsstring
How to say it, in plain words: Say cheerfully, Whisper, like a secret, Read slowly, like a bedtime story. Up to 2,000 characters.
response_formatstring
mp3 (the default), wav, or pcm: raw 24 kHz, 16-bit, mono samples.
speednumber
Not supported: ask for the pace in instructions instead. Only 1 is accepted.
modelsarray
Smart routing for this request: other speech models to try, in order, if the model is down. See Smart routing.
The audio comes back when all of it is made. Measured at full length (5,000 characters): the Flash models take about a minute, and Gemini 2.5 Pro TTS about five and a half. Give your HTTP client a timeout of 10 minutes (the OpenAI SDKs’ default), or send long text in parts.

Voices

VoiceSoundsVoiceSounds
ZephyrBrightPuckUpbeat
CharonInformativeKoreFirm
FenrirExcitableLedaYouthful
OrusFirmAoedeBreezy
CallirrhoeEasy-goingAutonoeBright
EnceladusBreathyIapetusClear
UmbrielEasy-goingAlgiebaSmooth
DespinaSmoothErinomeClear
AlgenibGravellyRasalgethiInformative
LaomedeiaUpbeatAchernarSoft
AlnilamFirmSchedarEven
GacruxMaturePulcherrimaForward
AchirdFriendlyZubenelgenubiCasual
VindemiatrixGentleSadachbiaLively
SadaltagerKnowledgeableSulafatWarm

Code written for OpenAI keeps working: its voice names map to a Gemini voice of a similar character.

OpenAI voiceSpeaks as
alloyKore
ashCharon
balladAlgieba
coralAoede
echoPuck
fableFenrir
novaLeda
onyxOrus
sageZephyr
shimmerCallirrhoe
verseUmbriel
marinDespina
cedarIapetus

Style and pacing

  • Describe the delivery the way you’d brief a voice actor: mood, pace, volume, accent.
  • Keep the direction in instructions and only the words to say in input, so the direction isn’t read out.
  • The language is picked up from the text, so a Spanish sentence is spoken in Spanish.

Paying for speech

Speech is billed on tokens: the text you send at the input price, and the audio you get back at the output price. Each second of audio is a set number of tokens, which the table lists for each model, so its last column is what a minute costs. Each reply says what it was billed on in its x-zurelay-input-tokens and x-zurelay-output-tokens headers, and every request is in your request log. Requests that fail cost nothing. With smart routing on, another speech model may answer when the one you asked for is down: x-zurelay-model names it, and it’s billed at its price.

Speech isn’t stored: save the file you get back. The same text and voice give a slightly different reading each time.

Try voices and styles in the dashboard’s Playground.

Questions, or something missing? Ask our support team and we’ll answer by email, usually within a minute.