Every input in one model
Text, images, audio and video go into the same request, with no separate transcription or vision service to wire up.
The Qwen 3.8 Omni Flash API: audio, video and images in, text out, 50% below list
qwen3.8-omni-flashfrom openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key="YOUR_ZURELAY_KEY",)stream = client.chat.completions.create( model="qwen3.8-omni-flash", messages=[{"role": "user", "content": "Explain quicksort in two sentences."}], stream=True,)for chunk in stream: if chunk.choices: print(chunk.choices[0].delta.content or "", end="")Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.
Pricing
Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.
| Rate | zurelay | Alibaba list | You save |
|---|---|---|---|
Input per 1M tokens | $0.075 | 50%off | |
Output per 1M tokens | $0.235 | 50%off | |
Cached input per 1M tokens | $0.007 | — | — |
One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.
Move the sliders to your monthly usage.
$73.20 a year
Overview
Qwen 3.8 Omni Flash is Alibaba's multimodal Qwen 3.8 model for audio and video understanding: it takes text, images, audio and video in one request and answers in text. The zurelay Qwen 3.8 Omni Flash API serves it through one OpenAI-compatible endpoint with a 1M-token context window, at 50% below Alibaba's international list price.
Alibaba released Qwen 3.8 Omni Flash on September 18, 2026. It is built for audio and video understanding and content analysis: you send text, images, audio and video, and it returns text. The Qwen 3.8 Omni Flash API holds a 1M-token context window and writes up to 131,072 output tokens per response. On zurelay the model ID is qwen3.8-omni-flash, and it costs $0.075 per 1M input tokens and $0.235 per 1M output tokens.
Audio understanding covers 113 languages and dialects. Audio counts 7 tokens per second, so a minute of speech is 420 input tokens and an hour is about 25,200, and long recordings fit easily in the context window. zurelay bills every input token at $0.075 per 1M, audio included. Alibaba also lists tool calling and a thinking mode that is on by default, with adjustable reasoning effort.
The model writes text only. Alibaba offers spoken replies through a separate realtime model, which zurelay does not serve, so plan on text output, or add your own text-to-speech step.
qwen/qwen3.8-omni-flash and qwen-3-8-omni-flash work as aliases. You pay $0.075 and $0.235 per 1M tokens, against Alibaba's international list price of $0.15 and $0.47 for text, image and video input and text output.
Strengths
Text, images, audio and video go into the same request, with no separate transcription or vision service to wire up.
At 7 tokens per second of audio, an hour of speech is about 25,200 tokens, a small slice of the 1M-token window.
Alibaba lists audio understanding in 113 languages and dialects.
Alibaba lists Omni Flash at Qwen 3.8 Flash's price for text, image and video input, and on zurelay audio tokens bill at the same $0.075 per 1M as text.
Use cases
Summarize calls, pull out action items and tag topics straight from the audio, without a transcription step first.
Ask questions about video content, flag scenes or write descriptions from the picture and the soundtrack together.
Understand speech across many languages and answer in the language your app needs.
Classify and route large batches of audio clips, videos and images at Flash-tier cost.
Get started
No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.
Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.
Change the base URL. Everything else in your code stays the same.
https://api.zurelay.com/v1Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.
qwen3.8-omni-flashWorks with the tools you already use
Compare
$0.07 / $0.22 per 1M tokens
Pick Qwen 3.8 Flash when your inputs are text, images and video only; it is the open-weight Flash-Next model built for speed.
Qwen 3.8 Flash API$0.06 / $0.12 per 1M tokens
Pick MiMo V2.6 Flash to compare Xiaomi's low-cost open-weight model, which also reads audio and video, on your own recordings.
MiMo V2.6 Flash API$1.20 / $3.60 per 1M tokens
Pick Qwen 3.8 Max for harder reasoning over text, images and video, where Alibaba's flagship is worth its higher price.
Qwen 3.8 Max APIFAQ
Anything else, write to sales@zurelay.com and an engineer will answer.
On zurelay, qwen3.8-omni-flash costs $0.075 per 1M input tokens and $0.235 per 1M output tokens, with cached input at $0.007. Alibaba's international list price is $0.15 input and $0.47 output, so you save 50%. Audio tokens bill at the same $0.075 per 1M as text.
Yes. zurelay serves qwen3.8-omni-flash at 50% below Alibaba's international list price. It is the same model, not a smaller substitute. You pay as you go from prepaid credit, with no subscription.
Set the OpenAI SDK's base URL to https://api.zurelay.com/v1, use your zurelay API key and pass model "qwen3.8-omni-flash". Put audio, images and video in the message content next to your text. zurelay also accepts qwen/qwen3.8-omni-flash and qwen-3-8-omni-flash.
Audio counts 7 tokens per second. A minute of audio is 420 input tokens, and an hour is about 25,200, so even long recordings use a small part of the 1M-token context window.
No. It returns text only. Alibaba's spoken output comes from a separate realtime model, which zurelay does not offer. Pair Omni Flash with your own text-to-speech step if your app needs a voice reply.
Qwen 3.8 Omni Flash has a 1M-token context window and returns up to 131,072 output tokens per response. Alibaba caps a single input at 991,808 tokens in non-thinking mode and 983,616 tokens in thinking mode.
Audio and video understanding and content analysis: summarizing calls and meetings, reviewing video, and tagging large batches of mixed media. Audio works in 113 languages and dialects.
Streaming works with stream: true, as with the OpenAI API. You can set a requests-per-minute cap and a monthly budget on each zurelay API key. Failed attempts are retried automatically and never billed.
Yes. Requests go to Qwen 3.8 Omni Flash, never a smaller or substitute model. Responses are sampled, so wording varies between runs, as it does on Alibaba's own API. zurelay is independent and not affiliated with Alibaba Cloud.
Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.