Alibaba
Operational· 100% uptime, 7 days

Qwen 3.8 Omni Flash API

The Qwen 3.8 Omni Flash API: audio, video and images in, text out, 50% below list

Price per 1M tokens50%off
Input
$0.075$0.15
Output
$0.235$0.47
Cached input
$0.007
Model IDqwen3.8-omni-flash
chat.py
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key="YOUR_ZURELAY_KEY",
)
stream = client.chat.completions.create(
model="qwen3.8-omni-flash",
messages=[{"role": "user", "content": "Explain quicksort in two sentences."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")

Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.

Context window
1M tokens
Max output
131K tokens
Input
Text, Image, Audio, Video
Output
Text
Released
Sep 18, 2026
Uptime, 7 days
100%

Pricing

Qwen 3.8 Omni Flash API pricing. A fraction of the list price.

Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.

RatezurelayAlibaba listYou save
Input
per 1M tokens
$0.075$0.1550%off
Output
per 1M tokens
$0.235$0.4750%off
Cached input
per 1M tokens
$0.007——

One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.

Savings calculator

Move the sliders to your monthly usage.

Input tokens per month50M
Output tokens per month10M
zurelay
$6.10
Alibaba list price
$12.20
You save every month$6.10

$73.20 a year

Overview

What is Qwen 3.8 Omni Flash?

Qwen 3.8 Omni Flash is Alibaba's multimodal Qwen 3.8 model for audio and video understanding: it takes text, images, audio and video in one request and answers in text. The zurelay Qwen 3.8 Omni Flash API serves it through one OpenAI-compatible endpoint with a 1M-token context window, at 50% below Alibaba's international list price.

Alibaba released Qwen 3.8 Omni Flash on September 18, 2026. It is built for audio and video understanding and content analysis: you send text, images, audio and video, and it returns text. The Qwen 3.8 Omni Flash API holds a 1M-token context window and writes up to 131,072 output tokens per response. On zurelay the model ID is qwen3.8-omni-flash, and it costs $0.075 per 1M input tokens and $0.235 per 1M output tokens.

Audio understanding covers 113 languages and dialects. Audio counts 7 tokens per second, so a minute of speech is 420 input tokens and an hour is about 25,200, and long recordings fit easily in the context window. zurelay bills every input token at $0.075 per 1M, audio included. Alibaba also lists tool calling and a thinking mode that is on by default, with adjustable reasoning effort.

The model writes text only. Alibaba offers spoken replies through a separate realtime model, which zurelay does not serve, so plan on text output, or add your own text-to-speech step.

qwen/qwen3.8-omni-flash and qwen-3-8-omni-flash work as aliases. You pay $0.075 and $0.235 per 1M tokens, against Alibaba's international list price of $0.15 and $0.47 for text, image and video input and text output.

Strengths

Where Qwen 3.8 Omni Flash shines. And what teams build with it.

01

Every input in one model

Text, images, audio and video go into the same request, with no separate transcription or vision service to wire up.

02

Long recordings fit

At 7 tokens per second of audio, an hour of speech is about 25,200 tokens, a small slice of the 1M-token window.

03

Many languages

Alibaba lists audio understanding in 113 languages and dialects.

04

Flash pricing

Alibaba lists Omni Flash at Qwen 3.8 Flash's price for text, image and video input, and on zurelay audio tokens bill at the same $0.075 per 1M as text.

Use cases

  • Meeting and call analysis

    Summarize calls, pull out action items and tag topics straight from the audio, without a transcription step first.

  • Video review

    Ask questions about video content, flag scenes or write descriptions from the picture and the soundtrack together.

  • Multilingual audio

    Understand speech across many languages and answer in the language your app needs.

  • Media tagging at volume

    Classify and route large batches of audio clips, videos and images at Flash-tier cost.

Get started

Call Qwen 3.8 Omni Flash in three steps.

No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.

  1. 1

    Create a key

    Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.

  2. 2

    Point your SDK at zurelay

    Change the base URL. Everything else in your code stays the same.

    https://api.zurelay.com/v1
  3. 3

    Use qwen3.8-omni-flash

    Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.

    qwen3.8-omni-flash

FAQ

Qwen 3.8 Omni Flash API questions.

Anything else, write to sales@zurelay.com and an engineer will answer.

How much does the Qwen 3.8 Omni Flash API cost on zurelay?

On zurelay, qwen3.8-omni-flash costs $0.075 per 1M input tokens and $0.235 per 1M output tokens, with cached input at $0.007. Alibaba's international list price is $0.15 input and $0.47 output, so you save 50%. Audio tokens bill at the same $0.075 per 1M as text.

Is there a cheaper Qwen 3.8 Omni Flash API than Alibaba's?

Yes. zurelay serves qwen3.8-omni-flash at 50% below Alibaba's international list price. It is the same model, not a smaller substitute. You pay as you go from prepaid credit, with no subscription.

How do I call the Qwen 3.8 Omni Flash API with the OpenAI SDK?

Set the OpenAI SDK's base URL to https://api.zurelay.com/v1, use your zurelay API key and pass model "qwen3.8-omni-flash". Put audio, images and video in the message content next to your text. zurelay also accepts qwen/qwen3.8-omni-flash and qwen-3-8-omni-flash.

How many tokens does audio use on Qwen 3.8 Omni Flash?

Audio counts 7 tokens per second. A minute of audio is 420 input tokens, and an hour is about 25,200, so even long recordings use a small part of the 1M-token context window.

Can Qwen 3.8 Omni Flash reply with speech?

No. It returns text only. Alibaba's spoken output comes from a separate realtime model, which zurelay does not offer. Pair Omni Flash with your own text-to-speech step if your app needs a voice reply.

What is the Qwen 3.8 Omni Flash context window?

Qwen 3.8 Omni Flash has a 1M-token context window and returns up to 131,072 output tokens per response. Alibaba caps a single input at 991,808 tokens in non-thinking mode and 983,616 tokens in thinking mode.

What is Qwen 3.8 Omni Flash best at?

Audio and video understanding and content analysis: summarizing calls and meetings, reviewing video, and tagging large batches of mixed media. Audio works in 113 languages and dialects.

Does the Qwen 3.8 Omni Flash API support streaming and per-key rate limits?

Streaming works with stream: true, as with the OpenAI API. You can set a requests-per-minute cap and a monthly budget on each zurelay API key. Failed attempts are retried automatically and never billed.

Is it the same Qwen 3.8 Omni Flash model Alibaba serves?

Yes. Requests go to Qwen 3.8 Omni Flash, never a smaller or substitute model. Responses are sampled, so wording varies between runs, as it does on Alibaba's own API. zurelay is independent and not affiliated with Alibaba Cloud.

More models on zurelay

All models

Stop paying list price.
Start saving today.

Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.