6B active parameters
Only 6B of its 125B parameters run per token, which keeps replies fast and the cost per call low.
The Qwen 3.8 Flash API: fast, low-cost Qwen for high-volume work, 53% below list
qwen3.8-flashfrom openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key="YOUR_ZURELAY_KEY",)stream = client.chat.completions.create( model="qwen3.8-flash", messages=[{"role": "user", "content": "Explain quicksort in two sentences."}], stream=True,)for chunk in stream: if chunk.choices: print(chunk.choices[0].delta.content or "", end="")Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.
Pricing
Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.
| Rate | zurelay | Alibaba list | You save |
|---|---|---|---|
Input per 1M tokens | $0.07 | 53%off | |
Output per 1M tokens | $0.22 | 53%off | |
Cached input per 1M tokens | $0.007 | — | — |
One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.
Move the sliders to your monthly usage.
$78.00 a year
Overview
Qwen 3.8 Flash is Alibaba's fast, low-cost Qwen 3.8 model, a 125B-parameter mixture-of-experts with only 6B active per token and a 1M-token context window. The zurelay Qwen 3.8 Flash API serves it through one OpenAI-compatible endpoint at 53% below Alibaba's international list price, for coding helpers, agents and high-volume pipelines.
Alibaba released Qwen 3.8 Flash on August 26, 2026 as the fast, low-cost model of the Qwen 3.8 family. It is the open-weight Flash-Next model: a mixture-of-experts with 125B total parameters and just 6B active per token, which previews the Qwen4 architecture. The small active size is what keeps it quick and cheap to run.
The Qwen 3.8 Flash API reads text, images and video and writes text, with a 1M-token context window and up to 131,072 output tokens per response. Alibaba points to coding help, agent workflows and visual understanding, with examples such as fixing code on its own, operating desktop applications and analyzing charts and long videos.
It supports a thinking mode, function calling, structured outputs and context caching. Alibaba charges one rate at every prompt length, up to the full 1M tokens.
On zurelay the model ID is qwen3.8-flash, and qwen/qwen3.8-flash and qwen-3-8-flash work as aliases. You pay $0.07 per 1M input tokens and $0.22 per 1M output tokens, against Alibaba's international list price of $0.15 and $0.47.
Strengths
Only 6B of its 125B parameters run per token, which keeps replies fast and the cost per call low.
Long documents, codebases and agent traces fit in one prompt, with room for 131,072 output tokens.
Reads screenshots, charts and long videos in the same request as your text, so one low-cost model covers text and visual steps.
Flash-Next previews the Qwen4 architecture, and its weights are open, so you can self-host the same model later.
Use cases
Run code review, test generation and small autonomous fixes, where latency and cost per call matter.
Use Flash for the frequent tool-calling steps of an agent system and keep Qwen 3.8 Max for planning or final review.
Read screenshots and decide the next click or keystroke in computer-use agents.
Pull numbers from charts and summarize long videos at volume.
Get started
No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.
Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.
Change the base URL. Everything else in your code stays the same.
https://api.zurelay.com/v1Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.
qwen3.8-flashWorks with the tools you already use
Compare
$1.20 / $3.60 per 1M tokens
Pick Qwen 3.8 Max, Alibaba's 2.4T-parameter flagship, for the hardest coding and analysis work, where quality matters more than cost.
Qwen 3.8 Max API$0.075 / $0.235 per 1M tokens
Pick Qwen 3.8 Omni Flash when your inputs include audio as well as images and video; Alibaba lists it at the same price for text, image and video input.
Qwen 3.8 Omni Flash API$0.06 / $0.12 per 1M tokens
Pick MiMo V2.6 Flash to compare Xiaomi's low-cost open-weight model, which also takes audio, on your own evals.
MiMo V2.6 Flash APIFAQ
Anything else, write to sales@zurelay.com and an engineer will answer.
On zurelay, qwen3.8-flash costs $0.07 per 1M input tokens and $0.22 per 1M output tokens, with cached input at $0.007. Alibaba's international list price is $0.15 input and $0.47 output, so you save 53%. The rate is the same at every prompt length.
Yes. zurelay serves qwen3.8-flash at 53% below Alibaba's international list price. It is the same model, not a smaller substitute. You pay as you go for the tokens you use, with no subscription.
Set the OpenAI SDK's base URL to https://api.zurelay.com/v1, use your zurelay API key and pass model "qwen3.8-flash". zurelay also accepts qwen/qwen3.8-flash and qwen-3-8-flash. Tool calls and streaming work the same way they do with OpenAI models.
Qwen 3.8 Flash has a 1M-token context window and returns up to 131,072 output tokens per response. Alibaba caps a single input at 991,808 tokens in non-thinking mode and 983,616 tokens in thinking mode.
Both have a 1M-token context, up to 131,072 output tokens and image and video input. Max is the 2.4T-parameter flagship with 95B active per token, built for the hardest coding and professional work. Flash runs 6B active parameters at a much lower list price and is the better default for volume, sub-agents and latency-sensitive routes.
Its weights are open: Qwen 3.8 Flash is the Flash-Next model, a 125B-parameter mixture-of-experts with 6B active per token. You can self-host it, and the API is the quick way to use it without running your own GPUs.
Fast, high-volume work: coding help, agent steps that run many times, and visual tasks such as reading charts, screenshots and long videos. Alibaba also points to fixing code on its own and operating desktop applications.
Streaming works with stream: true, as with the OpenAI API. You can set a requests-per-minute cap and a monthly budget on each zurelay API key. Failed attempts are retried automatically and never billed.
Yes. Requests go to Qwen 3.8 Flash, never a smaller or substitute model. Responses are sampled, so wording varies between runs, as it does on Alibaba's own API. zurelay is independent and not affiliated with Alibaba Cloud.
Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.