Alibaba
Operational· 100% uptime, 7 days

Qwen 3.8 Flash API

The Qwen 3.8 Flash API: fast, low-cost Qwen for high-volume work, 53% below list

Price per 1M tokens53%off
Input
$0.07$0.15
Output
$0.22$0.47
Cached input
$0.007
Model IDqwen3.8-flash
chat.py
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key="YOUR_ZURELAY_KEY",
)
stream = client.chat.completions.create(
model="qwen3.8-flash",
messages=[{"role": "user", "content": "Explain quicksort in two sentences."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")

Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.

Context window
1M tokens
Max output
131K tokens
Input
Text, Image, Video
Output
Text
Released
Aug 26, 2026
Uptime, 7 days
100%

Pricing

Qwen 3.8 Flash API pricing. A fraction of the list price.

Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.

RatezurelayAlibaba listYou save
Input
per 1M tokens
$0.07$0.1553%off
Output
per 1M tokens
$0.22$0.4753%off
Cached input
per 1M tokens
$0.007——

One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.

Savings calculator

Move the sliders to your monthly usage.

Input tokens per month50M
Output tokens per month10M
zurelay
$5.70
Alibaba list price
$12.20
You save every month$6.50

$78.00 a year

Overview

What is Qwen 3.8 Flash?

Qwen 3.8 Flash is Alibaba's fast, low-cost Qwen 3.8 model, a 125B-parameter mixture-of-experts with only 6B active per token and a 1M-token context window. The zurelay Qwen 3.8 Flash API serves it through one OpenAI-compatible endpoint at 53% below Alibaba's international list price, for coding helpers, agents and high-volume pipelines.

Alibaba released Qwen 3.8 Flash on August 26, 2026 as the fast, low-cost model of the Qwen 3.8 family. It is the open-weight Flash-Next model: a mixture-of-experts with 125B total parameters and just 6B active per token, which previews the Qwen4 architecture. The small active size is what keeps it quick and cheap to run.

The Qwen 3.8 Flash API reads text, images and video and writes text, with a 1M-token context window and up to 131,072 output tokens per response. Alibaba points to coding help, agent workflows and visual understanding, with examples such as fixing code on its own, operating desktop applications and analyzing charts and long videos.

It supports a thinking mode, function calling, structured outputs and context caching. Alibaba charges one rate at every prompt length, up to the full 1M tokens.

On zurelay the model ID is qwen3.8-flash, and qwen/qwen3.8-flash and qwen-3-8-flash work as aliases. You pay $0.07 per 1M input tokens and $0.22 per 1M output tokens, against Alibaba's international list price of $0.15 and $0.47.

Strengths

Where Qwen 3.8 Flash shines. And what teams build with it.

01

6B active parameters

Only 6B of its 125B parameters run per token, which keeps replies fast and the cost per call low.

02

1M-token context

Long documents, codebases and agent traces fit in one prompt, with room for 131,072 output tokens.

03

Images and video too

Reads screenshots, charts and long videos in the same request as your text, so one low-cost model covers text and visual steps.

04

A preview of Qwen4

Flash-Next previews the Qwen4 architecture, and its weights are open, so you can self-host the same model later.

Use cases

  • Coding helpers

    Run code review, test generation and small autonomous fixes, where latency and cost per call matter.

  • Sub-agents

    Use Flash for the frequent tool-calling steps of an agent system and keep Qwen 3.8 Max for planning or final review.

  • Desktop and UI automation

    Read screenshots and decide the next click or keystroke in computer-use agents.

  • Chart and video analysis

    Pull numbers from charts and summarize long videos at volume.

Get started

Call Qwen 3.8 Flash in three steps.

No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.

  1. 1

    Create a key

    Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.

  2. 2

    Point your SDK at zurelay

    Change the base URL. Everything else in your code stays the same.

    https://api.zurelay.com/v1
  3. 3

    Use qwen3.8-flash

    Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.

    qwen3.8-flash

FAQ

Qwen 3.8 Flash API questions.

Anything else, write to sales@zurelay.com and an engineer will answer.

How much does the Qwen 3.8 Flash API cost on zurelay?

On zurelay, qwen3.8-flash costs $0.07 per 1M input tokens and $0.22 per 1M output tokens, with cached input at $0.007. Alibaba's international list price is $0.15 input and $0.47 output, so you save 53%. The rate is the same at every prompt length.

Is there a cheaper Qwen 3.8 Flash API than Alibaba's?

Yes. zurelay serves qwen3.8-flash at 53% below Alibaba's international list price. It is the same model, not a smaller substitute. You pay as you go for the tokens you use, with no subscription.

How do I call the Qwen 3.8 Flash API with the OpenAI SDK?

Set the OpenAI SDK's base URL to https://api.zurelay.com/v1, use your zurelay API key and pass model "qwen3.8-flash". zurelay also accepts qwen/qwen3.8-flash and qwen-3-8-flash. Tool calls and streaming work the same way they do with OpenAI models.

What is the Qwen 3.8 Flash context window?

Qwen 3.8 Flash has a 1M-token context window and returns up to 131,072 output tokens per response. Alibaba caps a single input at 991,808 tokens in non-thinking mode and 983,616 tokens in thinking mode.

Should I use Qwen 3.8 Flash or Qwen 3.8 Max?

Both have a 1M-token context, up to 131,072 output tokens and image and video input. Max is the 2.4T-parameter flagship with 95B active per token, built for the hardest coding and professional work. Flash runs 6B active parameters at a much lower list price and is the better default for volume, sub-agents and latency-sensitive routes.

Is Qwen 3.8 Flash open source?

Its weights are open: Qwen 3.8 Flash is the Flash-Next model, a 125B-parameter mixture-of-experts with 6B active per token. You can self-host it, and the API is the quick way to use it without running your own GPUs.

What is Qwen 3.8 Flash best at?

Fast, high-volume work: coding help, agent steps that run many times, and visual tasks such as reading charts, screenshots and long videos. Alibaba also points to fixing code on its own and operating desktop applications.

Does the Qwen 3.8 Flash API support streaming and per-key rate limits?

Streaming works with stream: true, as with the OpenAI API. You can set a requests-per-minute cap and a monthly budget on each zurelay API key. Failed attempts are retried automatically and never billed.

Is it the same Qwen 3.8 Flash model Alibaba serves?

Yes. Requests go to Qwen 3.8 Flash, never a smaller or substitute model. Responses are sampled, so wording varies between runs, as it does on Alibaba's own API. zurelay is independent and not affiliated with Alibaba Cloud.

More models on zurelay

All models

Stop paying list price.
Start saving today.

Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.