Alibaba
Operational· 100% uptime, 7 days

Qwen 3.8 Max API

The Qwen 3.8 Max API: Alibaba's flagship Qwen model, 40% below list price

Price per 1M tokens40%off
Input
$1.20$2.00
Output
$3.60$6.00
Cached input
$0.15
Model IDqwen3.8-max
chat.py
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key="YOUR_ZURELAY_KEY",
)
stream = client.chat.completions.create(
model="qwen3.8-max",
messages=[{"role": "user", "content": "Explain quicksort in two sentences."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")

Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.

Context window
1M tokens
Max output
131K tokens
Input
Text, Image, Video
Output
Text
Released
Aug 3, 2026
Uptime, 7 days
100%

Pricing

Qwen 3.8 Max API pricing. A fraction of the list price.

Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.

RatezurelayAlibaba listYou save
Input
per 1M tokens
$1.20$2.0040%off
Output
per 1M tokens
$3.60$6.0040%off
Cached input
per 1M tokens
$0.15——

One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.

Savings calculator

Move the sliders to your monthly usage.

Input tokens per month50M
Output tokens per month10M
zurelay
$96.00
Alibaba list price
$160.00
You save every month$64.00

$768.00 a year

Overview

What is Qwen 3.8 Max?

Qwen 3.8 Max is Alibaba's flagship Qwen 3.8 model, a 2.4-trillion-parameter mixture-of-experts with a 1M-token context window and native vision. The zurelay Qwen 3.8 Max API serves it through one OpenAI-compatible endpoint at 40% below Alibaba's international list price, for coding, agent and long-document work.

Alibaba released Qwen 3.8 Max on August 3, 2026 as the flagship of the Qwen 3.8 family. It is a mixture-of-experts model with 2.4T total parameters and 95B active per token. The Qwen 3.8 Max API reads text, images and video, writes text and holds a 1M-token context window, with up to 131,072 output tokens per response. On zurelay the model ID is qwen3.8-max, and it costs $1.20 per 1M input tokens and $3.60 per 1M output tokens.

Alibaba describes it as a major step up in coding and office productivity. It says the model can code on its own for more than ten days to deliver a complete project, and that it handles hundreds of professional tasks in fields such as law, finance and design. Visual understanding is built into the model, so screenshots, charts and video go into the same request as your text.

Qwen 3.8 Max runs in thinking or non-thinking mode, and Alibaba charges the same for both. Thinking mode suits multi-step coding and analysis, while non-thinking mode answers faster for chat and simple extraction. Alibaba lists function calling, structured outputs and context caching, and its rate is the same at every prompt length up to the full 1M tokens.

On zurelay, qwen3.8-max answers as the qwen3.8-max-0902 snapshot, Alibaba's September 2, 2026 version of the model. If your code already sends qwen/qwen3.8-max, qwen-3-8-max or qwen3.8-max-0902, zurelay accepts those names too. You pay $1.20 and $3.60 per 1M tokens, against Alibaba's international list price of $2.00 and $6.00.

Strengths

Where Qwen 3.8 Max shines. And what teams build with it.

01

Flagship scale

Alibaba's top Qwen 3.8 model: a mixture-of-experts with 2.4T total parameters and 95B active per token.

02

Long autonomous coding

Alibaba reports a major leap in coding and says the model can work on its own for more than ten days to deliver a complete project.

03

Vision built in

Reads images and video natively, so charts, screenshots and recordings sit in the same prompt as your text.

04

One price for both modes

Thinking and non-thinking modes cost the same, so you choose per request by latency and depth, not by budget.

Use cases

  • Coding agents

    Run long refactors, feature builds and debug loops where the model plans, calls tools and checks its own work.

  • Professional documents

    Read contracts, filings and reports in law, finance and other fields in full, inside the 1M-token window.

  • Visual analysis

    Ask questions over charts, screenshots, slides and video, and get answers back as text or structured JSON.

  • Office automation

    Run multi-step workflows that pull data from documents, call your tools and return structured outputs.

Get started

Call Qwen 3.8 Max in three steps.

No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.

  1. 1

    Create a key

    Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.

  2. 2

    Point your SDK at zurelay

    Change the base URL. Everything else in your code stays the same.

    https://api.zurelay.com/v1
  3. 3

    Use qwen3.8-max

    Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.

    qwen3.8-max

FAQ

Qwen 3.8 Max API questions.

Anything else, write to sales@zurelay.com and an engineer will answer.

How much does the Qwen 3.8 Max API cost on zurelay?

On zurelay, qwen3.8-max costs $1.20 per 1M input tokens and $3.60 per 1M output tokens, with cached input at $0.15. Alibaba's international list price is $2.00 input and $6.00 output, so you save 40%. Thinking and non-thinking modes cost the same, and the rate holds at every prompt length.

Is there a cheaper Qwen 3.8 Max API than Alibaba's?

Yes. zurelay serves qwen3.8-max at 40% below Alibaba's international list price. It is the same model, answering as the qwen3.8-max-0902 snapshot, not a smaller substitute. You pay from prepaid credit, with no subscription.

How do I call the Qwen 3.8 Max API with the OpenAI SDK?

Set the OpenAI SDK's base URL to https://api.zurelay.com/v1, use your zurelay API key and pass model "qwen3.8-max". zurelay also accepts qwen/qwen3.8-max, qwen-3-8-max and qwen3.8-max-0902. Messages, tool calls and streaming use the standard OpenAI request format.

What is the Qwen 3.8 Max context window?

Qwen 3.8 Max has a 1M-token context window and returns up to 131,072 output tokens per response. Alibaba caps a single input at 991,808 tokens in non-thinking mode and 983,616 tokens in thinking mode, which leaves room for the reply.

Does Qwen 3.8 Max have a thinking mode?

Yes. It runs in thinking or non-thinking mode, and Alibaba prices both the same. Use thinking for hard coding and multi-step analysis, and non-thinking for chat and quick extraction, where a fast reply matters more.

Can Qwen 3.8 Max read images and video?

Yes. It takes images and video alongside text in the same request and returns text. It does not generate images or video; for images, use an image model such as Nano Banana 2 on the same key.

What is Qwen 3.8 Max best at?

Alibaba built it for coding and office productivity: long autonomous coding runs, professional tasks in fields such as law, finance and design, and work that mixes text with images or video. For high-volume, simpler jobs, Qwen 3.8 Flash is often enough at a fraction of the list price.

Does the Qwen 3.8 Max API support streaming and per-key rate limits?

Yes. Set stream: true to receive tokens as they are generated, and add stream_options.include_usage to get token counts in the final chunk. Each zurelay API key can have its own monthly budget and requests-per-minute cap. Failed attempts are retried automatically and never billed.

Is Qwen 3.8 Max on zurelay the same model Alibaba serves?

Yes. Requests run on Qwen 3.8 Max itself, as the qwen3.8-max-0902 snapshot, never a smaller or substitute model. Output is sampled, so wording varies between runs on any API. zurelay is an independent service and is not affiliated with or endorsed by Alibaba Cloud.

More models on zurelay

All models

Stop paying list price.
Start saving today.

Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.