Flagship scale
Alibaba's top Qwen 3.8 model: a mixture-of-experts with 2.4T total parameters and 95B active per token.
The Qwen 3.8 Max API: Alibaba's flagship Qwen model, 40% below list price
qwen3.8-maxfrom openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key="YOUR_ZURELAY_KEY",)stream = client.chat.completions.create( model="qwen3.8-max", messages=[{"role": "user", "content": "Explain quicksort in two sentences."}], stream=True,)for chunk in stream: if chunk.choices: print(chunk.choices[0].delta.content or "", end="")Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.
Pricing
Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.
| Rate | zurelay | Alibaba list | You save |
|---|---|---|---|
Input per 1M tokens | $1.20 | 40%off | |
Output per 1M tokens | $3.60 | 40%off | |
Cached input per 1M tokens | $0.15 | — | — |
One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.
Move the sliders to your monthly usage.
$768.00 a year
Overview
Qwen 3.8 Max is Alibaba's flagship Qwen 3.8 model, a 2.4-trillion-parameter mixture-of-experts with a 1M-token context window and native vision. The zurelay Qwen 3.8 Max API serves it through one OpenAI-compatible endpoint at 40% below Alibaba's international list price, for coding, agent and long-document work.
Alibaba released Qwen 3.8 Max on August 3, 2026 as the flagship of the Qwen 3.8 family. It is a mixture-of-experts model with 2.4T total parameters and 95B active per token. The Qwen 3.8 Max API reads text, images and video, writes text and holds a 1M-token context window, with up to 131,072 output tokens per response. On zurelay the model ID is qwen3.8-max, and it costs $1.20 per 1M input tokens and $3.60 per 1M output tokens.
Alibaba describes it as a major step up in coding and office productivity. It says the model can code on its own for more than ten days to deliver a complete project, and that it handles hundreds of professional tasks in fields such as law, finance and design. Visual understanding is built into the model, so screenshots, charts and video go into the same request as your text.
Qwen 3.8 Max runs in thinking or non-thinking mode, and Alibaba charges the same for both. Thinking mode suits multi-step coding and analysis, while non-thinking mode answers faster for chat and simple extraction. Alibaba lists function calling, structured outputs and context caching, and its rate is the same at every prompt length up to the full 1M tokens.
On zurelay, qwen3.8-max answers as the qwen3.8-max-0902 snapshot, Alibaba's September 2, 2026 version of the model. If your code already sends qwen/qwen3.8-max, qwen-3-8-max or qwen3.8-max-0902, zurelay accepts those names too. You pay $1.20 and $3.60 per 1M tokens, against Alibaba's international list price of $2.00 and $6.00.
Strengths
Alibaba's top Qwen 3.8 model: a mixture-of-experts with 2.4T total parameters and 95B active per token.
Alibaba reports a major leap in coding and says the model can work on its own for more than ten days to deliver a complete project.
Reads images and video natively, so charts, screenshots and recordings sit in the same prompt as your text.
Thinking and non-thinking modes cost the same, so you choose per request by latency and depth, not by budget.
Use cases
Run long refactors, feature builds and debug loops where the model plans, calls tools and checks its own work.
Read contracts, filings and reports in law, finance and other fields in full, inside the 1M-token window.
Ask questions over charts, screenshots, slides and video, and get answers back as text or structured JSON.
Run multi-step workflows that pull data from documents, call your tools and return structured outputs.
Get started
No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.
Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.
Change the base URL. Everything else in your code stays the same.
https://api.zurelay.com/v1Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.
qwen3.8-maxWorks with the tools you already use
Compare
$0.07 / $0.22 per 1M tokens
Pick Qwen 3.8 Flash for the same 1M-token context and image and video input at a much lower list price, when speed and volume matter more than peak quality.
Qwen 3.8 Flash API$0.40 / $1.20 per 1M tokens
Pick HY4 Preview, Tencent's new agent and coding model with a 1M-token context, to compare against Qwen 3.8 Max on your own evals.
HY4 Preview API$1.27 / $6.37 per 1M tokens
Pick Kimi K3 for Moonshot AI's long-horizon coding model, with image input and open weights.
Kimi K3 APIFAQ
Anything else, write to sales@zurelay.com and an engineer will answer.
On zurelay, qwen3.8-max costs $1.20 per 1M input tokens and $3.60 per 1M output tokens, with cached input at $0.15. Alibaba's international list price is $2.00 input and $6.00 output, so you save 40%. Thinking and non-thinking modes cost the same, and the rate holds at every prompt length.
Yes. zurelay serves qwen3.8-max at 40% below Alibaba's international list price. It is the same model, answering as the qwen3.8-max-0902 snapshot, not a smaller substitute. You pay from prepaid credit, with no subscription.
Set the OpenAI SDK's base URL to https://api.zurelay.com/v1, use your zurelay API key and pass model "qwen3.8-max". zurelay also accepts qwen/qwen3.8-max, qwen-3-8-max and qwen3.8-max-0902. Messages, tool calls and streaming use the standard OpenAI request format.
Qwen 3.8 Max has a 1M-token context window and returns up to 131,072 output tokens per response. Alibaba caps a single input at 991,808 tokens in non-thinking mode and 983,616 tokens in thinking mode, which leaves room for the reply.
Yes. It runs in thinking or non-thinking mode, and Alibaba prices both the same. Use thinking for hard coding and multi-step analysis, and non-thinking for chat and quick extraction, where a fast reply matters more.
Yes. It takes images and video alongside text in the same request and returns text. It does not generate images or video; for images, use an image model such as Nano Banana 2 on the same key.
Alibaba built it for coding and office productivity: long autonomous coding runs, professional tasks in fields such as law, finance and design, and work that mixes text with images or video. For high-volume, simpler jobs, Qwen 3.8 Flash is often enough at a fraction of the list price.
Yes. Set stream: true to receive tokens as they are generated, and add stream_options.include_usage to get token counts in the final chunk. Each zurelay API key can have its own monthly budget and requests-per-minute cap. Failed attempts are retried automatically and never billed.
Yes. Requests run on Qwen 3.8 Max itself, as the qwen3.8-max-0902 snapshot, never a smaller or substitute model. Output is sampled, so wording varies between runs on any API. zurelay is an independent service and is not affiliated with or endorsed by Alibaba Cloud.
Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.