Xiaomi
Disrupted· 89.6% uptime, 7 days

MiMo V2.6 Flash API

The MiMo V2.6 Flash API: Xiaomi's low-cost omnimodal model for high-volume work

Price per 1M tokens57%off
Input
$0.06$0.14
Output
$0.12$0.28
Cached input
$0.0012
Model IDmimo-v2.6-flash
chat.py
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key="YOUR_ZURELAY_KEY",
)
stream = client.chat.completions.create(
model="mimo-v2.6-flash",
messages=[{"role": "user", "content": "Explain quicksort in two sentences."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")

Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.

Context window
1M tokens
Max output
128K tokens
Input
Text, Image, Video, Audio
Output
Text
Released
Sep 22, 2026
Uptime, 7 days
89.6%

Pricing

MiMo V2.6 Flash API pricing. A fraction of the list price.

Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.

RatezurelayXiaomi listYou save
Input
per 1M tokens
$0.06$0.1457%off
Output
per 1M tokens
$0.12$0.2857%off
Cached input
per 1M tokens
$0.0012——

One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.

Savings calculator

Move the sliders to your monthly usage.

Input tokens per month50M
Output tokens per month10M
zurelay
$4.20
Xiaomi list price
$9.80
You save every month$5.60

$67.20 a year

Overview

What is MiMo V2.6 Flash?

MiMo V2.6 Flash is the efficient model in Xiaomi's MiMo-V2.6 series: 15B active parameters, 1M tokens of context, and text, image, audio and video input. The MiMo V2.6 Flash API on zurelay is for teams running high-volume or latency-sensitive workloads who want near-Pro agent results at 57% below Xiaomi's list price.

MiMo V2.6 Flash is the efficient model in Xiaomi's MiMo-V2.6 series, released alongside Pro on September 22, 2026. The MiMo V2.6 Flash API takes text, images, video and audio, returns text, and supports a 1M-token context window with up to 128K output tokens. Xiaomi positions it as the best balance for high-frequency calls and large-scale tasks.

It is a sparse mixture-of-experts model with 309B total parameters and 15B active per token, spread across 48 layers that mix sliding-window and global attention. The small active size keeps it quick and cheap to serve. The weights are open under the MIT license.

Flash was trained the same way as Pro: one mixed reinforcement learning run across coding, general agents, visual tasks and cybersecurity. It stays close to Pro on agent work. Xiaomi reports 87.6 on Terminal Bench 2.1 against Pro's 89.9, 80.8 on OSWorld-Verified against 82.0, and 95.1 on CyberGym, slightly above Pro.

Its list price is about a third of Pro's. On zurelay you call it with any OpenAI SDK using the model ID mimo-v2.6-flash, and pay $0.06 per 1M input tokens and $0.12 per 1M output tokens, against Xiaomi's list price of $0.14 and $0.28.

Strengths

Where MiMo V2.6 Flash shines. And what teams build with it.

01

Near-Pro agent scores

Xiaomi's own results put Flash within a few points of Pro on Terminal Bench 2.1, DeepSWE v1.1 and OSWorld-Verified.

02

Low cost per call

15B active parameters and a list price about a third of Pro's make it practical for high-volume pipelines.

03

Omnimodal input

Text, images, video and audio fit in one request, the same input range as Pro.

04

1M-token context

Long documents, repositories and agent traces fit in one prompt, with room for 128K output tokens.

Use cases

  • High-volume extraction and tagging

    Tag, route or pull fields from large streams of text, images or audio at a low cost per call.

  • Sub-agents

    Use Flash for the frequent tool-calling steps in an agent system and keep Pro for planning or final review.

  • Media processing

    Ask questions about audio clips, video and screenshots directly, without a separate transcription or vision step.

  • Coding helpers

    Run code review, test generation and terminal helpers where latency and cost per call matter.

Get started

Call MiMo V2.6 Flash in three steps.

No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.

  1. 1

    Create a key

    Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.

  2. 2

    Point your SDK at zurelay

    Change the base URL. Everything else in your code stays the same.

    https://api.zurelay.com/v1
  3. 3

    Use mimo-v2.6-flash

    Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.

    mimo-v2.6-flash

FAQ

MiMo V2.6 Flash API questions.

Anything else, write to sales@zurelay.com and an engineer will answer.

How much does the MiMo V2.6 Flash API cost?

On zurelay, MiMo V2.6 Flash costs $0.06 per 1M input tokens and $0.12 per 1M output tokens, with cached input at $0.0012. Xiaomi's list price is $0.14 input and $0.28 output, so you save 57%.

Is there a cheaper Xiaomi MiMo API?

zurelay serves mimo-v2.6-flash at 57% below Xiaomi's list price. You pay as you go for the tokens you use, and credit never expires.

How do I call the MiMo V2.6 Flash API with the OpenAI SDK?

Set the OpenAI SDK's base URL to https://api.zurelay.com/v1, use your zurelay API key and pass model "mimo-v2.6-flash". Tool calls and streaming work the same way they do with OpenAI models.

What is the MiMo V2.6 Flash context window?

MiMo V2.6 Flash has a 1M-token context window and returns up to 128K output tokens per response, the same limits as Pro.

What is MiMo V2.6 Flash best at?

High-frequency and large-scale work: agents that make many calls, extraction and tagging jobs, and media analysis. Xiaomi trained it on coding, general agents, visual tasks and cybersecurity, and it scores close to Pro on Xiaomi's agent benchmarks.

Should I use MiMo V2.6 Flash or MiMo V2.6 Pro?

Both have the same context window, input types and features. Pro has 1.02T total and 42B active parameters and scores a little higher on most of Xiaomi's agent benchmarks. Flash has 309B total and 15B active, costs about a third as much at list price and is the better default for volume.

Is MiMo V2.6 Flash open source?

Yes. Xiaomi released the weights on Hugging Face under the MIT license. You can use the API today and move to self-hosting later on the same model.

Does the MiMo V2.6 Flash API support streaming and per-key rate limits?

Streaming works with stream: true, as with the OpenAI API. You can set a requests-per-minute cap and a monthly budget on each zurelay API key.

Is it the same MiMo V2.6 Flash model Xiaomi serves?

Yes. Requests go to MiMo V2.6 Flash, never a substitute model. Responses are sampled, so wording varies between runs, as it does on Xiaomi's own platform. zurelay is independent and not affiliated with Xiaomi.

More models on zurelay

All models

Stop paying list price.
Start saving today.

Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.