Agentic software engineering
MiniMax reports 56.22% on SWE-Pro and 57.0% on Terminal Bench 2, plus 55.6% on VIBE-Pro for end-to-end project delivery.
The MiniMax M2.7 API for coding and agents, 57% below MiniMax's list price
minimax-m2.7from openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key="YOUR_ZURELAY_KEY",)stream = client.chat.completions.create( model="minimax-m2.7", messages=[{"role": "user", "content": "Explain quicksort in two sentences."}], stream=True,)for chunk in stream: if chunk.choices: print(chunk.choices[0].delta.content or "", end="")Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.
Pricing
Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.
| Rate | zurelay | MiniMax list | You save |
|---|---|---|---|
Input per 1M tokens | $0.13 | 57%off | |
Output per 1M tokens | $0.51 | 57%off | |
Cached input per 1M tokens | $0.013 | — | — |
One rate at every prompt length, up to the full 205K context window. Streaming, tool calls and structured outputs cost nothing extra.
Move the sliders to your monthly usage.
$184.80 a year
Overview
MiniMax M2.7 is MiniMax's model for agentic coding and professional work, released on March 18, 2026, with about 230B total and 10B active parameters. The zurelay MiniMax M2.7 API serves it through one OpenAI-compatible endpoint with a 204,800-token context window, at 57% below MiniMax's list price.
MiniMax released M2.7 on March 18, 2026 as the MiniMax-M2.7 and M2.7-highspeed API models, and published the weights on Hugging Face in April 2026. It is a Mixture-of-Experts model with about 230B total parameters, of which about 10B are active per token. MiniMax calls it its first model to take part in its own evolution: during development it helped debug training runs and modify its own agent scaffold. On zurelay the MiniMax M2.7 API model ID is minimax-m2.7, and it costs $0.13 per 1M input tokens and $0.51 per 1M output tokens.
MiniMax's reported results center on software engineering and office work. M2.7 scores 56.22% on SWE-Pro, 55.6% on VIBE-Pro for end-to-end project delivery, 57.0% on Terminal Bench 2 and 76.5 on SWE Multilingual. On GDPval-AA it reached an Elo of 1495, which MiniMax called the highest among open-source models at launch. MiniMax also reports 97% skill adherence across more than 40 complex skills, and native support for Agent Teams, where several agents collaborate on one task.
M2.7 is a text-in, text-out reasoning model. It thinks before it answers and interleaves reasoning with tool calls, carrying its reasoning forward across steps of an agent loop. Its context window is 204,800 tokens. MiniMax recommends streaming for its reasoning models and suggests temperature 1.0 and top_p 0.95.
The weights come under MiniMax's non-commercial license, so commercial use of the weights needs written authorization from MiniMax. MiniMax released M3 on June 1, 2026. M2.7 remains on MiniMax's API, and MiniMax has not announced a retirement date for it.
Strengths
MiniMax reports 56.22% on SWE-Pro and 57.0% on Terminal Bench 2, plus 55.6% on VIBE-Pro for end-to-end project delivery.
Only about 10B of its roughly 230B parameters run per token. MiniMax designed the M2 series end to end for agent deployment.
MiniMax reports 97% adherence across more than 40 complex skills, each longer than 2,000 tokens.
The weights are public on Hugging Face under MiniMax's non-commercial license, so you can inspect and test the model yourself.
Use cases
Bug fixing, multi-file changes and terminal tasks inside agent harnesses.
Read logs to find bugs and review code for security issues, work MiniMax highlights for M2.7.
Produce reports, analyses and other professional documents, the kind of work GDPval-AA measures.
Run several agents that split and coordinate one task, using M2.7's native Agent Teams support.
Get started
No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.
Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.
Change the base URL. Everything else in your code stays the same.
https://api.zurelay.com/v1Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.
minimax-m2.7Works with the tools you already use
Compare
$0.32 / $1.59 per 1M tokens
Pick Gemini 3.7 Flash when you need image, video, audio or PDF input or a context window over 1M tokens.
Gemini 3.7 Flash API$0.12 / $1.04 per 1M tokens
Pick Gemini 3.5 Flash-Lite for fast, high-volume steps with multimodal input and a 1M-token context.
Gemini 3.5 Flash-Lite API$0.037 / $0.15 per 1M tokens
Pick DeepSeek V4.1 Flash for a 1M-token context, image input and responses up to 384K tokens.
DeepSeek V4.1 Flash APIFAQ
Anything else, write to sales@zurelay.com and an engineer will answer.
On zurelay, minimax-m2.7 costs $0.13 per 1M input tokens and $0.51 per 1M output tokens, with cached input at $0.013. MiniMax's list price is $0.30 input and $1.20 output per 1M, so you save 57%. Reasoning tokens count as output.
Yes. zurelay serves minimax-m2.7 at 57% below MiniMax's list price for the standard MiniMax-M2.7 model. You pay from prepaid credit with no subscription, and failed requests are never billed.
Install the official openai package, set base_url to https://api.zurelay.com/v1 and use your zurelay API key. Then send a Chat Completions request with model set to "minimax-m2.7". zurelay also accepts MiniMax's own name, MiniMax-M2.7, so existing code only needs a new base URL and key.
MiniMax lists a 204,800-token context window for MiniMax-M2.7. The model takes text input and writes text; it does not read images.
The weights are public on Hugging Face, but under MiniMax's non-commercial license rather than an open-source license. Commercial use of the weights requires written authorization from MiniMax. Personal use and non-profit research are allowed.
Agentic coding and professional work: fixing bugs across a repository, terminal tasks, log analysis, and multi-step office deliverables. It suits long agent loops, because it keeps reasoning between tool calls.
MiniMax sells M2.7 in two API versions that it says give identical results: MiniMax-M2.7 and a faster, higher-priced M2.7-highspeed. zurelay's minimax-m2.7 is compared against the standard MiniMax-M2.7 list price.
Yes. Set stream: true, and add stream_options.include_usage to get token counts in the final chunk. Each zurelay API key can have its own monthly budget and requests-per-minute cap. Failed requests are retried on another route and never billed.
Yes, it is MiniMax's M2.7 model. Responses are sampled, so wording varies between runs, and zurelay does not promise byte-identical output. zurelay is independent and is not affiliated with or endorsed by MiniMax.
Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.