A usable 1M-token window
Z.ai trained GLM-5.2 for months on long coding-agent tasks so quality holds deep into the context. A whole service with its tests and docs fits in one request.
The GLM-5.2 API with 1M tokens of context for long coding tasks, 73% below list
glm-5.2from openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key="YOUR_ZURELAY_KEY",)stream = client.chat.completions.create( model="glm-5.2", messages=[{"role": "user", "content": "Explain quicksort in two sentences."}], stream=True,)for chunk in stream: if chunk.choices: print(chunk.choices[0].delta.content or "", end="")Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.
Pricing
Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.
| Rate | zurelay | Z.ai list | You save |
|---|---|---|---|
Input per 1M tokens | $0.38 | 73%off | |
Output per 1M tokens | $1.21 | 73%off | |
Cached input per 1M tokens | $0.038 | — | — |
One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.
Move the sliders to your monthly usage.
$994.80 a year
Overview
GLM-5.2 is Z.ai's open-weight model for long-horizon coding and the first GLM with a 1M-token context window. The zurelay GLM-5.2 API serves the same model through one OpenAI-compatible endpoint at 73% below Z.ai's list price.
GLM-5.2 is Z.ai's model for long-horizon tasks, released on June 16, 2026 after a few days of early access for GLM Coding Plan users. The GLM-5.2 API takes text, returns text, holds up to 1M tokens of context and writes up to 128K output tokens per request. That is five times the 200K context of GLM-5.1, and Z.ai trained the model for months on long coding-agent work so the whole window stays usable.
It is a Mixture-of-Experts model with 744B total parameters and 40B active per token, released with open weights under the MIT license. A technique Z.ai calls IndexShare reuses one sparse-attention indexer across every four layers, which cuts per-token compute by 2.9 times at 1M context. Z.ai also improved the multi-token prediction layer used for speculative decoding, raising acceptance length by up to 20%.
On coding benchmarks Z.ai reports 81.0 on Terminal-Bench 2.1 and 62.1 on SWE-bench Pro, up from 62.0 and 58.4 for GLM-5.1. In practice Z.ai points to project-level work: auditing a full codebase, running long refactors to completion, following a team's lint, build and commit rules, and debugging Android apps on a real device with ADB and logcat.
Thinking is on by default, and you can turn it off. With thinking on, Z.ai offers two reasoning_effort levels, high and max, with max as the default and the setting Z.ai recommends for coding. The model supports function calling, structured JSON output, streaming and context caching. On zurelay you call it at https://api.zurelay.com/v1 with the model ID glm-5.2, for $0.38 per 1M input tokens and $1.21 per 1M output tokens against Z.ai's list price of $1.40 and $4.40.
Strengths
Z.ai trained GLM-5.2 for months on long coding-agent tasks so quality holds deep into the context. A whole service with its tests and docs fits in one request.
It breaks a goal into stages, tracks dependencies and verifies as it goes. That suits module decoupling, API migrations and cross-language rewrites.
Z.ai reports more reliable adherence to code style, architecture boundaries and commit conventions over long sessions, with fewer out-of-scope changes.
The weights are open under MIT, so you can build on zurelay today and self-host the same model later if you need to.
Use cases
Load a real repository and ask for an architecture map, API contracts, data flows and technical debt in one pass.
Hand the model a bounded refactor, let it plan and change files across the project, then run the tests to confirm the result.
Build Android clients or WeChat Mini Programs from existing web code, then debug from ADB and logcat output.
Turn a paper's model, loss functions and data pipeline into a runnable project and check the metrics against the paper.
Get started
No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.
Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.
Change the base URL. Everything else in your code stays the same.
https://api.zurelay.com/v1Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.
glm-5.2Works with the tools you already use
Compare
$0.38 / $1.21 per 1M tokens
Pick GLM-5.3 for harder coding and security work; it builds on the same base model, and Z.ai lists both at the same price.
GLM-5.3 API$1.27 / $6.37 per 1M tokens
Pick Kimi K3 when you need image input or Moonshot's larger 2.8T-parameter model for long-horizon agents.
Kimi K3 API$0.037 / $0.15 per 1M tokens
Pick DeepSeek V4.1 Flash for high-volume text and image jobs where the lowest cost per token matters most.
DeepSeek V4.1 Flash APIFAQ
Anything else, write to sales@zurelay.com and an engineer will answer.
GLM-5.2 costs $0.38 per 1M input tokens and $1.21 per 1M output tokens on zurelay, with cached input at $0.038. Z.ai's list price is $1.40 input and $4.40 output, so you save 73%. One rate applies at every prompt length.
Yes. zurelay serves the same glm-5.2 model at 73% below Z.ai's list price. You pay from prepaid credit for the tokens you use, with no subscription or coding plan.
Install the official OpenAI SDK, set the base URL to https://api.zurelay.com/v1 and use your zurelay API key. Then pass model "glm-5.2" to chat.completions.create. Your existing prompts, tool definitions and streaming code work unchanged.
GLM-5.2 has a 1M-token context window and writes up to 128K output tokens per request, up from a 200K context on GLM-5.1. It takes text input only and returns text.
Z.ai built GLM-5.2 for long-horizon coding: auditing whole codebases, running long refactors to completion and holding to a team's engineering rules. It reports 81.0 on Terminal-Bench 2.1 and 62.1 on SWE-bench Pro, up from 62.0 and 58.4 on GLM-5.1.
Yes. Thinking is on by default, and unlike GLM-5.3, GLM-5.2 still lets you switch it off for quick, direct answers. With thinking on, reasoning_effort runs at high or max; max is the default and the one Z.ai recommends for coding.
GLM-5.3 uses the same base model with more post-training, and Z.ai reports a 50% gain on its in-house Code Bench with fewer output tokens per task. Both have a 1M-token context, 128K output and the same Z.ai list price. Choose GLM-5.2 if you need thinking switched off or MIT-licensed weights.
Yes. Set stream: true to receive tokens as they are generated, and add stream_options.include_usage to get token counts in the final chunk. You can give each zurelay API key its own monthly budget and requests-per-minute cap. Failed requests are retried on another route and never billed.
Yes. Requests go to glm-5.2 itself, never a smaller or substitute model. Output is sampled, so wording varies from run to run, just as it does on Z.ai's own API. zurelay is an independent service and is not affiliated with Z.ai.
Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.