Z.ai
Operational· 100% uptime, 7 days

GLM-5.2 API

The GLM-5.2 API with 1M tokens of context for long coding tasks, 73% below list

Price per 1M tokens73%off
Input
$0.38$1.40
Output
$1.21$4.40
Cached input
$0.038
Model IDglm-5.2
chat.py
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key="YOUR_ZURELAY_KEY",
)
stream = client.chat.completions.create(
model="glm-5.2",
messages=[{"role": "user", "content": "Explain quicksort in two sentences."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")

Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.

Context window
1M tokens
Max output
128K tokens
Input
Text
Output
Text
Released
Jun 16, 2026
Uptime, 7 days
100%

Pricing

GLM-5.2 API pricing. A fraction of the list price.

Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.

RatezurelayZ.ai listYou save
Input
per 1M tokens
$0.38$1.4073%off
Output
per 1M tokens
$1.21$4.4073%off
Cached input
per 1M tokens
$0.038——

One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.

Savings calculator

Move the sliders to your monthly usage.

Input tokens per month50M
Output tokens per month10M
zurelay
$31.10
Z.ai list price
$114.00
You save every month$82.90

$994.80 a year

Overview

What is GLM-5.2?

GLM-5.2 is Z.ai's open-weight model for long-horizon coding and the first GLM with a 1M-token context window. The zurelay GLM-5.2 API serves the same model through one OpenAI-compatible endpoint at 73% below Z.ai's list price.

GLM-5.2 is Z.ai's model for long-horizon tasks, released on June 16, 2026 after a few days of early access for GLM Coding Plan users. The GLM-5.2 API takes text, returns text, holds up to 1M tokens of context and writes up to 128K output tokens per request. That is five times the 200K context of GLM-5.1, and Z.ai trained the model for months on long coding-agent work so the whole window stays usable.

It is a Mixture-of-Experts model with 744B total parameters and 40B active per token, released with open weights under the MIT license. A technique Z.ai calls IndexShare reuses one sparse-attention indexer across every four layers, which cuts per-token compute by 2.9 times at 1M context. Z.ai also improved the multi-token prediction layer used for speculative decoding, raising acceptance length by up to 20%.

On coding benchmarks Z.ai reports 81.0 on Terminal-Bench 2.1 and 62.1 on SWE-bench Pro, up from 62.0 and 58.4 for GLM-5.1. In practice Z.ai points to project-level work: auditing a full codebase, running long refactors to completion, following a team's lint, build and commit rules, and debugging Android apps on a real device with ADB and logcat.

Thinking is on by default, and you can turn it off. With thinking on, Z.ai offers two reasoning_effort levels, high and max, with max as the default and the setting Z.ai recommends for coding. The model supports function calling, structured JSON output, streaming and context caching. On zurelay you call it at https://api.zurelay.com/v1 with the model ID glm-5.2, for $0.38 per 1M input tokens and $1.21 per 1M output tokens against Z.ai's list price of $1.40 and $4.40.

Strengths

Where GLM-5.2 shines. And what teams build with it.

01

A usable 1M-token window

Z.ai trained GLM-5.2 for months on long coding-agent tasks so quality holds deep into the context. A whole service with its tests and docs fits in one request.

02

Long refactors that finish

It breaks a goal into stages, tracks dependencies and verifies as it goes. That suits module decoupling, API migrations and cross-language rewrites.

03

Follows engineering rules

Z.ai reports more reliable adherence to code style, architecture boundaries and commit conventions over long sessions, with fewer out-of-scope changes.

04

MIT-licensed weights

The weights are open under MIT, so you can build on zurelay today and self-host the same model later if you need to.

Use cases

  • Codebase audits

    Load a real repository and ask for an architecture map, API contracts, data flows and technical debt in one pass.

  • Long-running refactors

    Hand the model a bounded refactor, let it plan and change files across the project, then run the tests to confirm the result.

  • Mobile and Mini Program builds

    Build Android clients or WeChat Mini Programs from existing web code, then debug from ADB and logcat output.

  • Research reproduction

    Turn a paper's model, loss functions and data pipeline into a runnable project and check the metrics against the paper.

Get started

Call GLM-5.2 in three steps.

No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.

  1. 1

    Create a key

    Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.

  2. 2

    Point your SDK at zurelay

    Change the base URL. Everything else in your code stays the same.

    https://api.zurelay.com/v1
  3. 3

    Use glm-5.2

    Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.

    glm-5.2

FAQ

GLM-5.2 API questions.

Anything else, write to sales@zurelay.com and an engineer will answer.

How much does the GLM-5.2 API cost?

GLM-5.2 costs $0.38 per 1M input tokens and $1.21 per 1M output tokens on zurelay, with cached input at $0.038. Z.ai's list price is $1.40 input and $4.40 output, so you save 73%. One rate applies at every prompt length.

Is there a cheaper GLM-5.2 API than Z.ai's?

Yes. zurelay serves the same glm-5.2 model at 73% below Z.ai's list price. You pay from prepaid credit for the tokens you use, with no subscription or coding plan.

How do I call the GLM-5.2 API with the OpenAI SDK?

Install the official OpenAI SDK, set the base URL to https://api.zurelay.com/v1 and use your zurelay API key. Then pass model "glm-5.2" to chat.completions.create. Your existing prompts, tool definitions and streaming code work unchanged.

What is the GLM-5.2 context window?

GLM-5.2 has a 1M-token context window and writes up to 128K output tokens per request, up from a 200K context on GLM-5.1. It takes text input only and returns text.

What is GLM-5.2 best at?

Z.ai built GLM-5.2 for long-horizon coding: auditing whole codebases, running long refactors to completion and holding to a team's engineering rules. It reports 81.0 on Terminal-Bench 2.1 and 62.1 on SWE-bench Pro, up from 62.0 and 58.4 on GLM-5.1.

Can I turn off thinking on GLM-5.2?

Yes. Thinking is on by default, and unlike GLM-5.3, GLM-5.2 still lets you switch it off for quick, direct answers. With thinking on, reasoning_effort runs at high or max; max is the default and the one Z.ai recommends for coding.

Should I use GLM-5.2 or GLM-5.3?

GLM-5.3 uses the same base model with more post-training, and Z.ai reports a 50% gain on its in-house Code Bench with fewer output tokens per task. Both have a 1M-token context, 128K output and the same Z.ai list price. Choose GLM-5.2 if you need thinking switched off or MIT-licensed weights.

Does the GLM-5.2 API support streaming and per-key limits?

Yes. Set stream: true to receive tokens as they are generated, and add stream_options.include_usage to get token counts in the final chunk. You can give each zurelay API key its own monthly budget and requests-per-minute cap. Failed requests are retried on another route and never billed.

Is GLM-5.2 on zurelay the same model Z.ai serves?

Yes. Requests go to glm-5.2 itself, never a smaller or substitute model. Output is sampled, so wording varies from run to run, just as it does on Z.ai's own API. zurelay is an independent service and is not affiliated with Z.ai.

More models on zurelay

All models

Stop paying list price.
Start saving today.

Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.