Stronger agentic coding
Z.ai reports a 50% gain over GLM-5.2 on its in-house Code Bench, and DeepSWE v1.1 up from 46.2 to 66.9. It was trained on long tasks that mirror real engineering work.
The GLM-5.3 API for coding agents and code security, 73% below Z.ai's list price
glm-5.3from openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key="YOUR_ZURELAY_KEY",)stream = client.chat.completions.create( model="glm-5.3", messages=[{"role": "user", "content": "Explain quicksort in two sentences."}], stream=True,)for chunk in stream: if chunk.choices: print(chunk.choices[0].delta.content or "", end="")Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.
Pricing
Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.
| Rate | zurelay | Z.ai list | You save |
|---|---|---|---|
Input per 1M tokens | $0.38 | 73%off | |
Output per 1M tokens | $1.21 | 73%off | |
Cached input per 1M tokens | $0.038 | — | — |
One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.
Move the sliders to your monthly usage.
$994.80 a year
Overview
GLM-5.3 is Z.ai's flagship model for complex software engineering and long-running agents, with a 1M-token context window and up to 128K output tokens. The zurelay GLM-5.3 API serves the same model through one OpenAI-compatible endpoint, billed per token at 73% below Z.ai's list price.
GLM-5.3 is Z.ai's flagship model for complex coding and agent work. Z.ai announced it on August 14, 2026, first for its GLM Coding Plan, then opened the GLM-5.3 API and published the weights on Hugging Face later that month. The model reads and writes text, holds up to 1M tokens of context and writes up to 128K output tokens per request.
GLM-5.3 uses the same base model as GLM-5.2: a Mixture-of-Experts model with 744B total parameters and 40B active per token. Every gain comes from post-training on long, realistic work environments, some of which represent several days of work for an experienced engineer. Z.ai reports a 50% gain over GLM-5.2 on its in-house Z.ai Code Bench, and Terminal-Bench 3.0 rising from 4.6 to 28.3. It also spends fewer tokens: at max effort it averages about 75K output tokens per Code Bench task, against 96K for GLM-5.2.
Reasoning is always on. Z.ai's API takes a reasoning_effort of low, high or max, with max as the default and the setting Z.ai recommends for coding. Thinking can no longer be switched off, so requests that set thinking to disabled on older GLM models should move to low effort. The model supports function calling, structured JSON output, streaming and context caching. Z.ai also reports a jump in security work: 84.5% on the CyberGym vulnerability discovery benchmark, up from 77.2% for GLM-5.2.
On zurelay you call GLM-5.3 with the OpenAI SDK you already use: set the base URL to https://api.zurelay.com/v1 and the model to glm-5.3. You pay $0.38 per 1M input tokens and $1.21 per 1M output tokens, against Z.ai's list price of $1.40 and $4.40, with cached input at $0.038. The rate is the same at every prompt length, and you pay from prepaid credit with no subscription.
Strengths
Z.ai reports a 50% gain over GLM-5.2 on its in-house Code Bench, and DeepSWE v1.1 up from 46.2 to 66.9. It was trained on long tasks that mirror real engineering work.
At max effort GLM-5.3 averages about 75K output tokens per Code Bench task, against 96K for GLM-5.2, while scoring higher. At high effort it uses around 50K.
GLM-5.3 scores 84.5% on CyberGym, up from 77.2% for GLM-5.2. Working with security teams, Z.ai says it found 1,097 medium-to-high severity issues in real-world codebases.
A large repository, long build logs or a full agent trace fits in one request, with up to 128K output tokens for long patches and reports.
Use cases
Run repository-scale refactors, bug fixes and test loops in agent harnesses that call tools over many steps.
Drive shell-based agents that build, test and debug in a sandbox. Z.ai reports its largest gains on Terminal-Bench 3.0.
Review source code for vulnerabilities and draft fixes before release. Z.ai built GLM-5.3 for white-box code review and vulnerability discovery.
Ask about architecture, call chains or a large diff across an entire project in one 1M-token request.
Get started
No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.
Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.
Change the base URL. Everything else in your code stays the same.
https://api.zurelay.com/v1Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.
glm-5.3Works with the tools you already use
Compare
$0.38 / $1.21 per 1M tokens
Pick GLM-5.2 when you need thinking switched off or MIT-licensed weights; Z.ai lists both models at the same price.
GLM-5.2 API$1.27 / $6.37 per 1M tokens
Pick Kimi K3 when your agent needs image input or you want Moonshot's larger 2.8T-parameter model for long-horizon coding.
Kimi K3 API$3.50 / $17.50 per 1M tokens
Pick Claude Fable 5 when top coding results matter more than cost; Z.ai's own Code Bench still puts it ahead of GLM-5.3.
Claude Fable 5 APIFAQ
Anything else, write to sales@zurelay.com and an engineer will answer.
GLM-5.3 costs $0.38 per 1M input tokens and $1.21 per 1M output tokens on zurelay, with cached input at $0.038. Z.ai's list price is $1.40 input and $4.40 output, so you save 73%. One rate applies at every prompt length.
Yes. zurelay serves the same glm-5.3 model at 73% below Z.ai's list price. You top up prepaid credit and pay only for the tokens you use, with no subscription or coding plan.
Install the official OpenAI SDK, set the base URL to https://api.zurelay.com/v1 and use your zurelay API key. Then pass model "glm-5.3" to chat.completions.create. Set reasoning_effort to low, high or max; max is the default and the one Z.ai recommends for coding.
GLM-5.3 has a 1M-token context window and writes up to 128K output tokens per request. It takes text input only and returns text. If you need image input, look at Kimi K3.
Z.ai built GLM-5.3 for complex software engineering and long-running agents. It reports a 50% gain over GLM-5.2 on its in-house Code Bench and top open-model results on Terminal-Bench 3.0. It is also strong at code security work, with 84.5% on the CyberGym vulnerability discovery benchmark.
GLM-5.3 keeps GLM-5.2's base model, 1M-token context and 128K output limit, and Z.ai lists both at the same price. The changes come from post-training: stronger coding, better security results and fewer output tokens per task. GLM-5.3 always reasons at low, high or max effort, while GLM-5.2 can switch thinking off.
The weights are on Hugging Face under the GLM-5.3 License. It allows commercial use, modification and redistribution. Only Model-as-a-Service operators with very large annual revenue must first pass a Z.ai security review.
Yes. Set stream: true to receive tokens as they are generated, and add stream_options.include_usage to get token counts in the final chunk. You can give each zurelay API key its own monthly budget and requests-per-minute cap. Failed requests are retried on another route and never billed.
Yes. Requests go to glm-5.3 itself, never a smaller or substitute model. Output is sampled, so wording varies from run to run, just as it does on Z.ai's own API. zurelay is an independent service and is not affiliated with Z.ai.
Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.