Z.ai
Operational· 100% uptime, 7 days

GLM-5.3 API

The GLM-5.3 API for coding agents and code security, 73% below Z.ai's list price

Price per 1M tokens73%off
Input
$0.38$1.40
Output
$1.21$4.40
Cached input
$0.038
Model IDglm-5.3
chat.py
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key="YOUR_ZURELAY_KEY",
)
stream = client.chat.completions.create(
model="glm-5.3",
messages=[{"role": "user", "content": "Explain quicksort in two sentences."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")

Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.

Context window
1M tokens
Max output
128K tokens
Input
Text
Output
Text
Released
Aug 14, 2026
Uptime, 7 days
100%

Pricing

GLM-5.3 API pricing. A fraction of the list price.

Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.

RatezurelayZ.ai listYou save
Input
per 1M tokens
$0.38$1.4073%off
Output
per 1M tokens
$1.21$4.4073%off
Cached input
per 1M tokens
$0.038——

One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.

Savings calculator

Move the sliders to your monthly usage.

Input tokens per month50M
Output tokens per month10M
zurelay
$31.10
Z.ai list price
$114.00
You save every month$82.90

$994.80 a year

Overview

What is GLM-5.3?

GLM-5.3 is Z.ai's flagship model for complex software engineering and long-running agents, with a 1M-token context window and up to 128K output tokens. The zurelay GLM-5.3 API serves the same model through one OpenAI-compatible endpoint, billed per token at 73% below Z.ai's list price.

GLM-5.3 is Z.ai's flagship model for complex coding and agent work. Z.ai announced it on August 14, 2026, first for its GLM Coding Plan, then opened the GLM-5.3 API and published the weights on Hugging Face later that month. The model reads and writes text, holds up to 1M tokens of context and writes up to 128K output tokens per request.

GLM-5.3 uses the same base model as GLM-5.2: a Mixture-of-Experts model with 744B total parameters and 40B active per token. Every gain comes from post-training on long, realistic work environments, some of which represent several days of work for an experienced engineer. Z.ai reports a 50% gain over GLM-5.2 on its in-house Z.ai Code Bench, and Terminal-Bench 3.0 rising from 4.6 to 28.3. It also spends fewer tokens: at max effort it averages about 75K output tokens per Code Bench task, against 96K for GLM-5.2.

Reasoning is always on. Z.ai's API takes a reasoning_effort of low, high or max, with max as the default and the setting Z.ai recommends for coding. Thinking can no longer be switched off, so requests that set thinking to disabled on older GLM models should move to low effort. The model supports function calling, structured JSON output, streaming and context caching. Z.ai also reports a jump in security work: 84.5% on the CyberGym vulnerability discovery benchmark, up from 77.2% for GLM-5.2.

On zurelay you call GLM-5.3 with the OpenAI SDK you already use: set the base URL to https://api.zurelay.com/v1 and the model to glm-5.3. You pay $0.38 per 1M input tokens and $1.21 per 1M output tokens, against Z.ai's list price of $1.40 and $4.40, with cached input at $0.038. The rate is the same at every prompt length, and you pay from prepaid credit with no subscription.

Strengths

Where GLM-5.3 shines. And what teams build with it.

01

Stronger agentic coding

Z.ai reports a 50% gain over GLM-5.2 on its in-house Code Bench, and DeepSWE v1.1 up from 46.2 to 66.9. It was trained on long tasks that mirror real engineering work.

02

Fewer tokens per task

At max effort GLM-5.3 averages about 75K output tokens per Code Bench task, against 96K for GLM-5.2, while scoring higher. At high effort it uses around 50K.

03

Vulnerability discovery

GLM-5.3 scores 84.5% on CyberGym, up from 77.2% for GLM-5.2. Working with security teams, Z.ai says it found 1,097 medium-to-high severity issues in real-world codebases.

04

1M tokens of context

A large repository, long build logs or a full agent trace fits in one request, with up to 128K output tokens for long patches and reports.

Use cases

  • Coding agents

    Run repository-scale refactors, bug fixes and test loops in agent harnesses that call tools over many steps.

  • Terminal automation

    Drive shell-based agents that build, test and debug in a sandbox. Z.ai reports its largest gains on Terminal-Bench 3.0.

  • Secure code review

    Review source code for vulnerabilities and draft fixes before release. Z.ai built GLM-5.3 for white-box code review and vulnerability discovery.

  • Whole-codebase questions

    Ask about architecture, call chains or a large diff across an entire project in one 1M-token request.

Get started

Call GLM-5.3 in three steps.

No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.

  1. 1

    Create a key

    Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.

  2. 2

    Point your SDK at zurelay

    Change the base URL. Everything else in your code stays the same.

    https://api.zurelay.com/v1
  3. 3

    Use glm-5.3

    Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.

    glm-5.3

FAQ

GLM-5.3 API questions.

Anything else, write to sales@zurelay.com and an engineer will answer.

How much does the GLM-5.3 API cost on zurelay?

GLM-5.3 costs $0.38 per 1M input tokens and $1.21 per 1M output tokens on zurelay, with cached input at $0.038. Z.ai's list price is $1.40 input and $4.40 output, so you save 73%. One rate applies at every prompt length.

Is there a cheaper GLM API than Z.ai's own?

Yes. zurelay serves the same glm-5.3 model at 73% below Z.ai's list price. You top up prepaid credit and pay only for the tokens you use, with no subscription or coding plan.

How do I call the GLM-5.3 API with the OpenAI SDK?

Install the official OpenAI SDK, set the base URL to https://api.zurelay.com/v1 and use your zurelay API key. Then pass model "glm-5.3" to chat.completions.create. Set reasoning_effort to low, high or max; max is the default and the one Z.ai recommends for coding.

What is the GLM-5.3 context window?

GLM-5.3 has a 1M-token context window and writes up to 128K output tokens per request. It takes text input only and returns text. If you need image input, look at Kimi K3.

What is GLM-5.3 best at?

Z.ai built GLM-5.3 for complex software engineering and long-running agents. It reports a 50% gain over GLM-5.2 on its in-house Code Bench and top open-model results on Terminal-Bench 3.0. It is also strong at code security work, with 84.5% on the CyberGym vulnerability discovery benchmark.

How is GLM-5.3 different from GLM-5.2?

GLM-5.3 keeps GLM-5.2's base model, 1M-token context and 128K output limit, and Z.ai lists both at the same price. The changes come from post-training: stronger coding, better security results and fewer output tokens per task. GLM-5.3 always reasons at low, high or max effort, while GLM-5.2 can switch thinking off.

Is GLM-5.3 open source?

The weights are on Hugging Face under the GLM-5.3 License. It allows commercial use, modification and redistribution. Only Model-as-a-Service operators with very large annual revenue must first pass a Z.ai security review.

Does the GLM-5.3 API support streaming and per-key limits?

Yes. Set stream: true to receive tokens as they are generated, and add stream_options.include_usage to get token counts in the final chunk. You can give each zurelay API key its own monthly budget and requests-per-minute cap. Failed requests are retried on another route and never billed.

Is GLM-5.3 on zurelay the same model Z.ai serves?

Yes. Requests go to glm-5.3 itself, never a smaller or substitute model. Output is sampled, so wording varies from run to run, just as it does on Z.ai's own API. zurelay is an independent service and is not affiliated with Z.ai.

More models on zurelay

All models

Stop paying list price.
Start saving today.

Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.