Near-Opus coding scores
On Z.ai's in-house Code Bench at max effort it scores 29.0 against 29.5 for Claude Opus 4.8, and it beats GLM-5.2 at every effort level.
The GLM-5.3 Flash API: Z.ai's low-cost GLM for coding and agents, 60% below Z.ai's list price
glm-5.3-flashfrom openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key="YOUR_ZURELAY_KEY",)stream = client.chat.completions.create( model="glm-5.3-flash", messages=[{"role": "user", "content": "Explain quicksort in two sentences."}], stream=True,)for chunk in stream: if chunk.choices: print(chunk.choices[0].delta.content or "", end="")Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK. Chat API docs
Pricing
Pay as you go from prepaid credit, with no subscription. Top up from $10. Requests that fail are never billed.
| Rate | Zurelay | Z.ai list | You save |
|---|---|---|---|
Input per 1M tokens | $0.06 | 60%off | |
Output per 1M tokens | $0.20 | 60%off | |
Cached input per 1M tokens | $0.006 | — | — |
One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.
Move the sliders to your monthly usage.
$90.00 a year
Overview
GLM-5.3 Flash is Z.ai's low-cost GLM, with 18B of its 320B parameters active per token. Z.ai says it outperforms GLM-5.2 and approaches Claude Opus 4.8 on coding and agent benchmarks. The Zurelay GLM-5.3 Flash API serves it through one OpenAI-compatible endpoint with a 1M-token context window, at 60% below Z.ai's list price.
Z.ai released GLM-5.3 Flash on August 26, 2026, as a smaller, cheaper sibling of its GLM-5.3 flagship. The GLM-5.3 Flash API on Zurelay takes text and writes text, with a 1M-token context window and up to 128K output tokens per request. Z.ai built it for coding and agent work: long tool-calling runs, frontend and game code, and office and research workflows. On Zurelay the model ID is glm-5.3-flash, and it costs $0.06 per 1M input tokens and $0.20 per 1M output tokens.
It has its own base model, trained on a new 30T-token corpus, with 320B total parameters and 18B active per token. It is also the first GLM to combine sparse and linear attention, which Z.ai says cuts attention compute by 3.0 times and the KV cache by 4.4 times against GLM-5.3, so long prompts stay cheap to serve. The weights are on Hugging Face under the MIT license.
Z.ai reports that GLM-5.3 Flash consistently outperforms GLM-5.2 across six coding and agent benchmarks, often by a wide margin: 63.4 against 46.2 on DeepSWE v1.1 and 48.8 against 26.2 on AutomationBench. On Z.ai's in-house Code Bench at max effort it nearly matches Claude Opus 4.8, 29.0 against 29.5. Z.ai also reports a score of 57 on the Artificial Analysis Intelligence Index, a level it says used to cost about ten times as much.
Reasoning is always on. Z.ai's API takes a reasoning_effort of low, high or max, with max as the default and the setting Z.ai recommends, and thinking can't be switched off. The model supports function calling, structured JSON output, streaming and context caching. You pay $0.06 per 1M input tokens and $0.20 per 1M output tokens, against Z.ai's list price of $0.15 and $0.50, with cached input at $0.006.
Strengths
On Z.ai's in-house Code Bench at max effort it scores 29.0 against 29.5 for Claude Opus 4.8, and it beats GLM-5.2 at every effort level.
Z.ai reports it ahead of GLM-5.2 on all six of its coding and agent benchmarks: 63.4 against 46.2 on DeepSWE v1.1 and 48.8 against 26.2 on AutomationBench.
Hybrid sparse and linear attention cuts attention compute by 3.0 times and the KV cache by 4.4 times against GLM-5.3, across a 1M-token window.
The weights are open under MIT, so you can build on Zurelay today and self-host the same model later if you need to.
Use cases
Run long tool-calling loops that read code, edit files and run tests, at about a tenth of GLM-5.3's list price.
Write web pages, interfaces, games and 3D scenes from a written spec, then iterate on the code with test results and error messages.
Plan research and analysis, call tools and review the results. Z.ai built it for office documents and financial research workflows.
Put whole codebases, contracts or report sets into the 1M-token window and ask across all of them at once.
Get started
No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to Zurelay.
Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.
Change the base URL. Everything else in your code stays the same.
https://api.zurelay.com/v1Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.
glm-5.3-flashWorks with the tools you already use
Compare
$0.38 / $1.21 per 1M tokens
Pick GLM-5.3, Z.ai's larger text-only flagship, for the hardest coding and security work; it scores 66.9 on DeepSWE v1.1 against 63.4.
GLM-5.3 API$0.07 / $0.22 per 1M tokens
Pick Qwen 3.8 Flash for Alibaba's low-cost open-weight model when you also need image and video input.
Qwen 3.8 Flash API$0.03 / $0.12 per 1M tokens
Pick DeepSeek V4 Flash for an even cheaper reasoning model with a 1M-token context, and compare the two on your own evals.
DeepSeek V4 Flash APIGLM-5.3 Flash costs $0.06 per 1M input tokens and $0.20 per 1M output tokens on Zurelay, with cached input at $0.006. Z.ai's list price is $0.15 input and $0.50 output, so you save 60%. One rate applies at every prompt length.
Yes. Zurelay serves the same glm-5.3-flash model at 60% below Z.ai's list price. You pay from prepaid credit for the tokens you use, with no subscription or coding plan.
Install the official OpenAI SDK, set the base URL to https://api.zurelay.com/v1 and use your Zurelay API key. Then pass model "glm-5.3-flash" to chat.completions.create. Z.ai recommends temperature 1, top_p 0.95 and a reasoning_effort of max.
GLM-5.3 Flash has a 1M-token context window and writes up to 128K output tokens per request. On Zurelay it takes text input and returns text.
Not on Zurelay yet. Z.ai's own API takes images and video, but Zurelay serves GLM-5.3 Flash for text only, and a request with an image is turned away with a clear error before anything is billed. For images and video at a similar price, use Qwen 3.8 Flash.
Coding and agent work: long tool-calling runs, frontend and game development and 3D scenes. Z.ai also built it for office documents and financial research, and reports it ahead of GLM-5.2 across six coding and agent benchmarks.
GLM-5.3 is Z.ai's larger, text-only flagship and scores higher on DeepSWE v1.1, 66.9 against 63.4. GLM-5.3 Flash is a smaller model with 18B active parameters, and Z.ai lists it at about a tenth of GLM-5.3's price. Both have a 1M-token context, 128K output and always-on reasoning at low, high or max.
Yes. Set stream: true to receive tokens as they are generated, and add stream_options.include_usage to get token counts in the final chunk. You can give each Zurelay API key its own monthly budget and requests-per-minute cap. Failed requests are retried on another route and never billed.
Yes. Requests go to glm-5.3-flash itself. Zurelay doesn't swap in a smaller or substitute model. Only smart routing, which you turn on yourself, can answer with a close alternative when the model is down, and the x-zurelay-model header names the model that answered. Output is sampled, so wording varies from run to run, just as it does on Z.ai's own API. Zurelay is an independent service and is not affiliated with Z.ai.
Create a key in seconds, point the SDK you already use at Zurelay, and every request costs up to 90% less from the first token.