Built for agents
Tencent designed HY4 Preview for multi-step agent workflows, tool calling and long chains of execution.
The HY4 Preview API: Tencent's agent and coding model, 52% below Tencent's price
hy4-previewfrom openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key="YOUR_ZURELAY_KEY",)stream = client.chat.completions.create( model="hy4-preview", messages=[{"role": "user", "content": "Explain quicksort in two sentences."}], stream=True,)for chunk in stream: if chunk.choices: print(chunk.choices[0].delta.content or "", end="")Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.
Pricing
Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.
| Rate | zurelay | Tencent list | You save |
|---|---|---|---|
Input per 1M tokens | $0.40 | 52%off | |
Output per 1M tokens | $1.20 | 52%off | |
Cached input per 1M tokens | $0.02 | — | — |
One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.
Move the sliders to your monthly usage.
$416.52 a year
Overview
HY4 Preview is Tencent's new HY model (formerly Hunyuan), with 770B total parameters and 49B active, built for agents, coding and productivity work. The zurelay HY4 Preview API serves it through one OpenAI-compatible endpoint with a 1M-token context window, at 52% below Tencent's official price.
Tencent released HY4 Preview on August 28, 2026, as an open-source model. HY is the new name for Tencent's Hunyuan models. HY4 Preview has 770B total parameters, of which 49B are active per token, and Tencent built it for agents, coding and productivity: multi-step agent workflows, tool calling, code generation and long chains of execution.
Tencent reports gains in long-context development, where it says the model is better at understanding, planning, debugging and validating code, and in office work such as financial analysis, data analysis and work across several documents. In an internal blind evaluation with 163 experts and 203 engineering tasks, Tencent says HY4 Preview averaged 2.99 out of 4, slightly ahead of GLM-5.3 at 2.92 and Kimi K3 at 2.94.
It is a reasoning model: it thinks before it answers, and that thinking counts toward the output budget. The context window is 1,048,576 tokens and a response can run to 64K output tokens, so give it a generous max_tokens on hard prompts, or the answer may be cut short after the reasoning.
On zurelay the model ID is hy4-preview, and tencent/hy4-preview and hy4 work as aliases. You pay $0.40 per 1M input tokens and $1.20 per 1M output tokens, with cache hits at $0.02, against Tencent's official price of $0.834 and $2.501.
Strengths
Tencent designed HY4 Preview for multi-step agent workflows, tool calling and long chains of execution.
Tencent reports stronger understanding, planning, debugging and validation on long-context development tasks.
Tencent highlights complex working environments, financial and data analysis, and work that spans several documents.
1,048,576 tokens hold a large repository, a long agent trace or a stack of documents in one request.
Use cases
Plan, write, debug and validate code across a large repository, with tool calls in the loop.
Run long chains of tool calls, such as research, data pulls and ticket handling, where the model has to keep track of many steps.
Read reports and exported data, compare figures across several documents and draft the analysis.
Turn a short game idea into a playable prototype. Tencent shows the model doing this from a single natural-language request.
Get started
No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.
Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.
Change the base URL. Everything else in your code stays the same.
https://api.zurelay.com/v1Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.
hy4-previewWorks with the tools you already use
Compare
$1.27 / $6.37 per 1M tokens
Pick Kimi K3 for long-horizon coding with image input; Tencent's own blind evaluation scored the two close together.
Kimi K3 API$0.38 / $1.21 per 1M tokens
Pick GLM-5.3 to compare another coding and agent model that Tencent's blind evaluation placed just behind HY4 Preview.
GLM-5.3 API$1.20 / $3.60 per 1M tokens
Pick Qwen 3.8 Max for Alibaba's larger flagship, with image and video input and a 1M-token context.
Qwen 3.8 Max APIFAQ
Anything else, write to sales@zurelay.com and an engineer will answer.
On zurelay, hy4-preview costs $0.40 per 1M input tokens and $1.20 per 1M output tokens, with cache hits at $0.02. Tencent's official price is $0.834 input and $2.501 output, so you save 52%. Reasoning tokens bill as output.
Yes. zurelay serves hy4-preview at 52% below Tencent's official price. It is the same model, not a smaller substitute. You top up prepaid credit and pay only for the tokens you use, with no subscription.
Set the OpenAI SDK's base URL to https://api.zurelay.com/v1, use your zurelay API key and pass model "hy4-preview". zurelay also accepts tencent/hy4-preview and hy4. Set max_tokens high enough for both the reasoning and the answer, since the model thinks before it replies.
HY4 Preview holds 1,048,576 tokens of context and returns up to 64K output tokens per response. Reasoning and the answer share that output budget, so leave headroom on hard prompts.
Yes. It thinks before it answers, and those reasoning tokens count toward max_tokens and bill as output. Give it a generous max_tokens so the reasoning doesn't use up the budget before the answer starts.
Tencent built it for agents, coding and productivity: multi-step agent workflows, tool calling, code generation and long chains of execution. Tencent also reports gains in office work, such as financial and data analysis, and on complex research problems.
Yes. HY is the new name for Tencent's Hunyuan models, and HY4 Preview belongs to the HY4 series. Tencent released it as an open-source model on August 28, 2026.
Yes. Set stream: true to receive tokens as they are generated, and add stream_options.include_usage to get token counts in the final chunk. Each zurelay API key can have its own monthly budget and requests-per-minute cap. Failed attempts are retried automatically and never billed.
Yes. Requests go to hy4-preview itself, never a smaller or substitute model. Output is sampled, so wording varies from run to run, as it does on Tencent's own API. zurelay is an independent service and is not affiliated with or endorsed by Tencent.
Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.