Tencent
Operational· 100% uptime, 7 days

HY4 Preview API

The HY4 Preview API: Tencent's agent and coding model, 52% below Tencent's price

Price per 1M tokens52%off
Input
$0.40$0.834
Output
$1.20$2.501
Cached input
$0.02
Model IDhy4-preview
chat.py
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key="YOUR_ZURELAY_KEY",
)
stream = client.chat.completions.create(
model="hy4-preview",
messages=[{"role": "user", "content": "Explain quicksort in two sentences."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")

Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.

Context window
1M tokens
Max output
64K tokens
Input
Text
Output
Text
Released
Aug 28, 2026
Uptime, 7 days
100%

Pricing

HY4 Preview API pricing. A fraction of the list price.

Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.

RatezurelayTencent listYou save
Input
per 1M tokens
$0.40$0.83452%off
Output
per 1M tokens
$1.20$2.50152%off
Cached input
per 1M tokens
$0.02——

One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.

Savings calculator

Move the sliders to your monthly usage.

Input tokens per month50M
Output tokens per month10M
zurelay
$32.00
Tencent list price
$66.71
You save every month$34.71

$416.52 a year

Overview

What is HY4 Preview?

HY4 Preview is Tencent's new HY model (formerly Hunyuan), with 770B total parameters and 49B active, built for agents, coding and productivity work. The zurelay HY4 Preview API serves it through one OpenAI-compatible endpoint with a 1M-token context window, at 52% below Tencent's official price.

Tencent released HY4 Preview on August 28, 2026, as an open-source model. HY is the new name for Tencent's Hunyuan models. HY4 Preview has 770B total parameters, of which 49B are active per token, and Tencent built it for agents, coding and productivity: multi-step agent workflows, tool calling, code generation and long chains of execution.

Tencent reports gains in long-context development, where it says the model is better at understanding, planning, debugging and validating code, and in office work such as financial analysis, data analysis and work across several documents. In an internal blind evaluation with 163 experts and 203 engineering tasks, Tencent says HY4 Preview averaged 2.99 out of 4, slightly ahead of GLM-5.3 at 2.92 and Kimi K3 at 2.94.

It is a reasoning model: it thinks before it answers, and that thinking counts toward the output budget. The context window is 1,048,576 tokens and a response can run to 64K output tokens, so give it a generous max_tokens on hard prompts, or the answer may be cut short after the reasoning.

On zurelay the model ID is hy4-preview, and tencent/hy4-preview and hy4 work as aliases. You pay $0.40 per 1M input tokens and $1.20 per 1M output tokens, with cache hits at $0.02, against Tencent's official price of $0.834 and $2.501.

Strengths

Where HY4 Preview shines. And what teams build with it.

01

Built for agents

Tencent designed HY4 Preview for multi-step agent workflows, tool calling and long chains of execution.

02

Long-context coding

Tencent reports stronger understanding, planning, debugging and validation on long-context development tasks.

03

Office and analysis work

Tencent highlights complex working environments, financial and data analysis, and work that spans several documents.

04

1M tokens of context

1,048,576 tokens hold a large repository, a long agent trace or a stack of documents in one request.

Use cases

  • Coding agents

    Plan, write, debug and validate code across a large repository, with tool calls in the loop.

  • Agent workflows

    Run long chains of tool calls, such as research, data pulls and ticket handling, where the model has to keep track of many steps.

  • Financial and data analysis

    Read reports and exported data, compare figures across several documents and draft the analysis.

  • Game prototypes

    Turn a short game idea into a playable prototype. Tencent shows the model doing this from a single natural-language request.

Get started

Call HY4 Preview in three steps.

No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.

  1. 1

    Create a key

    Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.

  2. 2

    Point your SDK at zurelay

    Change the base URL. Everything else in your code stays the same.

    https://api.zurelay.com/v1
  3. 3

    Use hy4-preview

    Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.

    hy4-preview

FAQ

HY4 Preview API questions.

Anything else, write to sales@zurelay.com and an engineer will answer.

How much does the HY4 Preview API cost on zurelay?

On zurelay, hy4-preview costs $0.40 per 1M input tokens and $1.20 per 1M output tokens, with cache hits at $0.02. Tencent's official price is $0.834 input and $2.501 output, so you save 52%. Reasoning tokens bill as output.

Is there a cheaper HY4 Preview API than Tencent's?

Yes. zurelay serves hy4-preview at 52% below Tencent's official price. It is the same model, not a smaller substitute. You top up prepaid credit and pay only for the tokens you use, with no subscription.

How do I call the HY4 Preview API with the OpenAI SDK?

Set the OpenAI SDK's base URL to https://api.zurelay.com/v1, use your zurelay API key and pass model "hy4-preview". zurelay also accepts tencent/hy4-preview and hy4. Set max_tokens high enough for both the reasoning and the answer, since the model thinks before it replies.

What is the HY4 Preview context window?

HY4 Preview holds 1,048,576 tokens of context and returns up to 64K output tokens per response. Reasoning and the answer share that output budget, so leave headroom on hard prompts.

Is HY4 Preview a reasoning model?

Yes. It thinks before it answers, and those reasoning tokens count toward max_tokens and bill as output. Give it a generous max_tokens so the reasoning doesn't use up the budget before the answer starts.

What is HY4 Preview best at?

Tencent built it for agents, coding and productivity: multi-step agent workflows, tool calling, code generation and long chains of execution. Tencent also reports gains in office work, such as financial and data analysis, and on complex research problems.

Is Tencent HY the same as Hunyuan?

Yes. HY is the new name for Tencent's Hunyuan models, and HY4 Preview belongs to the HY4 series. Tencent released it as an open-source model on August 28, 2026.

Does the HY4 Preview API support streaming and per-key rate limits?

Yes. Set stream: true to receive tokens as they are generated, and add stream_options.include_usage to get token counts in the final chunk. Each zurelay API key can have its own monthly budget and requests-per-minute cap. Failed attempts are retried automatically and never billed.

Is HY4 Preview on zurelay the same model Tencent serves?

Yes. Requests go to hy4-preview itself, never a smaller or substitute model. Output is sampled, so wording varies from run to run, as it does on Tencent's own API. zurelay is an independent service and is not affiliated with or endorsed by Tencent.

More models on zurelay

All models

Stop paying list price.
Start saving today.

Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.