Anthropic
Operational· 100% uptime, 7 days

Claude Haiku 4.5 API

The Claude Haiku 4.5 API: Anthropic's fastest model, 62% below list price.

Price per 1M tokens62%off
Input
$0.38$1.00
Output
$1.90$5.00
Cached input
$0.038
Model IDclaude-haiku-4-5
chat.py
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key="YOUR_ZURELAY_KEY",
)
stream = client.chat.completions.create(
model="claude-haiku-4-5",
messages=[{"role": "user", "content": "Explain quicksort in two sentences."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")

Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK, and the Anthropic SDK and Claude Code.

Context window
200K tokens
Max output
64K tokens
Input
Text, Image
Output
Text
Released
Oct 15, 2025
Uptime, 7 days
100%

Pricing

Claude Haiku 4.5 API pricing. A fraction of the list price.

Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.

RatezurelayAnthropic listYou save
Input
per 1M tokens
$0.38$1.0062%off
Output
per 1M tokens
$1.90$5.0062%off
Cached input
per 1M tokens
$0.038——

One rate at every prompt length, up to the full 200K context window. Streaming, tool calls and structured outputs cost nothing extra.

Savings calculator

Move the sliders to your monthly usage.

Input tokens per month50M
Output tokens per month10M
zurelay
$38.00
Anthropic list price
$100.00
You save every month$62.00

$744.00 a year

Overview

What is Claude Haiku 4.5?

Claude Haiku 4.5 is Anthropic's fastest model, which Anthropic describes as having near-frontier intelligence. The zurelay Claude Haiku 4.5 API serves it through OpenAI-compatible and Anthropic-compatible endpoints at 62% below Anthropic's list price. It fits chat, sub-agents and high-volume work where speed and cost per call matter most.

Claude Haiku 4.5 was released on October 15, 2025 and is still Anthropic's current Haiku model. Anthropic describes it as the fastest model with near-frontier intelligence and rates it the fastest in its current lineup. At launch, Anthropic said it matched Claude Sonnet 4 on coding at one-third the cost and more than twice the speed. On zurelay, the Claude Haiku 4.5 API uses the model ID claude-haiku-4-5 and costs $0.38 per 1M input tokens and $1.90 per 1M output tokens.

The specs: a 200K-token context window, up to 64K output tokens per request, text and image input, and text output. Its reliable knowledge cutoff is February 2025, with training data through July 2025. Haiku 4.5 supports extended thinking: set thinking type enabled with a budget_tokens value of at least 1,024 and below max_tokens. It has no effort setting and no interleaved thinking between tool calls. It tracks its remaining context window on its own, which Anthropic calls context awareness.

Anthropic reported 73.3% on SWE-bench Verified at launch and said Haiku 4.5 beats Claude Sonnet 4 at some tasks, such as using computers. It targets real-time, low-latency work like chat assistants, customer service agents and pair programming. Anthropic also pitches it as a sub-agent: a larger Claude model breaks a problem into steps, then a team of Haiku 4.5 instances completes the subtasks in parallel.

Haiku 4.5 is active, with retirement no sooner than October 15, 2026, and Anthropic has not announced a deprecation as of September 30, 2026. It has the lowest list price in Anthropic's current lineup. Prompt caching needs at least 4,096 tokens in the cached prompt, more than current Sonnet and Fable models require, so cache long, stable system prompts and documents.

Strengths

Where Claude Haiku 4.5 shines. And what teams build with it.

01

Fastest in the lineup

Anthropic rates Haiku 4.5 its fastest current model, a good match for replies a user is waiting on.

02

Lowest price per call

It has the lowest list price among Anthropic's current models, and zurelay takes a further 62% off.

03

Strong coding for its size

Anthropic reported 73.3% on SWE-bench Verified at launch, with coding on par with Claude Sonnet 4.

04

Built for sub-agents

Fast and cheap enough to run many copies in parallel under a larger Claude model that plans and reviews.

Use cases

  • Chat and support assistants

    Real-time chat and customer service agents where a user waits on every reply.

  • Sub-agent workers

    Parallel workers that search, read, extract or edit while a larger Claude model plans the task and checks the results.

  • Pair programming

    Quick completions, explanations and small fixes inside an editor, where latency matters more than depth.

  • High-volume extraction

    Tagging, routing, summarizing and pulling structured fields from large batches of text or images.

Get started

Call Claude Haiku 4.5 in three steps.

No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.

  1. 1

    Create a key

    Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.

  2. 2

    Point your SDK at zurelay

    Change the base URL. Everything else in your code stays the same.

    https://api.zurelay.com/v1
  3. 3

    Use claude-haiku-4-5

    Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.

    claude-haiku-4-5

FAQ

Claude Haiku 4.5 API questions.

Anything else, write to sales@zurelay.com and an engineer will answer.

How much does the Claude Haiku 4.5 API cost on zurelay?

zurelay charges $0.38 per 1M input tokens and $1.90 per 1M output tokens, with cached input at $0.038. Anthropic's list price is $1.00 input and $5.00 output, so you save 62%. The rate is the same at every prompt length.

Is there a cheaper Claude API than Anthropic's?

Yes. Claude Haiku 4.5 already has the lowest list price in Anthropic's current lineup, and zurelay serves claude-haiku-4-5 at 62% below that. It is the same model, paid from prepaid credit with no subscription. You keep your code and change the base URL and key.

How do I call Claude Haiku 4.5 with the OpenAI SDK?

Create a zurelay API key, set the base URL to https://api.zurelay.com/v1 and set the model to claude-haiku-4-5. Chat completions, streaming and tool calls use the request shapes you know from OpenAI, and temperature and top_p work as usual.

Can I use Claude Haiku 4.5 with the Anthropic SDK or Claude Code?

Yes. Set the Anthropic SDK's base URL, or ANTHROPIC_BASE_URL for Claude Code, to https://api.zurelay.com with no /v1, because the client adds /v1/messages itself. Authenticate with your zurelay key and choose claude-haiku-4-5 as the model.

What is the Claude Haiku 4.5 context window?

Claude Haiku 4.5 has a 200K-token context window, roughly 150k words, and returns up to 64K output tokens per request. For larger inputs, use a model with a 1M-token window, such as Claude Sonnet 5.5.

Does Claude Haiku 4.5 support thinking?

Yes, in extended mode. In the Messages format, send thinking type enabled with a budget_tokens value of at least 1,024 and below max_tokens. Haiku 4.5 does not support adaptive thinking, the effort setting, or interleaved thinking between tool calls.

What is Claude Haiku 4.5 best at?

Fast, frequent calls: chat assistants, customer support, pair programming, sub-agents and bulk extraction. Anthropic reported coding on par with Claude Sonnet 4 at launch. For hard multi-step reasoning, a Sonnet or Opus model is a better fit.

Does the zurelay Claude Haiku 4.5 API support streaming, and what are the limits?

Yes. Streaming works on chat completions and on the Messages API; on chat completions, set stream_options.include_usage to get token counts at the end of the stream. Each key can have its own monthly budget and requests-per-minute cap. If a route fails, zurelay retries on another route, and failed requests are never billed.

Is Claude Haiku 4.5 on zurelay the same model as Anthropic's?

Yes. Your requests run on Claude Haiku 4.5 itself, and zurelay never swaps in a different or smaller model. Output is sampled, so exact wording varies between calls on any API. zurelay is an independent service and is not affiliated with or endorsed by Anthropic.

More models on zurelay

All models

Stop paying list price.
Start saving today.

Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.