Fastest in the lineup
Anthropic rates Haiku 4.5 its fastest current model, a good match for replies a user is waiting on.
The Claude Haiku 4.5 API: Anthropic's fastest model, 62% below list price.
claude-haiku-4-5from openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key="YOUR_ZURELAY_KEY",)stream = client.chat.completions.create( model="claude-haiku-4-5", messages=[{"role": "user", "content": "Explain quicksort in two sentences."}], stream=True,)for chunk in stream: if chunk.choices: print(chunk.choices[0].delta.content or "", end="")Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK, and the Anthropic SDK and Claude Code.
Pricing
Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.
| Rate | zurelay | Anthropic list | You save |
|---|---|---|---|
Input per 1M tokens | $0.38 | 62%off | |
Output per 1M tokens | $1.90 | 62%off | |
Cached input per 1M tokens | $0.038 | — | — |
One rate at every prompt length, up to the full 200K context window. Streaming, tool calls and structured outputs cost nothing extra.
Move the sliders to your monthly usage.
$744.00 a year
Overview
Claude Haiku 4.5 is Anthropic's fastest model, which Anthropic describes as having near-frontier intelligence. The zurelay Claude Haiku 4.5 API serves it through OpenAI-compatible and Anthropic-compatible endpoints at 62% below Anthropic's list price. It fits chat, sub-agents and high-volume work where speed and cost per call matter most.
Claude Haiku 4.5 was released on October 15, 2025 and is still Anthropic's current Haiku model. Anthropic describes it as the fastest model with near-frontier intelligence and rates it the fastest in its current lineup. At launch, Anthropic said it matched Claude Sonnet 4 on coding at one-third the cost and more than twice the speed. On zurelay, the Claude Haiku 4.5 API uses the model ID claude-haiku-4-5 and costs $0.38 per 1M input tokens and $1.90 per 1M output tokens.
The specs: a 200K-token context window, up to 64K output tokens per request, text and image input, and text output. Its reliable knowledge cutoff is February 2025, with training data through July 2025. Haiku 4.5 supports extended thinking: set thinking type enabled with a budget_tokens value of at least 1,024 and below max_tokens. It has no effort setting and no interleaved thinking between tool calls. It tracks its remaining context window on its own, which Anthropic calls context awareness.
Anthropic reported 73.3% on SWE-bench Verified at launch and said Haiku 4.5 beats Claude Sonnet 4 at some tasks, such as using computers. It targets real-time, low-latency work like chat assistants, customer service agents and pair programming. Anthropic also pitches it as a sub-agent: a larger Claude model breaks a problem into steps, then a team of Haiku 4.5 instances completes the subtasks in parallel.
Haiku 4.5 is active, with retirement no sooner than October 15, 2026, and Anthropic has not announced a deprecation as of September 30, 2026. It has the lowest list price in Anthropic's current lineup. Prompt caching needs at least 4,096 tokens in the cached prompt, more than current Sonnet and Fable models require, so cache long, stable system prompts and documents.
Strengths
Anthropic rates Haiku 4.5 its fastest current model, a good match for replies a user is waiting on.
It has the lowest list price among Anthropic's current models, and zurelay takes a further 62% off.
Anthropic reported 73.3% on SWE-bench Verified at launch, with coding on par with Claude Sonnet 4.
Fast and cheap enough to run many copies in parallel under a larger Claude model that plans and reviews.
Use cases
Real-time chat and customer service agents where a user waits on every reply.
Parallel workers that search, read, extract or edit while a larger Claude model plans the task and checks the results.
Quick completions, explanations and small fixes inside an editor, where latency matters more than depth.
Tagging, routing, summarizing and pulling structured fields from large batches of text or images.
Get started
No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.
Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.
Change the base URL. Everything else in your code stays the same.
https://api.zurelay.com/v1Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.
claude-haiku-4-5Works with the tools you already use
Compare
$0.80 / $4.00 per 1M tokens
Pick Claude Sonnet 5.5 when you need more reasoning, adaptive thinking or a 1M-token context window, at a higher list price.
Claude Sonnet 5.5 API$1.40 / $7.00 per 1M tokens
Pick Claude Opus 5.5 to plan and review complex work that Haiku 4.5 sub-agents then carry out.
Claude Opus 5.5 API$0.025 / $0.125 per 1M tokens
Try GPT-6 Luna on the same zurelay key for high-volume text jobs, and compare the two on your own evals.
GPT-6 Luna APIFAQ
Anything else, write to sales@zurelay.com and an engineer will answer.
zurelay charges $0.38 per 1M input tokens and $1.90 per 1M output tokens, with cached input at $0.038. Anthropic's list price is $1.00 input and $5.00 output, so you save 62%. The rate is the same at every prompt length.
Yes. Claude Haiku 4.5 already has the lowest list price in Anthropic's current lineup, and zurelay serves claude-haiku-4-5 at 62% below that. It is the same model, paid from prepaid credit with no subscription. You keep your code and change the base URL and key.
Create a zurelay API key, set the base URL to https://api.zurelay.com/v1 and set the model to claude-haiku-4-5. Chat completions, streaming and tool calls use the request shapes you know from OpenAI, and temperature and top_p work as usual.
Yes. Set the Anthropic SDK's base URL, or ANTHROPIC_BASE_URL for Claude Code, to https://api.zurelay.com with no /v1, because the client adds /v1/messages itself. Authenticate with your zurelay key and choose claude-haiku-4-5 as the model.
Claude Haiku 4.5 has a 200K-token context window, roughly 150k words, and returns up to 64K output tokens per request. For larger inputs, use a model with a 1M-token window, such as Claude Sonnet 5.5.
Yes, in extended mode. In the Messages format, send thinking type enabled with a budget_tokens value of at least 1,024 and below max_tokens. Haiku 4.5 does not support adaptive thinking, the effort setting, or interleaved thinking between tool calls.
Fast, frequent calls: chat assistants, customer support, pair programming, sub-agents and bulk extraction. Anthropic reported coding on par with Claude Sonnet 4 at launch. For hard multi-step reasoning, a Sonnet or Opus model is a better fit.
Yes. Streaming works on chat completions and on the Messages API; on chat completions, set stream_options.include_usage to get token counts at the end of the stream. Each key can have its own monthly budget and requests-per-minute cap. If a route fails, zurelay retries on another route, and failed requests are never billed.
Yes. Your requests run on Claude Haiku 4.5 itself, and zurelay never swaps in a different or smaller model. Output is sampled, so exact wording varies between calls on any API. zurelay is an independent service and is not affiliated with or endorsed by Anthropic.
Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.