Anthropic
Operational· 100% uptime, 7 days

Claude Opus 5 API

The Claude Opus 5 API for hard coding and agent work, at 65% below list price.

Price per 1M tokens65%off
Input
$1.75$5.00
Output
$8.75$25.00
Cached input
$0.175
Model IDclaude-opus-5
chat.py
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key="YOUR_ZURELAY_KEY",
)
stream = client.chat.completions.create(
model="claude-opus-5",
messages=[{"role": "user", "content": "Explain quicksort in two sentences."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")

Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK, and the Anthropic SDK and Claude Code.

Context window
1M tokens
Max output
128K tokens
Input
Text, Image
Output
Text
Released
Jul 24, 2026
Uptime, 7 days
100%

Pricing

Claude Opus 5 API pricing. A fraction of the list price.

Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.

RatezurelayAnthropic listYou save
Input
per 1M tokens
$1.75$5.0065%off
Output
per 1M tokens
$8.75$25.0065%off
Cached input
per 1M tokens
$0.175——

One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.

Savings calculator

Move the sliders to your monthly usage.

Input tokens per month50M
Output tokens per month10M
zurelay
$175.00
Anthropic list price
$500.00
You save every month$325.00

$3,900 a year

Overview

What is Claude Opus 5?

Claude Opus 5 is Anthropic's July 2026 Opus model for complex agentic coding and enterprise work. The zurelay Claude Opus 5 API serves the same model through OpenAI-compatible and Anthropic-compatible endpoints at 65% below Anthropic's list price. It suits teams that built prompts and agent harnesses on Opus 5 and want to keep that behavior.

Claude Opus 5 was released by Anthropic on July 24, 2026 as the successor to Claude Opus 4.8, at the same list price. Anthropic built it for complex agentic coding and enterprise work, with particular strength on long-horizon agentic tasks. On zurelay, the Claude Opus 5 API uses the model ID claude-opus-5 and costs $1.75 per 1M input tokens and $8.75 per 1M output tokens, compared with Anthropic's list price of $5.00 and $25.00.

Anthropic says Opus 5 performs much better than Opus 4.8 at the same cost, and reports leading results on coding and knowledge-work evaluations such as GDPval-AA. Its documentation points to the hard end of coding: multi-file features, larger refactors and end-to-end feature work, where it finishes the job instead of leaving stubs. It reviews code with high precision and recall, builds multi-sheet spreadsheets with real formulas and well-structured slide decks, and Anthropic calls it a meaningful step up for scientific research.

The specs: a 1M-token context window as both default and maximum, up to 128K output tokens per request, text and image input, and text output. Reliable knowledge runs through May 2026. Thinking is adaptive and on by default, a change from Opus 4.8, which answers without thinking unless asked. You can still turn thinking off at effort high or below. Effort runs from low to max with high as the default, and Anthropic notes that low and medium give strong quality at a fraction of the tokens. The minimum cacheable prompt is 512 tokens, down from 1,024 on Opus 4.8.

A few behaviors differ from earlier Opus versions. Opus 5 writes longer answers by default, checks its own work without being told, and hands work to subagents more readily. Remove old 'double-check your answer' instructions, ask for concise output where length matters, and cap subagent use on cost-sensitive routes. Opus 5 also runs safety classifiers that can decline a request, so handle a refusal in your code. Anthropic now lists it as a legacy model, with retirement not sooner than July 24, 2027.

Strengths

Where Claude Opus 5 shines. And what teams build with it.

01

Hard agentic coding

Strongest on multi-file features, larger refactors and end-to-end feature work. It completes full tasks instead of leaving stubs or placeholders.

02

Code review that holds at low effort

Finds real bugs at a high rate per pass with few false positives, and stays accurate at lower effort, so a quick review on every change is practical.

03

Low effort that still works

Anthropic says low and medium effort give strong quality at a fraction of the tokens and latency, which makes effort your main cost lever.

04

Consistent across 1M tokens

Instruction following, tool calling and reasoning stay consistent across the full 1M-token window, according to Anthropic's prompting guide.

Use cases

  • Autonomous coding agents

    Give it the complete task spec for a feature or refactor, let it run with tool calls, and stream progress to your IDE, CLI or CI job.

  • Pull request review

    Run a fast low-effort review on every pull request and a deeper high-effort pass before release.

  • Spreadsheets, decks and reports

    Generate multi-sheet spreadsheets with non-trivial formulas and structured slide decks through your own file tools, following a template you provide.

  • Charts, diagrams and UI screenshots

    Read charts, documents and diagrams, or rebuild a UI from a screenshot. Anthropic says tools to crop and check its work help more than extra thinking.

Get started

Call Claude Opus 5 in three steps.

No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.

  1. 1

    Create a key

    Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.

  2. 2

    Point your SDK at zurelay

    Change the base URL. Everything else in your code stays the same.

    https://api.zurelay.com/v1
  3. 3

    Use claude-opus-5

    Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.

    claude-opus-5

FAQ

Claude Opus 5 API questions.

Anything else, write to sales@zurelay.com and an engineer will answer.

How much does the Claude Opus 5 API cost on zurelay?

zurelay charges $1.75 per 1M input tokens and $8.75 per 1M output tokens, with cached input at $0.175. Anthropic's list price for Claude Opus 5 is $5.00 input and $25.00 output, so you save 65%. The rate is the same at every prompt length, and you pay from prepaid credit with no subscription.

Is there a cheaper Claude Opus 5 API than Anthropic's?

Yes. zurelay serves claude-opus-5 at 65% below Anthropic's list price. It is the same model, not a smaller substitute. You keep your code and change the base URL, the API key and, if needed, the model ID.

How do I call the Claude Opus 5 API with the OpenAI SDK?

Create a zurelay API key, set the base URL to https://api.zurelay.com/v1 and set the model to claude-opus-5. Chat completions, streaming and tool calls use the request shapes you already know. Leave out temperature, top_p and top_k, because Opus models from 4.7 on reject non-default sampling values.

Can I use Claude Opus 5 with the Anthropic SDK or Claude Code?

Yes. Set the Anthropic SDK base URL, or ANTHROPIC_BASE_URL for Claude Code, to https://api.zurelay.com with no /v1, because the client adds /v1/messages itself. Authenticate with your zurelay key and select claude-opus-5, for example with Claude Code's --model flag. Opus 5 delegates to subagents readily, and Claude Code 2.1.217 or later can cap subagent depth and concurrency.

What is the Claude Opus 5 context window?

Claude Opus 5 has a 1M-token context window, which is both the default and the maximum, and returns up to 128K output tokens per request. On its tokenizer, 1M tokens is roughly 555k English words. Stream long responses so the connection stays open while tokens arrive.

What is Claude Opus 5 best at?

Anthropic built it for complex agentic coding and enterprise work. It does best on hard, multi-step jobs: multi-file features, large refactors, code review, spreadsheet and slide work, and reading charts and diagrams. On easy single-turn edits the gap over older models is smaller, so test it where your workload is hardest.

How is Claude Opus 5 different from Claude Opus 4.8?

Opus 5 has the same list price as Opus 4.8 and, per Anthropic, performs much better on coding, knowledge work and research. Thinking is on by default, and turning it off only works at effort high or below. It also writes longer answers, verifies its own work unprompted and uses subagents more often, so prompts tuned for 4.8 may need a conciseness instruction and fewer verification steps.

Does the zurelay Claude Opus 5 API support streaming, and what are the limits?

Yes. Streaming is supported, and on chat completions you can set stream_options.include_usage to get token counts in the final chunk. Each key can carry its own monthly budget and requests-per-minute cap. If a route fails, zurelay retries the request on another route, and failed requests are never billed.

Is Claude Opus 5 on zurelay the same model as on Anthropic's API?

Yes. Requests to claude-opus-5 run on Claude Opus 5 itself, not a smaller or different model. Outputs are sampled, so exact wording can differ between calls on any API, including Anthropic's. zurelay is an independent service and is not affiliated with or endorsed by Anthropic.

More models on zurelay

All models

Stop paying list price.
Start saving today.

Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.