Anthropic
Operational· 100% uptime, 7 days

Claude Opus 4.8 API

The Claude Opus 4.8 API at 65% below Anthropic's list price, on one key.

Price per 1M tokens65%off
Input
$1.75$5.00
Output
$8.75$25.00
Cached input
$0.175
Model IDclaude-opus-4-8
chat.py
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key="YOUR_ZURELAY_KEY",
)
stream = client.chat.completions.create(
model="claude-opus-4-8",
messages=[{"role": "user", "content": "Explain quicksort in two sentences."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")

Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK, and the Anthropic SDK and Claude Code.

Context window
1M tokens
Max output
128K tokens
Input
Text, Image
Output
Text
Released
May 28, 2026
Uptime, 7 days
100%

Pricing

Claude Opus 4.8 API pricing. A fraction of the list price.

Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.

RatezurelayAnthropic listYou save
Input
per 1M tokens
$1.75$5.0065%off
Output
per 1M tokens
$8.75$25.0065%off
Cached input
per 1M tokens
$0.175——

One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.

Savings calculator

Move the sliders to your monthly usage.

Input tokens per month50M
Output tokens per month10M
zurelay
$175.00
Anthropic list price
$500.00
You save every month$325.00

$3,900 a year

Overview

What is Claude Opus 4.8?

Claude Opus 4.8 is the final Opus 4 model, released by Anthropic in May 2026 for agentic coding and knowledge work. The zurelay Claude Opus 4.8 API serves the same model through OpenAI-compatible and Anthropic-compatible endpoints at 65% below Anthropic's list price. Teams keep it for its pinned behavior, its Opus 4.7 request shape and thinking that stays off until you ask for it.

Claude Opus 4.8 was released by Anthropic on May 28, 2026, six weeks after Opus 4.7, at the same list price. Anthropic lists its particular strengths as long-horizon agentic work, knowledge work, vision and memory tasks. On zurelay, the Claude Opus 4.8 API uses the model ID claude-opus-4-8 and costs $1.75 per 1M input tokens and $8.75 per 1M output tokens, compared with Anthropic's list price of $5.00 and $25.00.

The headline change from Opus 4.7 is reliability. Anthropic says Opus 4.8 is around four times less likely to let flaws in code it wrote pass without comment, more likely to flag uncertainty about its own work, and less likely to make unsupported claims. Its prompting guide adds that it finds bugs with higher recall and precision than prior models in internal evals.

The specs: a 1M-token context window, up to 128K output tokens per request, text and image input, and text output. Reliable knowledge runs through January 2026. Thinking is adaptive and off unless you set it, and effort runs from low to max, including xhigh, with high as the default. Anthropic suggests starting at xhigh for coding and agentic work. The minimum cacheable prompt is 1,024 tokens. On Anthropic's Messages API, Opus 4.8 also added mid-conversation system messages, which update an agent's instructions without breaking the prompt cache.

Why pick Opus 4.8 over Opus 5? It takes exactly the same requests as Opus 4.7, so it is an easy step for 4.7 users. Without a thinking setting it answers directly, while Opus 5 thinks by default. It also keeps answers shorter and spawns fewer subagents than Opus 5, which helps when cost per request matters more than peak capability. Anthropic's current retirement date for Opus 4.8 is not sooner than May 28, 2027.

Strengths

Where Claude Opus 4.8 shines. And what teams build with it.

01

Honest about its own work

Anthropic says it flags uncertainty more often, makes fewer unsupported claims and is about four times less likely than Opus 4.7 to let flaws in its own code slip by.

02

Better bug finding

Higher recall and precision on code review than earlier models in Anthropic's evals. Ask it to report every finding and filter in a later step.

03

Long autonomous runs

Built for long agentic work such as complex refactors. Give it the full task spec up front and run it at high or xhigh effort.

04

Drop-in for Opus 4.7

Same request surface as Opus 4.7, so upgrading is a model ID change plus light prompt tuning.

Use cases

  • Autonomous coding and refactors

    Run long multi-step coding agents that plan, call tools and finish without constant correction.

  • Code review and QA

    Review pull requests and generated code, with the model flagging what it is unsure about instead of glossing over it.

  • Knowledge work

    Draft reports, analyses and structured extractions where a clear 'I am not sure' beats a confident guess.

  • Screenshots and documents

    Read high-resolution screenshots, charts and scanned pages. For screen-driving agents, Anthropic suggests 1080p screenshots as a good balance of performance and cost.

Get started

Call Claude Opus 4.8 in three steps.

No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.

  1. 1

    Create a key

    Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.

  2. 2

    Point your SDK at zurelay

    Change the base URL. Everything else in your code stays the same.

    https://api.zurelay.com/v1
  3. 3

    Use claude-opus-4-8

    Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.

    claude-opus-4-8

FAQ

Claude Opus 4.8 API questions.

Anything else, write to sales@zurelay.com and an engineer will answer.

How much does the Claude Opus 4.8 API cost on zurelay?

zurelay charges $1.75 per 1M input tokens and $8.75 per 1M output tokens, with cached input at $0.175. Anthropic's list price for Claude Opus 4.8 is $5.00 input and $25.00 output, so you save 65%. The rate is the same at every prompt length, and you pay from prepaid credit with no subscription.

Is there a cheaper Claude Opus 4.8 API than Anthropic's?

Yes. zurelay serves claude-opus-4-8 at 65% below Anthropic's list price. It is the same model, not a smaller substitute. You keep your code and change the base URL and API key.

How do I call the Claude Opus 4.8 API with the OpenAI SDK?

Create a zurelay API key, set the base URL to https://api.zurelay.com/v1 and set the model to claude-opus-4-8. Chat completions, streaming and tool calls use the request shapes you already know. Leave out temperature, top_p and top_k, because Opus 4.8 rejects non-default sampling values.

Can I use Claude Opus 4.8 with the Anthropic SDK or Claude Code?

Yes. Set the Anthropic SDK base URL, or ANTHROPIC_BASE_URL for Claude Code, to https://api.zurelay.com with no /v1, because the client adds /v1/messages itself. Authenticate with your zurelay key and select claude-opus-4-8. If you want thinking, send the adaptive thinking type; Opus 4.8 rejects the older budget_tokens form.

What is the Claude Opus 4.8 context window?

Claude Opus 4.8 has a 1M-token context window and returns up to 128K output tokens per request. On its tokenizer, 1M tokens is roughly 555k English words. At xhigh or max effort, Anthropic suggests a large output limit, starting around 64K tokens, and streaming the response.

What is Claude Opus 4.8 best at?

Long-horizon agentic work, knowledge work, vision and memory tasks, per Anthropic. It stands out on code review and on work where you want the model to say when it is unsure. It performs best when you give it the full task up front and run it at high or xhigh effort.

Should I use Claude Opus 4.8 or Claude Opus 5?

Both have the same list price. Opus 5 is stronger on hard coding, research and document work, but thinks by default, writes longer answers and delegates to subagents more. Opus 4.8 answers without thinking unless asked and matches the Opus 4.7 request shape, so it fits routes where you want shorter, cheaper replies or pinned behavior.

Does the zurelay Claude Opus 4.8 API support streaming, and what are the limits?

Yes. Streaming is supported, and on chat completions you can set stream_options.include_usage to get token counts in the final chunk. Each key can carry its own monthly budget and requests-per-minute cap. If a route fails, zurelay retries the request on another route, and failed requests are never billed.

Is Claude Opus 4.8 on zurelay the same model as on Anthropic's API?

Yes. Requests to claude-opus-4-8 run on Claude Opus 4.8 itself, not a smaller or different model. Outputs are sampled, so exact wording can differ between calls on any API, including Anthropic's. zurelay is an independent service and is not affiliated with or endorsed by Anthropic.

More models on zurelay

All models

Stop paying list price.
Start saving today.

Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.