Honest about its own work
Anthropic says it flags uncertainty more often, makes fewer unsupported claims and is about four times less likely than Opus 4.7 to let flaws in its own code slip by.
The Claude Opus 4.8 API at 65% below Anthropic's list price, on one key.
claude-opus-4-8from openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key="YOUR_ZURELAY_KEY",)stream = client.chat.completions.create( model="claude-opus-4-8", messages=[{"role": "user", "content": "Explain quicksort in two sentences."}], stream=True,)for chunk in stream: if chunk.choices: print(chunk.choices[0].delta.content or "", end="")Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK, and the Anthropic SDK and Claude Code.
Pricing
Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.
| Rate | zurelay | Anthropic list | You save |
|---|---|---|---|
Input per 1M tokens | $1.75 | 65%off | |
Output per 1M tokens | $8.75 | 65%off | |
Cached input per 1M tokens | $0.175 | — | — |
One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.
Move the sliders to your monthly usage.
$3,900 a year
Overview
Claude Opus 4.8 is the final Opus 4 model, released by Anthropic in May 2026 for agentic coding and knowledge work. The zurelay Claude Opus 4.8 API serves the same model through OpenAI-compatible and Anthropic-compatible endpoints at 65% below Anthropic's list price. Teams keep it for its pinned behavior, its Opus 4.7 request shape and thinking that stays off until you ask for it.
Claude Opus 4.8 was released by Anthropic on May 28, 2026, six weeks after Opus 4.7, at the same list price. Anthropic lists its particular strengths as long-horizon agentic work, knowledge work, vision and memory tasks. On zurelay, the Claude Opus 4.8 API uses the model ID claude-opus-4-8 and costs $1.75 per 1M input tokens and $8.75 per 1M output tokens, compared with Anthropic's list price of $5.00 and $25.00.
The headline change from Opus 4.7 is reliability. Anthropic says Opus 4.8 is around four times less likely to let flaws in code it wrote pass without comment, more likely to flag uncertainty about its own work, and less likely to make unsupported claims. Its prompting guide adds that it finds bugs with higher recall and precision than prior models in internal evals.
The specs: a 1M-token context window, up to 128K output tokens per request, text and image input, and text output. Reliable knowledge runs through January 2026. Thinking is adaptive and off unless you set it, and effort runs from low to max, including xhigh, with high as the default. Anthropic suggests starting at xhigh for coding and agentic work. The minimum cacheable prompt is 1,024 tokens. On Anthropic's Messages API, Opus 4.8 also added mid-conversation system messages, which update an agent's instructions without breaking the prompt cache.
Why pick Opus 4.8 over Opus 5? It takes exactly the same requests as Opus 4.7, so it is an easy step for 4.7 users. Without a thinking setting it answers directly, while Opus 5 thinks by default. It also keeps answers shorter and spawns fewer subagents than Opus 5, which helps when cost per request matters more than peak capability. Anthropic's current retirement date for Opus 4.8 is not sooner than May 28, 2027.
Strengths
Anthropic says it flags uncertainty more often, makes fewer unsupported claims and is about four times less likely than Opus 4.7 to let flaws in its own code slip by.
Higher recall and precision on code review than earlier models in Anthropic's evals. Ask it to report every finding and filter in a later step.
Built for long agentic work such as complex refactors. Give it the full task spec up front and run it at high or xhigh effort.
Same request surface as Opus 4.7, so upgrading is a model ID change plus light prompt tuning.
Use cases
Run long multi-step coding agents that plan, call tools and finish without constant correction.
Review pull requests and generated code, with the model flagging what it is unsure about instead of glossing over it.
Draft reports, analyses and structured extractions where a clear 'I am not sure' beats a confident guess.
Read high-resolution screenshots, charts and scanned pages. For screen-driving agents, Anthropic suggests 1080p screenshots as a good balance of performance and cost.
Get started
No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.
Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.
Change the base URL. Everything else in your code stays the same.
https://api.zurelay.com/v1Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.
claude-opus-4-8Works with the tools you already use
Compare
$1.75 / $8.75 per 1M tokens
Pick Claude Opus 5 at the same list price for stronger results on hard coding, research and document work, if you accept longer answers and thinking on by default.
Claude Opus 5 API$1.78 / $8.93 per 1M tokens
Stay on Claude Opus 4.7 only if your evals are pinned to its exact behavior; it takes the same requests but adds a larger tool-use system prompt to every request with tools.
Claude Opus 4.7 API$1.40 / $7.00 per 1M tokens
Pick Claude Opus 5.5, Anthropic's current Opus at a lower list price, for new projects that can run with thinking always on.
Claude Opus 5.5 APIFAQ
Anything else, write to sales@zurelay.com and an engineer will answer.
zurelay charges $1.75 per 1M input tokens and $8.75 per 1M output tokens, with cached input at $0.175. Anthropic's list price for Claude Opus 4.8 is $5.00 input and $25.00 output, so you save 65%. The rate is the same at every prompt length, and you pay from prepaid credit with no subscription.
Yes. zurelay serves claude-opus-4-8 at 65% below Anthropic's list price. It is the same model, not a smaller substitute. You keep your code and change the base URL and API key.
Create a zurelay API key, set the base URL to https://api.zurelay.com/v1 and set the model to claude-opus-4-8. Chat completions, streaming and tool calls use the request shapes you already know. Leave out temperature, top_p and top_k, because Opus 4.8 rejects non-default sampling values.
Yes. Set the Anthropic SDK base URL, or ANTHROPIC_BASE_URL for Claude Code, to https://api.zurelay.com with no /v1, because the client adds /v1/messages itself. Authenticate with your zurelay key and select claude-opus-4-8. If you want thinking, send the adaptive thinking type; Opus 4.8 rejects the older budget_tokens form.
Claude Opus 4.8 has a 1M-token context window and returns up to 128K output tokens per request. On its tokenizer, 1M tokens is roughly 555k English words. At xhigh or max effort, Anthropic suggests a large output limit, starting around 64K tokens, and streaming the response.
Long-horizon agentic work, knowledge work, vision and memory tasks, per Anthropic. It stands out on code review and on work where you want the model to say when it is unsure. It performs best when you give it the full task up front and run it at high or xhigh effort.
Both have the same list price. Opus 5 is stronger on hard coding, research and document work, but thinks by default, writes longer answers and delegates to subagents more. Opus 4.8 answers without thinking unless asked and matches the Opus 4.7 request shape, so it fits routes where you want shorter, cheaper replies or pinned behavior.
Yes. Streaming is supported, and on chat completions you can set stream_options.include_usage to get token counts in the final chunk. Each key can carry its own monthly budget and requests-per-minute cap. If a route fails, zurelay retries the request on another route, and failed requests are never billed.
Yes. Requests to claude-opus-4-8 run on Claude Opus 4.8 itself, not a smaller or different model. Outputs are sampled, so exact wording can differ between calls on any API, including Anthropic's. zurelay is an independent service and is not affiliated with or endorsed by Anthropic.
Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.