Sampling parameters still work
The newest Opus that accepts temperature, top_p and top_k, for pipelines that tune output variety with parameters instead of prompts.
A Claude Opus 4.6 API key at 65% below list, for workloads built on Opus 4.6.
claude-opus-4-6from openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key="YOUR_ZURELAY_KEY",)stream = client.chat.completions.create( model="claude-opus-4-6", messages=[{"role": "user", "content": "Explain quicksort in two sentences."}], stream=True,)for chunk in stream: if chunk.choices: print(chunk.choices[0].delta.content or "", end="")Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK, and the Anthropic SDK and Claude Code.
Pricing
Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.
| Rate | zurelay | Anthropic list | You save |
|---|---|---|---|
Input per 1M tokens | $1.75 | 65%off | |
Output per 1M tokens | $8.75 | 65%off | |
Cached input per 1M tokens | $0.175 | — | — |
One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.
Move the sliders to your monthly usage.
$3,900 a year
Overview
Claude Opus 4.6 is the February 2026 Opus model that brought adaptive thinking and the max effort level to the Opus line. The zurelay Claude Opus 4.6 API serves the same model through OpenAI-compatible and Anthropic-compatible endpoints at 65% below Anthropic's list price. It is the newest Opus that still accepts temperature and budget_tokens, and it uses the tokenizer from before Opus 4.7.
Claude Opus 4.6 was released by Anthropic on February 5, 2026. At launch Anthropic highlighted more careful planning, longer agentic runs, more reliable work in large codebases, and stronger code review and debugging than Opus 4.5. On zurelay, the Claude Opus 4.6 API uses the model ID claude-opus-4-6 and costs $1.75 per 1M input tokens and $8.75 per 1M output tokens, compared with Anthropic's list price of $5.00 and $25.00.
Opus 4.6 introduced adaptive thinking to the Opus line, where the model decides when deeper reasoning helps, along with a max effort level above high. Effort runs low, medium, high and max, with high as the default; xhigh only arrived with Opus 4.7. Thinking stays off unless you enable it. The older budget_tokens mode is deprecated on this model but still works, which makes Opus 4.6 a bridge for code that has not moved to adaptive thinking yet.
The specs: a 1M-token context window, up to 128K output tokens, text and image input, and text output. The 1M window launched as a beta and now runs at Anthropic's standard pricing. Images top out at 1,568 pixels on the long edge. Knowledge is most reliable through May 2025, with training data through August 2025. The minimum cacheable prompt is 4,096 tokens, higher than on later Opus models, so short prompts will not cache.
Why still run Opus 4.6? It is the newest Opus that accepts temperature, top_p and top_k, which all return an error from Opus 4.7 on. It uses the older tokenizer, while Opus 4.7 and later can use up to about 35% more tokens for the same text. On the Messages API it returns summarized thinking text by default, where later models return it empty unless you opt in. Anthropic's current retirement date for it is not sooner than February 5, 2027.
Strengths
The newest Opus that accepts temperature, top_p and top_k, for pipelines that tune output variety with parameters instead of prompts.
It predates the Opus 4.7 tokenizer, which can use up to about 35% more tokens on the same input.
Anthropic cited more careful planning, longer agentic runs and more reliable work in big repositories than Opus 4.5.
Use adaptive thinking with effort, or keep existing budget_tokens code running while you migrate.
Use cases
Keep extraction, generation or evaluation jobs that depend on temperature or top_p running without rewrites.
Review changes and trace bugs across large codebases with adaptive thinking turned on.
Load contracts, filings or research sets into the 1M-token window and ask questions across all of them.
Keep a fixed Opus 4.6 baseline to compare newer Opus models against on your own test sets.
Get started
No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.
Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.
Change the base URL. Everything else in your code stays the same.
https://api.zurelay.com/v1Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.
claude-opus-4-6Works with the tools you already use
Compare
$1.78 / $8.93 per 1M tokens
Pick Claude Opus 4.7 for high-resolution vision and xhigh effort, once your code no longer sends sampling parameters or budget_tokens.
Claude Opus 4.7 API$1.75 / $8.75 per 1M tokens
Pick Claude Opus 4.8, the final Opus 4 release, for stronger bug finding and more honest reporting of uncertainty at the same list price.
Claude Opus 4.8 API$1.40 / $7.00 per 1M tokens
Pick Claude Opus 5.5, Anthropic's current Opus at a lower list price, for new projects.
Claude Opus 5.5 APIFAQ
Anything else, write to sales@zurelay.com and an engineer will answer.
zurelay charges $1.75 per 1M input tokens and $8.75 per 1M output tokens, with cached input at $0.175. Anthropic's list price for Claude Opus 4.6 is $5.00 input and $25.00 output, so you save 65%. The rate is the same at every prompt length, and you pay from prepaid credit with no subscription.
Yes. zurelay serves claude-opus-4-6 at 65% below Anthropic's list price. It is the same model, not a smaller substitute. You keep your code and change the base URL and API key.
Sign up at zurelay, add prepaid credit and create a key in the dashboard. Set the OpenAI SDK base URL to https://api.zurelay.com/v1 and the model to claude-opus-4-6. The same key works for every other model in the catalog, and you can give it a monthly budget and a requests-per-minute cap.
Yes. Set the Anthropic SDK base URL, or ANTHROPIC_BASE_URL for Claude Code, to https://api.zurelay.com with no /v1, because the client adds /v1/messages itself. Authenticate with your zurelay key and select claude-opus-4-6. Prefilled assistant turns return an error on Opus 4.6, so steer output format with instructions or structured outputs instead.
Claude Opus 4.6 has a 1M-token context window and returns up to 128K output tokens per request. On its tokenizer, 1M tokens fits about 750k English words. Stream long responses so the connection stays open while tokens arrive.
Use it when your code relies on temperature, top_p, top_k or budget_tokens thinking, which later Opus models reject, or when you need behavior that matches earlier Opus 4.6 results. It also needs fewer tokens than Opus 4.7 and later for the same text. For new projects, a newer Opus is the better starting point.
No. As of September 2026, Anthropic lists Claude Opus 4.6 as an active legacy model, not deprecated, with retirement not sooner than February 5, 2027. Anthropic gives at least 60 days' notice before retiring a publicly released model. Testing a newer Opus now means a future retirement will not catch you out.
Yes. Streaming is supported, and on chat completions you can set stream_options.include_usage to get token counts in the final chunk. Each key can carry its own monthly budget and requests-per-minute cap. If a route fails, zurelay retries the request on another route, and failed requests are never billed.
Yes. Requests to claude-opus-4-6 run on Claude Opus 4.6 itself, not a smaller or different model. Outputs are sampled, so exact wording can differ between calls on any API, including Anthropic's. zurelay is an independent service and is not affiliated with or endorsed by Anthropic.
Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.