Hard software engineering
Anthropic reported its biggest gains over Opus 4.6 on the most difficult coding tasks, the kind that used to need close supervision.
Claude Opus 4.7 API pricing at 64% below list, for apps tuned on Opus 4.7.
claude-opus-4-7from openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key="YOUR_ZURELAY_KEY",)stream = client.chat.completions.create( model="claude-opus-4-7", messages=[{"role": "user", "content": "Explain quicksort in two sentences."}], stream=True,)for chunk in stream: if chunk.choices: print(chunk.choices[0].delta.content or "", end="")Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK, and the Anthropic SDK and Claude Code.
Pricing
Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.
| Rate | zurelay | Anthropic list | You save |
|---|---|---|---|
Input per 1M tokens | $1.78 | 64%off | |
Output per 1M tokens | $8.93 | 64%off | |
Cached input per 1M tokens | $0.178 | — | — |
One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.
Move the sliders to your monthly usage.
$3,860 a year
Overview
Claude Opus 4.7 is the April 2026 Opus model that introduced xhigh effort, high-resolution vision and a new tokenizer. The zurelay Claude Opus 4.7 API serves the same model through OpenAI-compatible and Anthropic-compatible endpoints, priced at 64% below Anthropic's list price. It suits teams whose prompts, evals or agents are tuned to Opus 4.7 and should not move yet.
Claude Opus 4.7 was released by Anthropic on April 16, 2026 as the successor to Opus 4.6, at the same list price. Anthropic described it as a clear step up from Opus 4.6 in advanced software engineering, with the biggest gains on the hardest tasks, and said it handles long-running work with rigor and finds ways to check its own output. On zurelay, Claude Opus 4.7 API pricing is $1.78 per 1M input tokens and $8.93 per 1M output tokens, compared with Anthropic's list price of $5.00 and $25.00.
Several things arrived with Opus 4.7 and carried into later Opus models. It was the first Claude model with high-resolution image support, taking images up to 2,576 pixels on the long edge, up from 1,568, with pointing coordinates that map 1:1 to image pixels. It added the xhigh effort level between high and max, which Anthropic recommends as the starting point for coding and agentic work. It also introduced a new tokenizer that can use roughly 1.0 to 1.35 times as many tokens as Opus 4.6 for the same text.
The specs: a 1M-token context window, up to 128K output tokens, text and image input, and text output. Reliable knowledge runs through January 2026. Thinking is adaptive only and off unless you enable it. The older budget_tokens mode and non-default temperature, top_p and top_k values return an error. The minimum cacheable prompt is 2,048 tokens. Opus 4.7 also shipped real-time cybersecurity safeguards that can decline prohibited or high-risk security requests.
Opus 4.8 accepts exactly the same requests, so moving up is a model ID change. Keep Opus 4.7 when agents, prompts or evals are validated on it and you want pinned behavior until you re-test. Anthropic's current retirement date for it is not sooner than April 16, 2027.
Strengths
Anthropic reported its biggest gains over Opus 4.6 on the most difficult coding tasks, the kind that used to need close supervision.
Reads screenshots, documents and charts at up to 2,576 pixels on the long edge, with coordinates that map 1:1 to image pixels.
It respects low and medium effort strictly and scopes work to what was asked, which suits tuned pipelines and structured extraction.
Five effort levels, including the xhigh level it introduced, let you set cost and depth per route.
Use cases
Keep an agent that passed your evals on Opus 4.7 running unchanged while you test newer Opus models on the same key.
Extract text, numbers and layout from dense screenshots, scanned pages and charts at full resolution.
Turn documents into JSON with predictable scope, using low or medium effort to keep simple calls cheap.
Hand off multi-step fixes and features that need rigor, and let the model check its output before it reports back.
Get started
No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.
Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.
Change the base URL. Everything else in your code stays the same.
https://api.zurelay.com/v1Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.
claude-opus-4-7Works with the tools you already use
Compare
$1.75 / $8.75 per 1M tokens
Pick Claude Opus 4.8 at the same list price: it takes the same requests, catches more flaws in its own code and uses a smaller tool-use system prompt.
Claude Opus 4.8 API$1.75 / $8.75 per 1M tokens
Pick Claude Opus 4.6 if you still need temperature, top_p, top_k or budget_tokens thinking, or its older tokenizer, which uses fewer tokens for the same text.
Claude Opus 4.6 API$1.75 / $8.75 per 1M tokens
Pick Claude Opus 5 at the same list price for stronger results on hard coding and research, with thinking on by default.
Claude Opus 5 APIFAQ
Anything else, write to sales@zurelay.com and an engineer will answer.
zurelay charges $1.78 per 1M input tokens and $8.93 per 1M output tokens, with cached input at $0.178. Anthropic's list price for Claude Opus 4.7 is $5.00 input and $25.00 output, so you save 64%. The rate is the same at every prompt length, and you pay from prepaid credit with no subscription.
Yes. zurelay serves claude-opus-4-7 at 64% below Anthropic's list price. It is the same model, not a smaller substitute. You keep your code and change the base URL and API key.
Create a zurelay API key, set the base URL to https://api.zurelay.com/v1 and set the model to claude-opus-4-7. Chat completions, streaming and tool calls use the request shapes you already know. Leave out temperature, top_p and top_k: Opus 4.7 rejects non-default sampling values, and some frameworks set them for you.
Yes. Set the Anthropic SDK base URL, or ANTHROPIC_BASE_URL for Claude Code, to https://api.zurelay.com with no /v1, because the client adds /v1/messages itself. Authenticate with your zurelay key and select claude-opus-4-7. For thinking, send the adaptive thinking type; budget_tokens returns an error on this model.
Claude Opus 4.7 has a 1M-token context window and returns up to 128K output tokens per request. On its tokenizer, 1M tokens is roughly 555k English words. At xhigh or max effort, give it a large output limit, starting around 64K tokens, and stream the response.
Hard software engineering and long-running tasks that need rigor, plus vision work on dense screenshots, charts and documents. It follows the scope of a prompt closely, which suits pipelines with carefully tuned instructions. For new projects, Opus 4.8 takes the same requests and is the stronger model.
Opus 4.7 introduced a new tokenizer, and Anthropic says the same text can map to roughly 1.0 to 1.35 times as many tokens, depending on content. Full-resolution images can also use up to about 3 times more image tokens, up to 4,784 per image. Re-check output limits and cost estimates when you switch, and downsample images you do not need at full detail.
Yes. Streaming is supported, and on chat completions you can set stream_options.include_usage to get token counts in the final chunk. Each key can carry its own monthly budget and requests-per-minute cap. If a route fails, zurelay retries the request on another route, and failed requests are never billed.
Yes. Requests to claude-opus-4-7 run on Claude Opus 4.7 itself, not a smaller or different model. Outputs are sampled, so exact wording can differ between calls on any API, including Anthropic's. zurelay is an independent service and is not affiliated with or endorsed by Anthropic.
Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.