Tier above Opus
Anthropic places Mythos-class models, including Fable, above its Opus class in capability. Use it for the hardest reasoning, planning and review steps in your stack.
The Claude Fable 5 API, Anthropic's first Fable model, at 65% below list price.
claude-fable-5from openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key="YOUR_ZURELAY_KEY",)stream = client.chat.completions.create( model="claude-fable-5", messages=[{"role": "user", "content": "Explain quicksort in two sentences."}], stream=True,)for chunk in stream: if chunk.choices: print(chunk.choices[0].delta.content or "", end="")Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK, and the Anthropic SDK and Claude Code.
Pricing
Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.
| Rate | zurelay | Anthropic list | You save |
|---|---|---|---|
Input per 1M tokens | $3.50 | 65%off | |
Output per 1M tokens | $17.50 | 65%off | |
Cached input per 1M tokens | $0.35 | — | — |
One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.
Move the sliders to your monthly usage.
$7,800 a year
Overview
Claude Fable 5 is Anthropic's first Fable model, part of the Mythos-class tier that Anthropic places above Opus in capability. The zurelay Claude Fable 5 API serves it through OpenAI-compatible and Anthropic-compatible endpoints at 65% below Anthropic's list price. It suits teams that tuned prompts and evals on Fable 5 and want to keep them running for less.
Claude Fable 5 launched on June 9, 2026 as the first model in Anthropic's Fable line. Anthropic calls it a Mythos-class model made safe for general use, and says Mythos-class models sit above its Opus class in capability. It is built for demanding reasoning and long-horizon agentic work. On zurelay, the Claude Fable 5 API uses the model ID claude-fable-5 and costs $3.50 per 1M input tokens and $17.50 per 1M output tokens.
The specs: a 1M-token context window by default, up to 128K output tokens per request, text and image input, and text output. Its reliable knowledge cutoff is January 2026. Adaptive thinking is always on and cannot be disabled, and the raw chain of thought is never returned. You steer depth, latency and cost with effort: low, medium, high (the default), xhigh or max. Anthropic says Fable 5 at lower effort still performs well and often beats earlier models running at xhigh.
Claude Fable 5.1 replaced it as the current Fable model on September 1, 2026, and Anthropic now lists Fable 5 as a legacy model. Legacy means it is still served but no longer updated, with retirement no sooner than June 9, 2027. There are practical reasons to stay on it: Fable 5 accepts forced tool choice, which Fable 5.1 rejects, and in agent loops it tends to batch several tool calls per turn where Fable 5.1 may issue one at a time.
A few API rules apply. Sampling parameters (temperature, top_p, top_k) must stay at their defaults, and prefilling the assistant reply returns an error. Anthropic runs safety classifiers on Fable 5 that can decline some requests in areas such as cybersecurity and biology, and says more than 95% of Fable sessions involve no fallback at all. Prompt caching works from a 512-token minimum prompt.
Strengths
Anthropic places Mythos-class models, including Fable, above its Opus class in capability. Use it for the hardest reasoning, planning and review steps in your stack.
Built for long-running agentic work. Anthropic says Fable 5 stays focused across millions of tokens in long-running tasks.
Adaptive thinking is always on, and effort sets how hard it works. Anthropic says lower effort on Fable 5 often beats earlier models at xhigh, so test low and medium first.
As a legacy model, Fable 5 gets no further updates, so prompts and evals tuned on it keep behaving the same until you choose to move.
Use cases
Keep production prompts and evals tuned on Fable 5 while cutting the per-token bill, then move to Fable 5.1 on your own schedule.
Workflows that force a specific tool call with tool_choice keep working on Fable 5. Fable 5.1 rejects forced tool use.
Load a big repository or document collection into the 1M-token window and ask for analysis, plans or migrations across all of it.
Anthropic names software engineering, knowledge work, vision and life sciences research as Fable 5's strongest areas.
Get started
No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.
Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.
Change the base URL. Everything else in your code stays the same.
https://api.zurelay.com/v1Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.
claude-fable-5Works with the tools you already use
Compare
$3.50 / $17.50 per 1M tokens
Pick Claude Fable 5.1 for new work: it is the current Fable model, stronger on long agentic runs, at the same list price per token with cheaper cache reads.
Claude Fable 5.1 API$1.40 / $7.00 per 1M tokens
Pick Claude Opus 5.5 for most workloads at a lower price; Anthropic suggests starting there and moving up to Fable only when your evals fall short.
Claude Opus 5.5 API$0.80 / $4.00 per 1M tokens
Pick Claude Sonnet 5.5 for fast, well-scoped tasks where latency and cost matter more than peak reasoning.
Claude Sonnet 5.5 APIFAQ
Anything else, write to sales@zurelay.com and an engineer will answer.
zurelay charges $3.50 per 1M input tokens and $17.50 per 1M output tokens, with cached input at $0.35. Anthropic's list price is $10.00 input and $50.00 output, so you save 65%. The rate is the same at every prompt length, up to the full 1M-token window.
Yes. zurelay serves claude-fable-5 at 65% below Anthropic's list price, paid from prepaid credit with no subscription. It is the same model, not a smaller substitute. You keep your code and change the base URL and key.
Create a zurelay API key, set the base URL to https://api.zurelay.com/v1 and set the model to claude-fable-5. Chat completions, streaming and tool calls use the request shapes you know from OpenAI. Leave temperature and top_p unset, because Fable 5 rejects non-default sampling values.
Yes. zurelay also serves the Anthropic Messages API. Set the Anthropic SDK's base URL, or ANTHROPIC_BASE_URL for Claude Code, to https://api.zurelay.com with no /v1, because the client adds /v1/messages itself. Authenticate with your zurelay key and choose claude-fable-5 as the model.
For new work, use Fable 5.1. Anthropic reports gains in long agentic coding, research, document work, vision and computer use at the same list price per token. Stay on Fable 5 if your code forces tool choice, which Fable 5.1 rejects, or if you need results that match evals you already ran.
Claude Fable 5 has a 1M-token context window and returns up to 128K output tokens per request. Thinking counts toward the output limit, so set a large max_tokens at high effort and stream long responses.
No. Anthropic lists Claude Fable 5 as a legacy model: still active, no longer updated, with retirement no sooner than June 9, 2027. Claude Fable 5.1 is its successor.
Yes. Streaming works on chat completions and on the Messages API; on chat completions, set stream_options.include_usage to get token counts at the end of the stream. Each key can have its own monthly budget and requests-per-minute cap. If a route fails, zurelay retries on another route, and failed requests are never billed.
Yes. Your requests run on Claude Fable 5 itself, and zurelay never swaps in a different or smaller model. Output is sampled, so exact wording varies between calls on any API. zurelay is an independent service and is not affiliated with or endorsed by Anthropic.
Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.