Every model you want. Up to 88% cheaper.
Anthropic, OpenAI, Google, xAI, Xiaomi, DeepSeek, Z.ai, Moonshot AI, Alibaba, Tencent, MiniMax and ByteDance models through one OpenAI-compatible API. Text and image models at a fraction of each lab’s list price, and video billed by the second.
Anthropic
Claude Opus 5.5
$1.40 / $7.00 per 1M tokens
Claude Opus 5.5 is Anthropic's current Opus model, built for long-running agentic coding and knowledge work. The zurelay Claude Opus 5.5 API serves it through an OpenAI-compatible or Anthropic-compatible endpoint at 65% below Anthropic's list price, for teams that want Opus-level results without Opus-level bills.
Claude Sonnet 5.5
$0.80 / $4.00 per 1M tokens
Claude Sonnet 5.5 is Anthropic's current Sonnet model, which Anthropic calls its best combination of speed and intelligence. The zurelay Claude Sonnet 5.5 API serves it through an OpenAI-compatible or Anthropic-compatible endpoint at 60% below list price, for teams running coding, agent and document work at volume.
Claude Fable 5.1
$3.50 / $17.50 per 1M tokens
Claude Fable 5.1 is Anthropic's current Fable model, from the tier above Opus, built for demanding reasoning and long-horizon agentic work. The zurelay Claude Fable 5.1 API serves it through OpenAI-compatible and Anthropic-compatible endpoints at 65% below Anthropic's list price, so your hardest jobs run for less.
Claude Fable 5
$3.50 / $17.50 per 1M tokens
Claude Fable 5 is Anthropic's first Fable model, part of the Mythos-class tier that Anthropic places above Opus in capability. The zurelay Claude Fable 5 API serves it through OpenAI-compatible and Anthropic-compatible endpoints at 65% below Anthropic's list price. It suits teams that tuned prompts and evals on Fable 5 and want to keep them running for less.
Claude Opus 5
$1.75 / $8.75 per 1M tokens
Claude Opus 5 is Anthropic's July 2026 Opus model for complex agentic coding and enterprise work. The zurelay Claude Opus 5 API serves the same model through OpenAI-compatible and Anthropic-compatible endpoints at 65% below Anthropic's list price. It suits teams that built prompts and agent harnesses on Opus 5 and want to keep that behavior.
Claude Opus 4.8
$1.75 / $8.75 per 1M tokens
Claude Opus 4.8 is the final Opus 4 model, released by Anthropic in May 2026 for agentic coding and knowledge work. The zurelay Claude Opus 4.8 API serves the same model through OpenAI-compatible and Anthropic-compatible endpoints at 65% below Anthropic's list price. Teams keep it for its pinned behavior, its Opus 4.7 request shape and thinking that stays off until you ask for it.
Claude Opus 4.7
$1.78 / $8.93 per 1M tokens
Claude Opus 4.7 is the April 2026 Opus model that introduced xhigh effort, high-resolution vision and a new tokenizer. The zurelay Claude Opus 4.7 API serves the same model through OpenAI-compatible and Anthropic-compatible endpoints, priced at 64% below Anthropic's list price. It suits teams whose prompts, evals or agents are tuned to Opus 4.7 and should not move yet.
Claude Opus 4.6
$1.75 / $8.75 per 1M tokens
Claude Opus 4.6 is the February 2026 Opus model that brought adaptive thinking and the max effort level to the Opus line. The zurelay Claude Opus 4.6 API serves the same model through OpenAI-compatible and Anthropic-compatible endpoints at 65% below Anthropic's list price. It is the newest Opus that still accepts temperature and budget_tokens, and it uses the tokenizer from before Opus 4.7.
Claude Haiku 4.5
$0.38 / $1.90 per 1M tokens
Claude Haiku 4.5 is Anthropic's fastest model, which Anthropic describes as having near-frontier intelligence. The zurelay Claude Haiku 4.5 API serves it through OpenAI-compatible and Anthropic-compatible endpoints at 62% below Anthropic's list price. It fits chat, sub-agents and high-volume work where speed and cost per call matter most.
OpenAI
GPT-6 Astra
$2.50 / $12.50 per 1M tokens
GPT-6 Astra is the top model in OpenAI's GPT-6 family, built for complex reasoning, coding, research and document work. The GPT-6 Astra API on zurelay runs the same model through one OpenAI-compatible endpoint, below OpenAI's list price. It is for teams whose hardest tasks justify the most capable model.
GPT-6.1 Sol
$0.50 / $2.50 per 1M tokens
GPT-6.1 Sol is OpenAI's newest Sol model, released on September 29, 2026. OpenAI says it nearly matches GPT-6 Astra on agentic coding, computer use and professional work at one-fifth of Astra's list price per token. The GPT-6.1 Sol API on zurelay serves it through one OpenAI-compatible endpoint, below OpenAI's list price.
GPT-6 Sol
$0.50 / $2.50 per 1M tokens
GPT-6 Sol is the middle model in OpenAI's GPT-6 family, built for complex coding and agentic workflows. The GPT-6 Sol API on zurelay runs the same model through one OpenAI-compatible endpoint, below OpenAI's list price. It suits teams that want Sol's full reasoning range, from none to max.
GPT-6 Luna
$0.025 / $0.125 per 1M tokens
GPT-6 Luna is the fastest and lowest-cost model in OpenAI's GPT-6 family, built for focused, high-volume tasks like summaries, extraction and quick answers. The GPT-6 Luna API on zurelay runs the same model through one OpenAI-compatible endpoint, below OpenAI's list price.
GPT-5.6 Sol
$1.14 / $5.70 per 1M tokens
GPT-5.6 Sol is the flagship of OpenAI's GPT-5.6 family, built for frontier reasoning and long-horizon agentic work. The GPT-5.6 Sol API on zurelay runs the same model through one OpenAI-compatible endpoint, below OpenAI's list price. You pay one flat rate per token, even on prompts over 272K tokens.
GPT-5.6 Terra
$0.54 / $3.26 per 1M tokens
GPT-5.6 Terra is the middle tier of OpenAI's GPT-5.6 family, built to balance intelligence and cost. The GPT-5.6 Terra API on zurelay runs the same model through one OpenAI-compatible endpoint, below OpenAI's list price. It fits production traffic that needs solid reasoning without paying for Sol.
GPT-5.6 Luna
$0.08 / $0.49 per 1M tokens
GPT-5.6 Luna is the fastest and lowest-cost tier of OpenAI's GPT-5.6 family, built for cost-sensitive, high-volume workloads. The GPT-5.6 Luna API on zurelay runs the same model through one OpenAI-compatible endpoint, below OpenAI's list price. It suits classification, routing, extraction and short summaries at scale.
GPT Image 2.5 Flare
$0.02 per image
GPT Image 2.5 Flare is OpenAI's fastest model for high-quality, everyday image generation, released on September 8, 2026. The GPT Image 2.5 Flare API on zurelay turns a text prompt into images through one OpenAI-compatible endpoint. Every image is made at high quality, at one flat price for any size.
GPT Image 2.5 Sunburst
$0.02 per image
GPT Image 2.5 Sunburst is OpenAI's most capable image model, tuned for quality and precision and released on September 8, 2026. The GPT Image 2.5 Sunburst API on zurelay turns a text prompt into detailed images through one OpenAI-compatible endpoint. Every image is made at high quality, at one flat price for any size up to 4K.
GPT Image 2
$0.02 per image
GPT Image 2 is OpenAI's image generation and editing model for production visuals such as infographics, posters and text-heavy layouts, released on April 21, 2026. The GPT Image 2 API on zurelay turns a text prompt into images through one OpenAI-compatible endpoint. Every image is made at high quality, at one flat price for any size up to 3840 pixels.
Nano Banana Pro
from $0.035 per image
Nano Banana Pro is Google DeepMind's image generation model, built on Gemini 3 Pro and served in the Gemini API as Gemini 3 Pro Image. The zurelay Nano Banana Pro API turns a text prompt into 1K, 2K or 4K images through an OpenAI-style images endpoint, for developers who need legible text, diagrams and production assets at a lower price per image.
Nano Banana 2
from $0.025 per image
Nano Banana 2 is Google's fast, general-purpose image model, served in the Gemini API as Gemini 3.1 Flash Image. The zurelay Nano Banana 2 API turns a text prompt into 1K, 2K or 4K images through an OpenAI-style images endpoint, at 63% below Google's price per image, and checks every image against Google's official size.
Gemini 3.7 Flash
$0.32 / $1.59 per 1M tokens
Gemini 3.7 Flash is Google's Flash model for coding, agentic tool use and multi-step work, released on August 13, 2026. The zurelay Gemini 3.7 Flash API serves it through one OpenAI-compatible endpoint with a 1M-token context window, at 58% below Google's list price.
Gemini 3.5 Flash-Lite
$0.12 / $1.04 per 1M tokens
Gemini 3.5 Flash-Lite is the fastest and most cost-effective model in Google's Gemini 3.5 series, released on July 21, 2026. The zurelay Gemini 3.5 Flash-Lite API serves it through one OpenAI-compatible endpoint for high-volume agent steps, search and document processing, at 58% below Google's list price.
Gemini 3.1 Flash-Lite
$0.10 / $0.59 per 1M tokens
Gemini 3.1 Flash-Lite is Google's low-latency, low-cost Gemini 3 model for high-volume work such as translation, moderation, extraction and model routing. The zurelay Gemini 3.1 Flash-Lite API serves it through one OpenAI-compatible endpoint with a 1M-token context window, at 61% below Google's list price.
xAI
Xiaomi
MiMo V2.6 Pro
$0.18 / $0.36 per 1M tokens
MiMo V2.6 Pro is Xiaomi's open-weight flagship, a 1.02-trillion-parameter model that reads text, images, audio and video and holds 1M tokens of context. The MiMo V2.6 Pro API on zurelay suits teams building agents, coding tools and multimodal apps who want Xiaomi's strongest model at 59% below its list price.
MiMo V2.6 Flash
$0.06 / $0.12 per 1M tokens
MiMo V2.6 Flash is the efficient model in Xiaomi's MiMo-V2.6 series: 15B active parameters, 1M tokens of context, and text, image, audio and video input. The MiMo V2.6 Flash API on zurelay is for teams running high-volume or latency-sensitive workloads who want near-Pro agent results at 57% below Xiaomi's list price.
DeepSeek
DeepSeek V4.1 Flash
$0.037 / $0.15 per 1M tokens
DeepSeek V4.1 Flash is DeepSeek's newest model and the one its own API now serves for Flash requests. It reads text and images, holds 1M tokens of context and runs in thinking or non-thinking mode. The zurelay DeepSeek V4.1 Flash API serves it through one OpenAI-compatible endpoint at 88% below DeepSeek's peak list price.
DeepSeek V4 Flash
$0.045 / $0.18 per 1M tokens
DeepSeek V4 Flash was the fast, low-cost half of DeepSeek's V4 launch in April 2026. DeepSeek has since retired it on its own API and answers the deepseek-v4-flash name with the newer V4.1 Flash. The zurelay DeepSeek V4 Flash API does the same, so existing code keeps working at 85% below list price.
DeepSeek V4 Flash 0731
$0.022 / $0.06 per 1M tokens
DeepSeek V4 Flash 0731 is the official July 31, 2026 release of DeepSeek V4 Flash, a 284B-parameter open-weight model tuned for agents and coding. DeepSeek's own API has moved on to V4.1 Flash, so the DeepSeek V4 Flash 0731 API on zurelay is how you keep this exact checkpoint in production.
DeepSeek V4 Pro 0813
$0.23 / $0.62 per 1M tokens
DeepSeek V4 Pro 0813 is the official August 13, 2026 release of DeepSeek V4 Pro, a 1.6-trillion-parameter open-weight model built for coding and agents. It is the model DeepSeek's own API serves as deepseek-v4-pro. The zurelay DeepSeek V4 Pro API runs it through one OpenAI-compatible endpoint at 84% below DeepSeek's peak list price.
Z.ai
GLM-5.3
$0.38 / $1.21 per 1M tokens
GLM-5.3 is Z.ai's flagship model for complex software engineering and long-running agents, with a 1M-token context window and up to 128K output tokens. The zurelay GLM-5.3 API serves the same model through one OpenAI-compatible endpoint, billed per token at 73% below Z.ai's list price.
GLM-5.2
$0.38 / $1.21 per 1M tokens
GLM-5.2 is Z.ai's open-weight model for long-horizon coding and the first GLM with a 1M-token context window. The zurelay GLM-5.2 API serves the same model through one OpenAI-compatible endpoint at 73% below Z.ai's list price.
Moonshot AI
Alibaba
Qwen 3.8 Max
$1.20 / $3.60 per 1M tokens
Qwen 3.8 Max is Alibaba's flagship Qwen 3.8 model, a 2.4-trillion-parameter mixture-of-experts with a 1M-token context window and native vision. The zurelay Qwen 3.8 Max API serves it through one OpenAI-compatible endpoint at 40% below Alibaba's international list price, for coding, agent and long-document work.
Qwen 3.8 Flash
$0.07 / $0.22 per 1M tokens
Qwen 3.8 Flash is Alibaba's fast, low-cost Qwen 3.8 model, a 125B-parameter mixture-of-experts with only 6B active per token and a 1M-token context window. The zurelay Qwen 3.8 Flash API serves it through one OpenAI-compatible endpoint at 53% below Alibaba's international list price, for coding helpers, agents and high-volume pipelines.
Qwen 3.8 Omni Flash
$0.075 / $0.235 per 1M tokens
Qwen 3.8 Omni Flash is Alibaba's multimodal Qwen 3.8 model for audio and video understanding: it takes text, images, audio and video in one request and answers in text. The zurelay Qwen 3.8 Omni Flash API serves it through one OpenAI-compatible endpoint with a 1M-token context window, at 50% below Alibaba's international list price.
Tencent
MiniMax
ByteDance
Seedance 2.5
from $0.14 per second
Seedance 2.5 is the newest model in ByteDance's Seedance video family. It turns a text prompt, or a prompt with images, into a video with generated sound, at up to 1080p and up to 30 seconds long. The zurelay Seedance 2.5 API starts at $0.14 per second of video, and a video is charged only once it's delivered.
Seedance 2.0
from $0.097 per second
Seedance 2.0 is ByteDance's video generation model for text-to-video and image-to-video, with generated sound. It makes 480p and 720p clips from 4 to 15 seconds long. The zurelay Seedance 2.0 API starts at $0.097 per second of video, and a video is charged only once it's delivered.
Seedance 2.0 Fast
from $0.074 per second
Seedance 2.0 Fast is a quicker, lower-cost version of ByteDance's Seedance 2.0 video model, trading some quality for speed and price. It makes 480p and 720p clips from 4 to 15 seconds long, with generated sound, from text or images. On zurelay it starts at $0.074 per second of video, charged only once a video is delivered.
Seedance 2.0 Mini
from $0.047 per second
Seedance 2.0 Mini is the lowest-cost model in ByteDance's Seedance 2.0 family, trading some quality for speed and price. It makes 480p and 720p clips from 4 to 15 seconds long, with generated sound, from text or images. On zurelay it costs from $0.047 per second of video, charged only once a video is delivered.
Stop paying list price.
Start saving today.
Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.