DeepSeek API pricing. Up to 90% below list.
DeepSeek's V4 models are open-weight models built for coding and agents. V4.1 Flash is the newest, and the one DeepSeek's own API now serves for Flash requests. V4 Pro is the largest, at 1.6 trillion parameters. Zurelay serves them through one OpenAI-compatible endpoint with 1M tokens of context, at the same price at every hour.
- DeepSeek models
- 4
- Below DeepSeek's list
- 43–90%
- From, per 1M input
- $0.029
Pricing
Every DeepSeek model. Live prices, one key.
Each DeepSeek model on Zurelay, with DeepSeek's list price beside ours.
| Model | Context | Input / 1M | Output / 1M | 10K requests | You save |
|---|---|---|---|---|---|
| DeepSeek V4.1 Flashdeepseek-v4.1-flash | 1M | $0.03 | $0.12 | $1.20 | 90%off |
| DeepSeek V4 Flashdeepseek-v4-flash | 1M | $0.03 | $0.12 | $1.20 | 90%off |
| DeepSeek V4 Flash 0731deepseek-v4-flash-0731 | 1M | $0.029 | $0.087 | $1.015 | 43%off |
| DeepSeek V4 Pro 0813deepseek-v4-pro-0813 | 1M | $0.23 | $0.62 | $7.70 | 84%off |
Prices in US dollars per 1M tokens, live from the catalog, with DeepSeek's list price struck through. 10K requests is 10,000 requests of 2,000 tokens in and 500 out, at each price. You save is on output tokens.
Which model
Which DeepSeek model to use. Same key for all of them.
Most work
DeepSeek V4.1 Flash
$0.03 / $0.12 per 1M tokens
DeepSeek's newest model. It reads text and images, holds 1M tokens of context and runs in thinking or non-thinking mode.
DeepSeek V4.1 Flash APIHarder coding and agents
DeepSeek V4 Pro 0813
$0.23 / $0.62 per 1M tokens
The August 13, 2026 release of DeepSeek V4 Pro, the 1.6-trillion-parameter model DeepSeek's API serves as deepseek-v4-pro.
DeepSeek V4 Pro 0813 APIA pinned checkpoint
DeepSeek V4 Flash 0731
$0.029 / $0.087 per 1M tokens
The July 31, 2026 release of V4 Flash, for agents that need this exact checkpoint in production.
DeepSeek V4 Flash 0731 APIAlso on Zurelay, for prompts, evals and agents tuned on them: DeepSeek V4 Flash.
Quickstart
Change two lines. Keep your code.
Point the OpenAI SDK, or any OpenAI-compatible tool, at https://api.zurelay.com/v1 and use your Zurelay key. Streaming, tool calling and structured output work as documented.
Add credit from $10, with no subscription. The quickstart has the rest.
from openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key="YOUR_ZURELAY_KEY",)stream = client.chat.completions.create( model="deepseek-v4.1-flash", messages=[{"role": "user", "content": "Explain quicksort in two sentences."}], stream=True,)for chunk in stream: if chunk.choices: print(chunk.choices[0].delta.content or "", end="")Verified
The DeepSeek you asked for. Checked every day.
Model capacity is often bought ahead, through prepaid credits, committed-spend contracts and volume tiers, and not all of it gets used. Sellers offer it below list price. Zurelay tests each route, pools several per model and sends every request to one that is available and verified.
Before a route serves requests, and every day after, it has to pass identity, dated-knowledge and reasoning checks that catch a different or older model sold under a newer name. Failed attempts move to another route, and you aren’t billed for them. How models are verified.
Live health for every model is on the status page.
How much does the DeepSeek API cost?
On Zurelay, DeepSeek models cost from $0.029 input and $0.087 output per 1M tokens (DeepSeek V4 Flash 0731) to $0.23 and $0.62 (DeepSeek V4 Pro 0813), 43–90% below DeepSeek's list prices. There's no subscription or minimum spend: you add prepaid credit from $10 and pay for the tokens you use.
Is it the real DeepSeek?
Yes. Each model is served over several routes run by independent sellers. Every route is tested before it serves requests and every day after, with identity, dated-knowledge and reasoning checks that catch a different or older model sold under the name you asked for. If an answer doesn't look right, send its x-request-id to support and we check the route that served it.
Why is it cheaper than DeepSeek's own API?
Sellers offer model capacity they bought ahead, through prepaid credits, committed-spend contracts and volume tiers, below list price. Zurelay tests it, pools several routes per model and sends each request to one that is available and verified. Requests aren't served under DeepSeek's own terms for its customers; our Terms explain how they are served.
How do I start calling DeepSeek through Zurelay?
Create a key, then point the OpenAI SDK, or any OpenAI-compatible tool, at https://api.zurelay.com/v1 with a model ID such as deepseek-v4.1-flash. Code that already uses the OpenAI SDK only needs the base URL, the key and the model name changed.
What is the saving measured against?
DeepSeek's standard list price. DeepSeek charges less in its off-peak hours; Zurelay's price is the same at every hour.
What does deepseek-v4-flash answer with?
DeepSeek retired V4 Flash on its own API and answers that name with the newer V4.1 Flash. Zurelay does the same, so existing code keeps working. To keep the original V4 Flash, use deepseek-v4-flash-0731.
Stop paying list price.
Start saving today.
Create a key in seconds, point the SDK you already use at Zurelay, and every request costs up to 90% less from the first token.