DeepSeek
Operational· 100% uptime, 7 days

DeepSeek V4 Pro 0813 API

The DeepSeek V4 Pro API: the 0813 release of DeepSeek's largest V4 model

Price per 1M tokens84%off
Input
$0.23$1.32
Output
$0.62$3.96
Cached input
$0.00759
Model IDdeepseek-v4-pro-0813
chat.py
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key="YOUR_ZURELAY_KEY",
)
stream = client.chat.completions.create(
model="deepseek-v4-pro-0813",
messages=[{"role": "user", "content": "Explain quicksort in two sentences."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")

Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.

Context window
1M tokens
Max output
384K tokens
Input
Text
Output
Text
Released
Aug 13, 2026
Uptime, 7 days
100%

Pricing

DeepSeek V4 Pro 0813 API pricing. A fraction of the list price.

Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.

RatezurelayDeepSeek listYou save
Input
per 1M tokens
$0.23$1.3283%off
Output
per 1M tokens
$0.62$3.9684%off
Cached input
per 1M tokens
$0.00759——

One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.

Savings calculator

Move the sliders to your monthly usage.

Input tokens per month50M
Output tokens per month10M
zurelay
$17.70
DeepSeek list price
$105.60
You save every month$87.90

$1,055 a year

Overview

What is DeepSeek V4 Pro 0813?

DeepSeek V4 Pro 0813 is the official August 13, 2026 release of DeepSeek V4 Pro, a 1.6-trillion-parameter open-weight model built for coding and agents. It is the model DeepSeek's own API serves as deepseek-v4-pro. The zurelay DeepSeek V4 Pro API runs it through one OpenAI-compatible endpoint at 84% below DeepSeek's peak list price.

DeepSeek V4 Pro 0813 is the general-availability release of DeepSeek V4 Pro. It shipped on August 13, 2026 and replaced the April preview. It is the largest model in the V4 family: a mixture-of-experts with 1.6T total parameters and 49B active per token, a 1M-token context window and up to 384K output tokens. The official DeepSeek V4 Pro API serves this exact version under the deepseek-v4-pro name.

DeepSeek focused the 0813 release on agent work in production. Its model card shows Terminal Bench 2.1 rising from 72.1 to 87.9 over the preview, DeepSWE from 12.8 to 62.7 and CyberGym from 52.7 to 83.3. The checkpoint keeps the preview's structure and adds a DSpark speculative decoding module for faster generation.

Reasoning effort has three levels. DeepSeek suggests low for simple tasks, high for daily agent workflows and max for complex tasks. The model supports tool calls and JSON output, takes text input only, and its weights are open under the MIT license.

On zurelay the model ID is deepseek-v4-pro-0813, and deepseek-v4-pro works as an alias. You pay $0.23 per 1M input tokens and $0.62 per 1M output tokens, one rate at all hours, while DeepSeek's own price changes between peak and off-peak hours.

Strengths

Where DeepSeek V4 Pro 0813 shines. And what teams build with it.

01

Big gains on agent benchmarks

Against the V4 Pro preview, Terminal Bench 2.1 went from 72.1 to 87.9 and DeepSWE from 12.8 to 62.7.

02

The largest V4 model

1.6T total parameters and 49B active per token, nearly four times the active size of V4 Flash.

03

Matches the lab's current Pro

DeepSeek still serves this checkpoint as deepseek-v4-pro, so you get the same version the lab ships today.

04

Reasoning depth per request

Low, high and max effort let you trade speed for depth on each call.

Use cases

  • Production coding agents

    Run repository-scale changes, terminal tasks and test-and-fix loops, the areas where 0813 improved most.

  • Security analysis

    DeepSeek reports 83.3 on CyberGym, up from 52.7 for the preview.

  • Research with tools

    Answer multi-step questions with search or code tools. DeepSeek reports Humanity's Last Exam with tools rising from 48.2 to 60.0.

  • Long-output generation

    Write long reports, specs or complete files with up to 384K output tokens per response.

Get started

Call DeepSeek V4 Pro 0813 in three steps.

No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.

  1. 1

    Create a key

    Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.

  2. 2

    Point your SDK at zurelay

    Change the base URL. Everything else in your code stays the same.

    https://api.zurelay.com/v1
  3. 3

    Use deepseek-v4-pro-0813

    Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.

    deepseek-v4-pro-0813

FAQ

DeepSeek V4 Pro 0813 API questions.

Anything else, write to sales@zurelay.com and an engineer will answer.

How much does the DeepSeek V4 Pro API cost?

On zurelay, DeepSeek V4 Pro 0813 costs $0.23 per 1M input tokens and $0.62 per 1M output tokens, with cached input at $0.00759. DeepSeek's peak-hour list price is $1.32 input and $3.96 output, and it charges half that off-peak. zurelay uses one rate at all hours, 84% below DeepSeek's peak list price.

Is there a cheaper DeepSeek V4 Pro API?

zurelay serves DeepSeek V4 Pro 0813 at 84% below DeepSeek's peak list price, with no peak-hour pricing to plan around. You pay as you go, and credit never expires.

How do I call the DeepSeek V4 Pro API with the OpenAI SDK?

Set the OpenAI SDK's base URL to https://api.zurelay.com/v1, use your zurelay API key and set model to "deepseek-v4-pro-0813". If your code already sends deepseek-v4-pro, zurelay accepts that name too.

What is the DeepSeek V4 Pro context window?

DeepSeek V4 Pro 0813 supports 1M tokens of context and up to 384K output tokens per response.

What changed in DeepSeek V4 Pro 0813 compared with the preview?

0813 keeps the preview's architecture, adds a DSpark speculative decoding module and was trained for much stronger agent performance. DeepSeek's model card shows Terminal Bench 2.1 up from 72.1 to 87.9, DeepSWE from 12.8 to 62.7 and NL2Repo from 38.5 to 61.5. DeepSeek also introduced three reasoning effort levels: low, high and max.

What is DeepSeek V4 Pro 0813 best at?

Agentic coding, terminal work, security tasks and tool-assisted research, where the 0813 gains are largest. DeepSeek says the improvements are most pronounced in production environments.

Does the DeepSeek V4 Pro API support streaming and per-key rate limits?

Streaming works with stream: true, as with the OpenAI API. You can set a requests-per-minute cap and a monthly budget on each zurelay API key.

Is it the same model DeepSeek serves?

Yes. DeepSeek's own API serves deepseek-v4-pro as DeepSeek-V4-Pro-0813, and that is the model zurelay runs. Responses are sampled, so wording varies between runs. zurelay is independent and not affiliated with DeepSeek.

More models on zurelay

All models

Stop paying list price.
Start saving today.

Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.