DeepSeek
Disrupted· 82.1% uptime, 7 days

DeepSeek V4 Flash 0731 API

The DeepSeek V4 Flash 0731 API: the pinned V4 Flash release, for reproducible agents

Price per 1M tokens61%off
Input
$0.022$0.076
Output
$0.06$0.153
Cached input
$0.00044
Model IDdeepseek-v4-flash-0731
chat.py
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key="YOUR_ZURELAY_KEY",
)
stream = client.chat.completions.create(
model="deepseek-v4-flash-0731",
messages=[{"role": "user", "content": "Explain quicksort in two sentences."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")

Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.

Context window
1M tokens
Max output
384K tokens
Input
Text
Output
Text
Released
Jul 31, 2026
Uptime, 7 days
82.1%

Pricing

DeepSeek V4 Flash 0731 API pricing. A fraction of the list price.

Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.

RatezurelayDeepSeek listYou save
Input
per 1M tokens
$0.022$0.07671%off
Output
per 1M tokens
$0.06$0.15361%off
Cached input
per 1M tokens
$0.00044——

One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.

Savings calculator

Move the sliders to your monthly usage.

Input tokens per month50M
Output tokens per month10M
zurelay
$1.70
DeepSeek list price
$5.33
You save every month$3.63

$43.56 a year

Overview

What is DeepSeek V4 Flash 0731?

DeepSeek V4 Flash 0731 is the official July 31, 2026 release of DeepSeek V4 Flash, a 284B-parameter open-weight model tuned for agents and coding. DeepSeek's own API has moved on to V4.1 Flash, so the DeepSeek V4 Flash 0731 API on zurelay is how you keep this exact checkpoint in production.

DeepSeek-V4-Flash-0731 is the official release of DeepSeek V4 Flash. It shipped on July 31, 2026 and replaced the April preview. It keeps the preview's architecture and size, 284B total parameters with 13B active per token, and was re-post-trained for much stronger agent performance. The DeepSeek V4 Flash 0731 API gives you that exact checkpoint under a fixed, dated model ID.

DeepSeek's model card shows the jump. On Terminal Bench 2.1 it scores 82.7, up from 61.8 for the V4 Flash preview and above the 72.1 of the much larger V4 Pro preview. On DeepSWE it scores 54.4, against 7.3 for the preview. The checkpoint ships with a DSpark speculative decoding module for faster generation.

It is a text model with a 1M-token context. Reasoning effort has three levels: low, high and max. DeepSeek recommends allowing up to 384K output tokens for high and max. The model supports tool calling, and the weights are open under the MIT license.

Why a pinned snapshot? DeepSeek retired V4 Flash on its own API on September 10, 2026 and now answers the deepseek-v4-flash name with V4.1 Flash. If your prompts, evals or agent harness were tuned on 0731, calling deepseek-v4-flash-0731 on zurelay keeps that behavior fixed. You pay $0.022 per 1M input tokens and $0.06 per 1M output tokens.

Strengths

Where DeepSeek V4 Flash 0731 shines. And what teams build with it.

01

Behavior that stays fixed

A dated checkpoint doesn't change under you, so evals, prompts and agent harnesses stay valid.

02

Large agent gains over the preview

DeepSeek's model card lists 82.7 on Terminal Bench 2.1 and 54.4 on DeepSWE, far above the April preview.

03

Small active size

Only 13B parameters are active per token, which keeps responses quick and serving cheap.

04

Three reasoning levels

DeepSeek suggests low for simple tasks, high for daily agent work and max for complex problems.

Use cases

  • Reproducible agent pipelines

    Keep a production agent on the exact model version it was tested with.

  • Evals and regression tests

    Compare prompt or harness changes against a model version that never changes.

  • Terminal and coding agents

    Run shell tasks, repository edits and tool-calling loops, where the 0731 agent training shows most.

  • Long-context text work

    Analyze long documents and codebases inside the 1M-token window.

Get started

Call DeepSeek V4 Flash 0731 in three steps.

No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.

  1. 1

    Create a key

    Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.

  2. 2

    Point your SDK at zurelay

    Change the base URL. Everything else in your code stays the same.

    https://api.zurelay.com/v1
  3. 3

    Use deepseek-v4-flash-0731

    Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.

    deepseek-v4-flash-0731

FAQ

DeepSeek V4 Flash 0731 API questions.

Anything else, write to sales@zurelay.com and an engineer will answer.

How much does the DeepSeek V4 Flash 0731 API cost?

On zurelay, DeepSeek V4 Flash 0731 costs $0.022 per 1M input tokens and $0.06 per 1M output tokens, with cached input at $0.00044. That is 61% below the list price of $0.076 input and $0.153 output.

Is there a cheaper DeepSeek V4 Flash API?

DeepSeek no longer serves the 0731 checkpoint on its own API; the deepseek-v4-flash name there now runs V4.1 Flash. zurelay keeps deepseek-v4-flash-0731 available at $0.022 input and $0.06 output per 1M tokens. You pay as you go, and credit never expires.

How do I call the DeepSeek V4 Flash 0731 API with the OpenAI SDK?

Set the OpenAI SDK's base URL to https://api.zurelay.com/v1, use your zurelay API key and set model to "deepseek-v4-flash-0731". Tool definitions and streaming work as they do with OpenAI models.

What is the difference between deepseek-v4-flash-0731 and deepseek-v4-flash?

deepseek-v4-flash-0731 always runs the July 31 V4 Flash checkpoint. deepseek-v4-flash follows DeepSeek's own API, which now serves that name with V4.1 Flash. Pin 0731 when you need behavior that doesn't change.

What is the DeepSeek V4 Flash 0731 context window?

It supports a 1M-token context. DeepSeek recommends allowing up to 384K output tokens when you use high or max reasoning effort.

What is DeepSeek V4 Flash 0731 best at?

Agent work: terminal tasks, repository-level coding and tool-calling loops. DeepSeek's model card shows it outperforming the V4 Pro preview on its agent benchmarks despite a far smaller active parameter count.

Does the DeepSeek V4 Flash 0731 API support streaming and per-key rate limits?

Streaming works with stream: true, as with the OpenAI API. You can set a requests-per-minute cap and a monthly budget on each zurelay API key.

Is it the same model DeepSeek released?

Yes. Requests run DeepSeek-V4-Flash-0731, the checkpoint DeepSeek published with open weights under the MIT license. Responses are sampled, so wording varies between runs. zurelay is independent and not affiliated with DeepSeek.

More models on zurelay

All models

Stop paying list price.
Start saving today.

Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.