Behavior that stays fixed
A dated checkpoint doesn't change under you, so evals, prompts and agent harnesses stay valid.
The DeepSeek V4 Flash 0731 API: the pinned V4 Flash release, for reproducible agents
deepseek-v4-flash-0731from openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key="YOUR_ZURELAY_KEY",)stream = client.chat.completions.create( model="deepseek-v4-flash-0731", messages=[{"role": "user", "content": "Explain quicksort in two sentences."}], stream=True,)for chunk in stream: if chunk.choices: print(chunk.choices[0].delta.content or "", end="")Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.
Pricing
Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.
| Rate | zurelay | DeepSeek list | You save |
|---|---|---|---|
Input per 1M tokens | $0.022 | 71%off | |
Output per 1M tokens | $0.06 | 61%off | |
Cached input per 1M tokens | $0.00044 | — | — |
One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.
Move the sliders to your monthly usage.
$43.56 a year
Overview
DeepSeek V4 Flash 0731 is the official July 31, 2026 release of DeepSeek V4 Flash, a 284B-parameter open-weight model tuned for agents and coding. DeepSeek's own API has moved on to V4.1 Flash, so the DeepSeek V4 Flash 0731 API on zurelay is how you keep this exact checkpoint in production.
DeepSeek-V4-Flash-0731 is the official release of DeepSeek V4 Flash. It shipped on July 31, 2026 and replaced the April preview. It keeps the preview's architecture and size, 284B total parameters with 13B active per token, and was re-post-trained for much stronger agent performance. The DeepSeek V4 Flash 0731 API gives you that exact checkpoint under a fixed, dated model ID.
DeepSeek's model card shows the jump. On Terminal Bench 2.1 it scores 82.7, up from 61.8 for the V4 Flash preview and above the 72.1 of the much larger V4 Pro preview. On DeepSWE it scores 54.4, against 7.3 for the preview. The checkpoint ships with a DSpark speculative decoding module for faster generation.
It is a text model with a 1M-token context. Reasoning effort has three levels: low, high and max. DeepSeek recommends allowing up to 384K output tokens for high and max. The model supports tool calling, and the weights are open under the MIT license.
Why a pinned snapshot? DeepSeek retired V4 Flash on its own API on September 10, 2026 and now answers the deepseek-v4-flash name with V4.1 Flash. If your prompts, evals or agent harness were tuned on 0731, calling deepseek-v4-flash-0731 on zurelay keeps that behavior fixed. You pay $0.022 per 1M input tokens and $0.06 per 1M output tokens.
Strengths
A dated checkpoint doesn't change under you, so evals, prompts and agent harnesses stay valid.
DeepSeek's model card lists 82.7 on Terminal Bench 2.1 and 54.4 on DeepSWE, far above the April preview.
Only 13B parameters are active per token, which keeps responses quick and serving cheap.
DeepSeek suggests low for simple tasks, high for daily agent work and max for complex problems.
Use cases
Keep a production agent on the exact model version it was tested with.
Compare prompt or harness changes against a model version that never changes.
Run shell tasks, repository edits and tool-calling loops, where the 0731 agent training shows most.
Analyze long documents and codebases inside the 1M-token window.
Get started
No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.
Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.
Change the base URL. Everything else in your code stays the same.
https://api.zurelay.com/v1Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.
deepseek-v4-flash-0731Works with the tools you already use
Compare
$0.037 / $0.15 per 1M tokens
Pick DeepSeek V4.1 Flash for DeepSeek's newest Flash model with image input, if you don't need this exact checkpoint.
DeepSeek V4.1 Flash API$0.23 / $0.62 per 1M tokens
Pick DeepSeek V4 Pro 0813 for the larger pinned V4 model, which scores 87.9 against 82.7 on Terminal Bench 2.1.
DeepSeek V4 Pro 0813 API$0.045 / $0.18 per 1M tokens
Pick deepseek-v4-flash if your code uses that name and you're fine with it following DeepSeek's current Flash model.
DeepSeek V4 Flash APIFAQ
Anything else, write to sales@zurelay.com and an engineer will answer.
On zurelay, DeepSeek V4 Flash 0731 costs $0.022 per 1M input tokens and $0.06 per 1M output tokens, with cached input at $0.00044. That is 61% below the list price of $0.076 input and $0.153 output.
DeepSeek no longer serves the 0731 checkpoint on its own API; the deepseek-v4-flash name there now runs V4.1 Flash. zurelay keeps deepseek-v4-flash-0731 available at $0.022 input and $0.06 output per 1M tokens. You pay as you go, and credit never expires.
Set the OpenAI SDK's base URL to https://api.zurelay.com/v1, use your zurelay API key and set model to "deepseek-v4-flash-0731". Tool definitions and streaming work as they do with OpenAI models.
deepseek-v4-flash-0731 always runs the July 31 V4 Flash checkpoint. deepseek-v4-flash follows DeepSeek's own API, which now serves that name with V4.1 Flash. Pin 0731 when you need behavior that doesn't change.
It supports a 1M-token context. DeepSeek recommends allowing up to 384K output tokens when you use high or max reasoning effort.
Agent work: terminal tasks, repository-level coding and tool-calling loops. DeepSeek's model card shows it outperforming the V4 Pro preview on its agent benchmarks despite a far smaller active parameter count.
Streaming works with stream: true, as with the OpenAI API. You can set a requests-per-minute cap and a monthly budget on each zurelay API key.
Yes. Requests run DeepSeek-V4-Flash-0731, the checkpoint DeepSeek published with open weights under the MIT license. Responses are sampled, so wording varies between runs. zurelay is independent and not affiliated with DeepSeek.
Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.