DeepSeek
Operational· 100% uptime, 7 days

DeepSeek V4 Flash API

The DeepSeek V4 Flash API name your code already uses, at 85% below list price

Price per 1M tokens85%off
Input
$0.045$0.30
Output
$0.18$1.20
Cached input
$0.0009
Model IDdeepseek-v4-flash
chat.py
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key="YOUR_ZURELAY_KEY",
)
stream = client.chat.completions.create(
model="deepseek-v4-flash",
messages=[{"role": "user", "content": "Explain quicksort in two sentences."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")

Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.

Context window
1M tokens
Max output
384K tokens
Input
Text
Output
Text
Released
Apr 24, 2026
Uptime, 7 days
100%

Pricing

DeepSeek V4 Flash API pricing. A fraction of the list price.

Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.

RatezurelayDeepSeek listYou save
Input
per 1M tokens
$0.045$0.3085%off
Output
per 1M tokens
$0.18$1.2085%off
Cached input
per 1M tokens
$0.0009——

One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.

Savings calculator

Move the sliders to your monthly usage.

Input tokens per month50M
Output tokens per month10M
zurelay
$4.05
DeepSeek list price
$27.00
You save every month$22.95

$275.40 a year

Overview

What is DeepSeek V4 Flash?

DeepSeek V4 Flash was the fast, low-cost half of DeepSeek's V4 launch in April 2026. DeepSeek has since retired it on its own API and answers the deepseek-v4-flash name with the newer V4.1 Flash. The zurelay DeepSeek V4 Flash API does the same, so existing code keeps working at 85% below list price.

The DeepSeek V4 Flash API launched on April 24, 2026 with the smaller model in the DeepSeek V4 family: 284B total parameters with 13B active, a 1M-token context, and thinking and non-thinking modes. DeepSeek described it as fast and cost-effective, with reasoning close to V4 Pro. On July 31, 2026 DeepSeek shipped the official release, a re-post-trained checkpoint named DeepSeek-V4-Flash-0731.

On September 10, 2026 DeepSeek released V4.1 Flash and retired V4 Flash on its API. For compatibility, DeepSeek now serves requests for deepseek-v4-flash with V4.1 Flash, billed at the Flash price. zurelay's deepseek-v4-flash follows the same rule, so the DeepSeek V4 Flash API name in your code keeps working and returns the current Flash model.

In practice that means a 1M-token context window, up to 384K output tokens, thinking and non-thinking modes, tool calls and JSON output. If you need the V4 Flash checkpoint itself for reproducible results, call deepseek-v4-flash-0731. If you are starting fresh or want image input, call deepseek-v4.1-flash directly.

On zurelay, deepseek-v4-flash costs $0.045 per 1M input tokens and $0.18 per 1M output tokens, one rate at all hours. Point any OpenAI SDK at https://api.zurelay.com/v1 and keep your model string as it is.

Strengths

Where DeepSeek V4 Flash shines. And what teams build with it.

01

No code changes

Apps that call deepseek-v4-flash keep working. Change the base URL and API key and nothing else.

02

The current Flash model behind the name

As on DeepSeek's own API, V4.1 Flash answers this name. DeepSeek reports V4.1 Flash ahead of V4 Pro in tests by multiple parties.

03

1M context, 384K output

There is room for whole repositories, long documents and long agent traces, plus very long generated answers.

04

A pinned option on the same key

Need the V4 Flash checkpoint itself? deepseek-v4-flash-0731 serves that fixed release.

Use cases

  • Migrating existing DeepSeek apps

    Move production traffic that uses the deepseek-v4-flash name without touching prompts or model strings.

  • Agent and tool-calling loops

    Run multi-step agents with tool calls and JSON output over long contexts.

  • Coding assistants

    Generate, review and explain code with thinking on, or get quick completions with it off.

  • Bulk processing

    Summarize, classify and extract from large volumes of text at a low per-token rate.

Get started

Call DeepSeek V4 Flash in three steps.

No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.

  1. 1

    Create a key

    Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.

  2. 2

    Point your SDK at zurelay

    Change the base URL. Everything else in your code stays the same.

    https://api.zurelay.com/v1
  3. 3

    Use deepseek-v4-flash

    Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.

    deepseek-v4-flash

FAQ

DeepSeek V4 Flash API questions.

Anything else, write to sales@zurelay.com and an engineer will answer.

How much does the DeepSeek V4 Flash API cost?

On zurelay, deepseek-v4-flash costs $0.045 per 1M input tokens and $0.18 per 1M output tokens, with cached input at $0.0009. The list price is $0.30 input and $1.20 output per 1M, DeepSeek's peak-hour Flash rate, so you save 85%. zurelay uses one rate at all hours.

Is there a cheaper DeepSeek V4 API?

zurelay serves the deepseek-v4-flash name at 85% below DeepSeek's peak list price, with no peak-hour pricing to plan around. You pay as you go, and credit never expires.

Which model answers deepseek-v4-flash requests?

DeepSeek retired V4 Flash on its API on September 10, 2026 and now serves the deepseek-v4-flash name with DeepSeek V4.1 Flash. zurelay does the same, so you get the current Flash model under the name you already use. To pin the V4 Flash checkpoint itself, use deepseek-v4-flash-0731.

How do I call the DeepSeek V4 Flash API with the OpenAI SDK?

Set the OpenAI SDK's base URL to https://api.zurelay.com/v1, use your zurelay API key and keep model "deepseek-v4-flash". Prompts, tool definitions and streaming code stay the same.

What is the DeepSeek V4 Flash context window?

1M tokens of context and up to 384K output tokens per response.

What is DeepSeek V4 Flash best at?

Fast, low-cost work at scale: coding help, tool-calling agents, extraction and summarization over long inputs. It is also the easiest way to move an app built on DeepSeek's V4 Flash name without code changes.

Does the DeepSeek V4 Flash API support streaming and per-key rate limits?

Streaming works with stream: true, as with the OpenAI API. You can set a requests-per-minute cap and a monthly budget on each zurelay API key.

Are responses the same as from DeepSeek's API?

Yes, in the same sense as DeepSeek's own API: both serve deepseek-v4-flash with DeepSeek V4.1 Flash. Responses are sampled, so wording varies between runs. zurelay is independent and not affiliated with DeepSeek.

More models on zurelay

All models

Stop paying list price.
Start saving today.

Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.