DeepSeek
Operational· 100% uptime, 7 days

DeepSeek V4.1 Flash API

The DeepSeek V4.1 Flash API: DeepSeek's newest model, with image input and 1M context

Price per 1M tokens88%off
Input
$0.037$0.30
Output
$0.15$1.20
Cached input
$0.00074
Model IDdeepseek-v4.1-flash
chat.py
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key="YOUR_ZURELAY_KEY",
)
stream = client.chat.completions.create(
model="deepseek-v4.1-flash",
messages=[{"role": "user", "content": "Explain quicksort in two sentences."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")

Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.

Context window
1M tokens
Max output
384K tokens
Input
Text, Image
Output
Text
Released
Sep 10, 2026
Uptime, 7 days
100%

Pricing

DeepSeek V4.1 Flash API pricing. A fraction of the list price.

Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.

RatezurelayDeepSeek listYou save
Input
per 1M tokens
$0.037$0.3088%off
Output
per 1M tokens
$0.15$1.2088%off
Cached input
per 1M tokens
$0.00074——

One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.

Savings calculator

Move the sliders to your monthly usage.

Input tokens per month50M
Output tokens per month10M
zurelay
$3.35
DeepSeek list price
$27.00
You save every month$23.65

$283.80 a year

Overview

What is DeepSeek V4.1 Flash?

DeepSeek V4.1 Flash is DeepSeek's newest model and the one its own API now serves for Flash requests. It reads text and images, holds 1M tokens of context and runs in thinking or non-thinking mode. The zurelay DeepSeek V4.1 Flash API serves it through one OpenAI-compatible endpoint at 88% below DeepSeek's peak list price.

DeepSeek released V4.1 Flash on September 10, 2026 as the smallest model in its new architecture family, with native visual understanding. It is the model behind DeepSeek's deepseek-flash API name, and DeepSeek also routes its retired deepseek-v4-flash name to it. The DeepSeek V4.1 Flash API accepts text and images, returns text, and supports a 1M-token context with up to 384K output tokens.

Under the hood it is a 552B-parameter mixture-of-experts backbone with a new causal encoder-decoder design. It activates 8B parameters per token when reading input and 16B when generating output. With FP4 KV caching, the cache takes 890 bytes per token, about a quarter of V4 Flash. DeepSeek says this lowers the cost of running long-horizon agents at scale.

DeepSeek reports that tests by multiple parties put V4.1 Flash ahead of V4 Pro on performance, cost, speed and total runtime. Thinking mode is the default on DeepSeek's API, and you can switch it off when you want direct answers without a reasoning trace. The model supports tool calls and JSON output, and the weights are open under the MIT license.

On zurelay the model ID is deepseek-v4.1-flash, and DeepSeek's own name deepseek-flash works too, so code written for DeepSeek's API only needs a new base URL and key. You pay $0.037 per 1M input tokens and $0.15 per 1M output tokens at any hour. DeepSeek's own rates change between peak and off-peak hours.

Strengths

Where DeepSeek V4.1 Flash shines. And what teams build with it.

01

Ahead of V4 Pro at Flash cost

DeepSeek reports V4.1 Flash beats V4 Pro on performance, cost, speed and total runtime in tests by multiple parties.

02

Cheap long context

A KV cache of 890 bytes per token, about a quarter of V4 Flash, keeps 1M-token prompts and long agent sessions affordable to serve.

03

Native image understanding

Screenshots, charts and scanned documents go in alongside text, with no separate vision model.

04

Drop-in for DeepSeek code

zurelay accepts DeepSeek's deepseek-flash model name, so existing clients switch with a base URL and key change.

Use cases

  • Long-running agents

    Run agents that keep long tool histories in context. The small KV cache was designed for exactly this.

  • Coding tools

    Use thinking mode for code generation, review and repository Q&A, and turn it off for quick completions.

  • Document and screenshot parsing

    Extract fields from scanned forms, charts and UI screenshots straight into JSON.

  • Bulk text processing

    Summarize, classify and translate large volumes of text at a low per-token rate.

Get started

Call DeepSeek V4.1 Flash in three steps.

No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.

  1. 1

    Create a key

    Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.

  2. 2

    Point your SDK at zurelay

    Change the base URL. Everything else in your code stays the same.

    https://api.zurelay.com/v1
  3. 3

    Use deepseek-v4.1-flash

    Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.

    deepseek-v4.1-flash

FAQ

DeepSeek V4.1 Flash API questions.

Anything else, write to sales@zurelay.com and an engineer will answer.

How much does the DeepSeek V4.1 Flash API cost?

On zurelay, DeepSeek V4.1 Flash costs $0.037 per 1M input tokens and $0.15 per 1M output tokens, with cached input at $0.00074. DeepSeek's peak-hour list price is $0.30 input and $1.20 output, and it charges half that off-peak. zurelay uses one rate at all hours, 88% below DeepSeek's peak list price.

Is there a cheaper DeepSeek API?

zurelay serves DeepSeek V4.1 Flash at 88% below DeepSeek's peak list price, with no peak-hour pricing to plan around. You pay as you go, and credit never expires.

How do I call the DeepSeek V4.1 Flash API with the OpenAI SDK?

Set the OpenAI SDK's base URL to https://api.zurelay.com/v1, use your zurelay API key and set model to "deepseek-v4.1-flash". If your code already uses DeepSeek's name deepseek-flash, zurelay accepts that too.

What is the DeepSeek V4.1 Flash context window?

DeepSeek V4.1 Flash supports 1M tokens of context and up to 384K output tokens per response.

How is DeepSeek V4.1 Flash different from DeepSeek V4 Flash?

V4.1 Flash uses a new architecture: a 552B-parameter causal encoder-decoder, against V4 Flash's 284B-parameter model. It adds native image input and cuts the KV cache to about a quarter of V4 Flash's. DeepSeek retired V4 Flash on its own API when V4.1 Flash launched.

What is DeepSeek V4.1 Flash best at?

Agents, coding and long-context work where cost matters. DeepSeek designed it for faster inference, higher throughput and cheaper long-horizon agents, and reports it ahead of V4 Pro in tests by multiple parties.

Does DeepSeek V4.1 Flash have a thinking mode?

Yes. It supports thinking and non-thinking modes, and thinking is the default on DeepSeek's own API. Use thinking for multi-step reasoning and coding, and non-thinking for quick, direct answers.

Does the DeepSeek V4.1 Flash API support streaming and per-key rate limits?

Streaming works with stream: true, as with the OpenAI API. You can set a requests-per-minute cap and a monthly budget on each zurelay API key.

Is it the same model DeepSeek serves?

Yes. Requests to deepseek-v4.1-flash run DeepSeek V4.1 Flash, the model DeepSeek serves as deepseek-flash. Responses are sampled, so wording varies between runs, as it does on DeepSeek's API. zurelay is independent and not affiliated with DeepSeek.

More models on zurelay

All models

Stop paying list price.
Start saving today.

Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.