We introduce
DeepSeek-V4.1-Flash
, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. The model natively processes images and text, and generates text autoregressively.
Architecture.
DeepSeek-V4.1-Flash adopts a
Causal Encoder-Decoder (CED)
architecture: a 40-layer Transformer organized as a 20-layer causal encoder followed by a 20-layer decoder. With CED, the decoder's global KV cache is projected from the final encoder hidden states rather than derived from each decoder layer's own hidden states. This allows the model to activate only
8B parameters per token during prefill
and
16B during decode
, substantially improving cost efficiency for input-heavy agentic workloads.
SWA Bounded Replay
reconstructs missing SWA KV states by replaying only the most recent
n
_win tokens, avoiding the need to persist SWA KV to SSD and reducing the persistent KV cache footprint to roughly
1/8
of that of DeepSeek-V4-Flash.
Compressed Sparse Attention 2 (CSA2).
DeepSeek-V4.1-Flash uses CSA2, which assigns each attention layer one of three static modes —
Full
,
Reindex
, or
Reuse
— to share main KV and indexer K across layers and reuse Top-K sparse-attention indices. In the decoder, a
Hierarchical Sparse Indexer
further restricts later indexing layers to a candidate pool constructed by the first Full Mode layer, bounding deeper indexer cost independently of context length. Combined with
FP4 main KV caching
(E2M1 format, one E4M3 scale per 16 channels), these designs reduce the global KV cache footprint to
890 bytes per token
— roughly
1/4
of DeepSeek-V4-Flash.
Additional architectural components
include Single-Pass mHC (revised residual-stream mixing with an efficient Mega-mHC kernel), Engram conditional memory (196B parameters, sparsely accessed via token-based lookup), and DSpark speculative decoding (semi-autoregressive draft generation with confidence-scheduled verification). The model uses 1 shared expert and 384 routed experts per MoE layer, activating 6 routed experts per token.
Multimodal architecture.
A vision encoder (DeepSeek-ViT, trained from scratch with 2D-RoPE and 3×3 pixel-unshuffle downsampling) and a two-layer MLP projector convert images into visual embeddings, processed jointly with text embeddings from the start of language-model pre-training.
Pre-training.
DeepSeek-V4.1-Flash is trained from scratch on a multimodal corpus comprising
45T tokens
, with sparse attention trained at a sequence length of 64K and context extended to 1M tokens at 34T tokens.
Post-training.
The post-training recipe follows the standard SFT → RL → on-policy distillation (OPD) paradigm without algorithmic modifications. All substantive changes lie instead in the data pipeline: large-scale automated synthesis of agent tasks and environments with progressive scaling of data, tasks, and rollouts. The model supports a
continuously controllable reasoning effort
setting (integer 1–100) that trades inference cost for accuracy.
Figure 1. (a) Performance of DeepSeek-V4.1-Flash and counterparts on agentic benchmarks. (b) Global KV cache size per token (bytes) across generations of DeepSeek models. DeepSeek-V4.1-Flash achieves approximately 4-fold and 437-fold reductions relative to DeepSeek-V4-Flash and DeepSeek-V1, respectively.
Evaluation Results
Base Model
All base models are evaluated in our internal framework under the same evaluation settings. Scores within 0.3 of each other are considered equivalent.
DeepSeek-V4.1-Flash supports a continuously controllable reasoning effort from 1 to 100. All instruct results below use the maximum effort setting (
reasoning_effort=100
). Evaluations use
temperature=1.0, top_p=0.95
.
For code agent benchmarks (Terminal-Bench 2.1/3.0/4.0, DeepSWE v1.1, NL2Repo-Bench, ProgramBench), the model is evaluated with the Minimal mode of DeepSeek Harness and a 1M-token context window. To align with official setup requirements, the mini-SWE harness is used for DeepSWE v1.1, and the Claude Code harness for SEC-Bench Pro. Visual agent benchmarks (Chartography, BabyVision, ZeroBench) use the Claude Code harness with a 512k-token context window. Agent's Last Exam and AutomationBench use their official scaffolds. All agentic evaluations use
temperature=1.0, top_p=0.95
.
Comparison with frontier models (Max reasoning effort)
Benchmark (Metric)
Opus-5.0
GPT-5.6 Sol
K3
GLM-5.3
DS-V4-Pro
DS-V4-Flash
DS-V4.1-Flash
Reasoning
GPQA Diamond (Pass@1)
93.4
94.1
92.9
88.1
92.4
89.9
90.9
HLE (Pass@1)
56.3
44.5
43.5
42.0†
42.7†
37.8†
36.8 (39.1†)
Codeforces (Rating)
—
—
—
—
3348
3289
3471
MathArena Apex (Pass@1)
—
—
65.6
—
65.3
58.6
65.6
Agentic
Terminal-Bench 2.1 (Pass@1)
89.1
88.8
88.3
88.2
87.9
82.7
90.6
Terminal-Bench 3.0 (Pass@1)
43.3
34.4
17.7
28.3
11.8
7.6
30.0
Terminal-Bench 4.0 (Pass@1)
51.8
39.9
12.6
37.9
12.4
7.0
31.2
DeepSWE v1.1 (Resolved)
74.0
73.0
67.5
66.9
62.7
54.4
74.2
ProgramBench (Almost@1)
37.0
23.0
17.5
19.0
15.5
—
20.3
NL2Repo-Bench (Score)
75.3
56.8
58.0
58.0
61.5
54.2
64.0
CyberGym (Pass@1)
—
84.5
80.0
84.5
83.3
76.7
88.1
SEC-Bench Pro (Pass@1)
—
74.3
—
—
56.4
30.9
62.8
ExploitGym (Pass@1)
22.1
33.7
—
15.0
5.4
1.8
15.3
HLE w/ tools (Pass@1)
63.6
—
59.8
62.5
60.0
51.5
63.9
AutomationBench (Pass@1)
50.3
45.8
46.7
48.8
43.2
37.7
54.8
Agent's Last Exam (Pass@1)
28.6
26.7
27.6
28.5
25.7
25.2
31.8
Chartography w/ tools (Pass@1)
84.0
79.9
68.1
—
—
—
78.9
BabyVision w/ tools (Pass@1)
94.1
88.9
85.7
—
—
—
89.6
ZeroBench-main w/ tools (Pass@5)
52.0
53.0
41.0
—
—
—
49.0
† Text-only subset of HLE.
Performance across agent scaffolds (DeepSWE v1.1 and Terminal-Bench 2.1, Max reasoning effort)
All scaffolds use N=8 samples per task on DeepSWE v1.1 and N=3 on Terminal-Bench 2.1, with Linux containers,
temperature=1.0
,
top_p=0.95
, a 1M-token context limit, and max_steps=500 per agent. Terminal-Bench 2.1 is evaluated without network access.
Benchmark (Metric)
Claude Code
Codex
OpenCode
Pi
mini-SWE
DSH Minimal
DSH Standard
DSH PTC
DeepSWE v1.1 (Resolved)
69.8
65.6
65.5
66.2
74.2
72.6
70.5
67.6
Terminal-Bench 2.1 (Pass@1)
88.0
84.1
85.0
86.1
90.3
90.6
85.8
85.8
Prompt Encoding
This release does not include a Jinja-format chat template. The
encoding
folder contains a self-contained Python reference implementation (
encoding.py
) with test cases for multi-turn conversations, tool calling, thinking mode, numeric reasoning effort, mid-conversation system messages, and interleaved image content.
For production use, we additionally release
deepseek-recipe
, a set of Rust libraries with Python bindings that provides the same prompt format as a maintained, protocol-aware toolkit. It converts Messages, Chat Completions, and Responses API requests into the Conversation format, encodes them into DeepSeek V4 and V4.1 prompts or token IDs, and parses model output back into complete or streamed responses — covering thinking, tool calls, images, and generation settings. Model inference, tool execution, and HTTP transport are left to the caller.
Minimal Inference
Please refer to the
inference
folder for instructions on weight conversion and running inference locally.
Recommended sampling parameters:
Parameter
Value
temperature
1.0
top_p
0.95 or 1.0
context_window
1M tokens
max_tokens
≥ 256K
Reproducing DeepSWE Benchmark Results
The
evaluation
folder contains step-by-step instructions for reproducing the DeepSWE v1.1 benchmark results, covering both the
dsh-minimal
agent and the official
mini-swe-agent
. The patch required to integrate
dsh-minimal
with
Pier
is also included there.
License
This repository and the model weights are licensed under the
MIT License
.
Citation
@misc{deepseekai2026deepseekv41flash,
title={DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression},
author={DeepSeek-AI},
year={2026},
}
Contact
If you have any questions, please raise an issue or contact us at
[email protected]
.
Runs of deepseek-ai DeepSeek-V4.1-Flash on huggingface.co
767.9K
Total runs
46.7K
24-hour runs
99.3K
3-day runs
161.8K
7-day runs
767.9K
30-day runs
More Information About DeepSeek-V4.1-Flash huggingface.co Model
DeepSeek-V4.1-Flash huggingface.co is an AI model on huggingface.co that provides DeepSeek-V4.1-Flash's model effect (), which can be used instantly with this deepseek-ai DeepSeek-V4.1-Flash model. huggingface.co supports a free trial of the DeepSeek-V4.1-Flash model, and also provides paid use of the DeepSeek-V4.1-Flash. Support call DeepSeek-V4.1-Flash model through api, including Node.js, Python, http.
DeepSeek-V4.1-Flash huggingface.co is an online trial and call api platform, which integrates DeepSeek-V4.1-Flash's modeling effects, including api services, and provides a free online trial of DeepSeek-V4.1-Flash, you can try DeepSeek-V4.1-Flash online for free by clicking the link below.
deepseek-ai DeepSeek-V4.1-Flash online free url in huggingface.co:
DeepSeek-V4.1-Flash is an open source model from GitHub that offers a free installation service, and any user can find DeepSeek-V4.1-Flash on GitHub to install. At the same time, huggingface.co provides the effect of DeepSeek-V4.1-Flash install, users can directly use DeepSeek-V4.1-Flash installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
DeepSeek-V4.1-Flash install url in huggingface.co: