DeepSeek-V4-Flash-0731
is the official release of
DeepSeek-V4-Flash
, superseding the preview version, with substantially enhanced agentic capabilities. It has the same model structure as
DeepSeek-V4-Flash-DSpark
, i.e. it comes with a speculative decoding module attached.
DeepSeek-V4-Flash-0731 outperforms DeepSeek-V4-Pro (Preview) on benchmarks listed below despite its far smaller activated parameter count, and is broadly competitive with the strongest proprietary models available.
Benchmark
DeepSeek-V4-Flash-0731
DeepSeek-V4-Flash (Preview)
DeepSeek-V4-Pro (Preview)
GLM-5.2
Opus-4.8
Terminal Bench 2.1
82.7
61.8
72.1
81.0
85.0
NL2Repo
54.2
39.4
38.5
48.9
69.7
Cybergym
76.7
38.7
52.7
-
83.1
DeepSWE
54.4
7.3
12.8
46.2
58.0
Toolathlon-Verified
70.3
49.7
55.9
59.9
76.2
Agents' Last Exam
25.2
15.8
16.5
23.8
25.7
AutomationBench Public
25.1
10.8
12.8
12.9
27.2
DSBench-FullStack †
68.7
37.0
41.8
61.8
71.6
DSBench-Hard †
59.6
25.8
31.1
54.5
71.7
Notes:
For the Code Agent tasks among the public benchmarks above, DeepSeek-V4-Flash-0731 is evaluated with the minimal mode of DeepSeek Harness (to be released) as the agent framework, using the
max
reasoning effort level with
temperature = 1.0, top_p = 0.95
.
† DSBench-FullStack is an internal full-stack development test set; DSBench-Hard is an internal test set of difficult coding-agent problems.
Chat Template
This release does not include a Jinja-format chat template. Instead, we provide a dedicated
encoding
folder with Python scripts and test cases demonstrating how to encode messages in OpenAI-compatible format into input strings for the model, and how to parse the model's text output. Please refer to the
encoding
folder for full documentation.
The
reasoning_effort
parameter now supports three levels —
low
,
high
, and
max
— which control how much deliberation the model spends before answering.
For example, the command below serves the model with vLLM on a single 4×GB300 node.
See the
vLLM recipe
for detailed instructions and other hardware configurations.
Please refer to the
inference
folder for detailed instructions on running DeepSeek-V4 locally, including model weight conversion and interactive chat demos.
For local deployment, we recommend setting the sampling parameters to
temperature = 1.0
, with
top_p = 0.95
for agentic scenarios and
top_p = 1.0
otherwise. For the
high
and
max
reasoning effort levels, we recommend a maximum output length of
384K
tokens.
License
This repository and the model weights are licensed under the
MIT License
.
DeepSeek-V4-Flash-0731 huggingface.co is an AI model on huggingface.co that provides DeepSeek-V4-Flash-0731's model effect (), which can be used instantly with this deepseek-ai DeepSeek-V4-Flash-0731 model. huggingface.co supports a free trial of the DeepSeek-V4-Flash-0731 model, and also provides paid use of the DeepSeek-V4-Flash-0731. Support call DeepSeek-V4-Flash-0731 model through api, including Node.js, Python, http.
DeepSeek-V4-Flash-0731 huggingface.co is an online trial and call api platform, which integrates DeepSeek-V4-Flash-0731's modeling effects, including api services, and provides a free online trial of DeepSeek-V4-Flash-0731, you can try DeepSeek-V4-Flash-0731 online for free by clicking the link below.
deepseek-ai DeepSeek-V4-Flash-0731 online free url in huggingface.co:
DeepSeek-V4-Flash-0731 is an open source model from GitHub that offers a free installation service, and any user can find DeepSeek-V4-Flash-0731 on GitHub to install. At the same time, huggingface.co provides the effect of DeepSeek-V4-Flash-0731 install, users can directly use DeepSeek-V4-Flash-0731 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
DeepSeek-V4-Flash-0731 install url in huggingface.co: