AxionML / DeepSeek-V4.1-Flash-NVFP4

huggingface.co
Total runs: 93
24-hour runs: 4
7-day runs: 93
30-day runs: 93
Model's Last Updated: September 29 2026
image-text-to-text

Introduction of DeepSeek-V4.1-Flash-NVFP4

Model Details of DeepSeek-V4.1-Flash-NVFP4

AxionML DeepSeek-V4.1-Flash-NVFP4

Mirrored by AxionML for open-source serving and deployment use cases. Part of AxionML's effort to provide ready-to-serve quantized models for the community.

Quantized by NVIDIA. The weights in this repository are an unmodified copy of nvidia/DeepSeek-V4.1-Flash-NVFP4 (revision 3431dde3247c13b5957f682b1e3c6fcae2566079 ). All credit for the quantization belongs to NVIDIA.

This is an NVFP4-quantized version of deepseek-ai/DeepSeek-V4.1-Flash (552B backbone plus 196B Engram conditional memory), quantized with NVIDIA Model Optimizer .

About NVFP4 quantization: NVFP4 on Blackwell couples a compact E2M1 FP4 codebook with blockwise FP8 (E4M3) scaling over 16-element micro-blocks, so that 4-bit stored values remain numerically useful for neural-network computation. The E2M1 codebook provides a small, nonuniform set of representable magnitudes up to ±6 and relies on saturating behavior rather than IEEE NaN/Inf encodings to maximize usable range per bit. Using an FP8 block scale (rather than power-of-two-only E8M0) enables fractional scales and error-minimizing scale selection. On Blackwell Tensor Cores, native FP4 multipliers exploit E2M1 simplicity while higher-precision FP32 accumulation protects dot-product accuracy.

Ready for commercial and non-commercial use under MIT .

Model Summary
Architecture Causal Encoder-Decoder MoE with Compressed Sparse Attention 2 ( DeepseekV41ForCausalLM )
Backbone Parameters 552B total
Activated Parameters 8B (prefill) / 16B (decode)
Engram Memory 196B
Experts 384 routed, 6 active, 40 layers
Input Text, image
Context Length 1M tokens
Checkpoint Size ~527 GB
Evaluation Results
Benchmark MXFP4 (source) NVFP4
GPQA Diamond 91.035 91.288
AA-LCR 78.563 78.438
SciCode 54.401 55.843
IFBench 76.667 77.267
MMMU-Pro 74.046 73.699
Terminal-Bench 2.1 81.60 82.16

Scores reported by NVIDIA for this checkpoint (vLLM). Baseline: deepseek-ai/DeepSeek-V4.1-Flash . temperature=1.0 , top_p=0.95 , reasoning_effort=100 .

Quantization Details
  • Quantization format: routed MoE experts ( w1 , w2 , w3 ) converted from source MXFP4 to NVFP4 W4A4 (group size 16); attention, shared experts, vision, Engram tables and MTP/DSpark keep their source precision (incl. MXFP8)
  • Weight conversion: lossless — all 16,986,931,200 weight blocks preserve their dequantized values; only block scales are rewritten
  • Calibration dataset: 1,024 samples from cnn_dailymail and Nemotron-Post-Training-Dataset-v2 (activation scales)
  • Tool: NVIDIA Model Optimizer v0.47.0rc0
Usage
Deploy with SGLang
python -m sglang.launch_server \
    --model-path AxionML/DeepSeek-V4.1-Flash-NVFP4 \
    --tp 4 \
    --context-length 1048576 \
    --reasoning-parser deepseek-v41 \
    --tool-call-parser deepseekv41 \
    --chunked-prefill-size 4096 \
    --max-running-requests 16
Deploy with vLLM
vllm serve AxionML/DeepSeek-V4.1-Flash-NVFP4 \
    --tensor-parallel-size 4 \
    --tokenizer-mode deepseek_v41 \
    --reasoning-parser deepseek_v41 \
    --language-model-only \
    --max-model-len 1048576 \
    --max-num-seqs 32 \
    --max-num-batched-tokens 8192 \
    --enable-chunked-prefill \
    --no-enable-prefix-caching

Validated upstream on 4x GB300 with lmsysorg/sglang:dev-cu13-dsv41 and vllm/vllm-openai:deepseekv41-flash-0909 . Pass a request-level reasoning_effort (e.g. "max" ) to enable thinking in SGLang. DSpark tensors are preserved but speculative decoding was not validated.

Limitations

The base model was trained on data that may contain toxic language and societal biases. The quantized model inherits these limitations. It may generate inaccurate, biased, or offensive content. Please refer to the original model card and the upstream quantized model card for full details.

Credits

Runs of AxionML DeepSeek-V4.1-Flash-NVFP4 on huggingface.co

93
Total runs
4
24-hour runs
93
3-day runs
93
7-day runs
93
30-day runs

More Information About DeepSeek-V4.1-Flash-NVFP4 huggingface.co Model

More DeepSeek-V4.1-Flash-NVFP4 license Visit here:

https://choosealicense.com/licenses/mit

DeepSeek-V4.1-Flash-NVFP4 huggingface.co

DeepSeek-V4.1-Flash-NVFP4 huggingface.co is an AI model on huggingface.co that provides DeepSeek-V4.1-Flash-NVFP4's model effect (), which can be used instantly with this AxionML DeepSeek-V4.1-Flash-NVFP4 model. huggingface.co supports a free trial of the DeepSeek-V4.1-Flash-NVFP4 model, and also provides paid use of the DeepSeek-V4.1-Flash-NVFP4. Support call DeepSeek-V4.1-Flash-NVFP4 model through api, including Node.js, Python, http.

DeepSeek-V4.1-Flash-NVFP4 huggingface.co Url

https://huggingface.co/AxionML/DeepSeek-V4.1-Flash-NVFP4

AxionML DeepSeek-V4.1-Flash-NVFP4 online free

DeepSeek-V4.1-Flash-NVFP4 huggingface.co is an online trial and call api platform, which integrates DeepSeek-V4.1-Flash-NVFP4's modeling effects, including api services, and provides a free online trial of DeepSeek-V4.1-Flash-NVFP4, you can try DeepSeek-V4.1-Flash-NVFP4 online for free by clicking the link below.

AxionML DeepSeek-V4.1-Flash-NVFP4 online free url in huggingface.co:

https://huggingface.co/AxionML/DeepSeek-V4.1-Flash-NVFP4

DeepSeek-V4.1-Flash-NVFP4 install

DeepSeek-V4.1-Flash-NVFP4 is an open source model from GitHub that offers a free installation service, and any user can find DeepSeek-V4.1-Flash-NVFP4 on GitHub to install. At the same time, huggingface.co provides the effect of DeepSeek-V4.1-Flash-NVFP4 install, users can directly use DeepSeek-V4.1-Flash-NVFP4 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

DeepSeek-V4.1-Flash-NVFP4 install url in huggingface.co:

https://huggingface.co/AxionML/DeepSeek-V4.1-Flash-NVFP4

Url of DeepSeek-V4.1-Flash-NVFP4

DeepSeek-V4.1-Flash-NVFP4 huggingface.co Url

Provider of DeepSeek-V4.1-Flash-NVFP4 huggingface.co

AxionML
ORGANIZATIONS

Other API from AxionML

huggingface.co

Total runs: 195
Run Growth: 195
Growth Rate: 100.00%
Updated:September 29 2026