Mirrored by
AxionML
for open-source serving and deployment use cases. Part of AxionML's effort to provide ready-to-serve quantized models for the community.
Quantized by NVIDIA.
The weights in this repository are an unmodified copy of
nvidia/DeepSeek-V4-Pro-0813-NVFP4
(revision
949138637e8e8fe190335be2396af62187dc8a83
). All credit for the quantization belongs to NVIDIA.
About NVFP4 quantization:
NVFP4 on Blackwell couples a compact E2M1 FP4 codebook with blockwise FP8 (E4M3) scaling over 16-element micro-blocks, so that 4-bit stored values remain numerically useful for neural-network computation. The E2M1 codebook provides a small, nonuniform set of representable magnitudes up to ±6 and relies on saturating behavior rather than IEEE NaN/Inf encodings to maximize usable range per bit. Using an FP8 block scale (rather than power-of-two-only E8M0) enables fractional scales and error-minimizing scale selection. On Blackwell Tensor Cores, native FP4 multipliers exploit E2M1 simplicity while higher-precision FP32 accumulation protects dot-product accuracy.
Ready for commercial and non-commercial use under
MIT
.
Scores reported by NVIDIA for this checkpoint (SGLang, B200). Baseline:
deepseek-ai/DeepSeek-V4-Pro-0813
.
temperature=1.0
,
top_p=1.0
, max reasoning effort.
Quantization Details
Quantization format:
routed MoE experts in NVFP4 (weights and activations); source MXFP4 expert weights are bit-cast losslessly to NVFP4 and only the block scales are rewritten; FP8 attention and shared experts, BF16 norms and embeddings unchanged
vLLM configuration validated upstream on 8x B200 with
vllm/vllm-openai:v0.27.1
and
v0.28.0
. DSpark heads are preserved, but speculative decoding was not validated for this NVFP4 checkpoint.
Limitations
The base model was trained on data that may contain toxic language and societal biases. The quantized model inherits these limitations. It may generate inaccurate, biased, or offensive content. Please refer to the
original model card
and the
upstream quantized model card
for full details.
DeepSeek-V4-Pro-0813-NVFP4 huggingface.co is an AI model on huggingface.co that provides DeepSeek-V4-Pro-0813-NVFP4's model effect (), which can be used instantly with this AxionML DeepSeek-V4-Pro-0813-NVFP4 model. huggingface.co supports a free trial of the DeepSeek-V4-Pro-0813-NVFP4 model, and also provides paid use of the DeepSeek-V4-Pro-0813-NVFP4. Support call DeepSeek-V4-Pro-0813-NVFP4 model through api, including Node.js, Python, http.
DeepSeek-V4-Pro-0813-NVFP4 huggingface.co is an online trial and call api platform, which integrates DeepSeek-V4-Pro-0813-NVFP4's modeling effects, including api services, and provides a free online trial of DeepSeek-V4-Pro-0813-NVFP4, you can try DeepSeek-V4-Pro-0813-NVFP4 online for free by clicking the link below.
AxionML DeepSeek-V4-Pro-0813-NVFP4 online free url in huggingface.co:
DeepSeek-V4-Pro-0813-NVFP4 is an open source model from GitHub that offers a free installation service, and any user can find DeepSeek-V4-Pro-0813-NVFP4 on GitHub to install. At the same time, huggingface.co provides the effect of DeepSeek-V4-Pro-0813-NVFP4 install, users can directly use DeepSeek-V4-Pro-0813-NVFP4 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
DeepSeek-V4-Pro-0813-NVFP4 install url in huggingface.co: