Quantized layers:
All routed + shared MoE expert projections. All other modules (attention, the MoE router gate, norms, embeddings, the output head, and the MTP block) are excluded and kept in original precision.
Weight quantization:
OCP MXFP4, Static
Activation quantization:
OCP MXFP4, Dynamic
Model Quantization
Quantized from
deepseek-ai/DeepSeek-V4-Pro
with
AMD Quark
. The pipeline
re-quantizes only the MoE expert weights and activations to MXFP4. All non-expert
modules are kept as-is via the exclude list.
Quantization script
from quark.torch import ModelQuantizer
from quark.torch.quantization.config.template import LLMTemplate
template = LLMTemplate.get('deepseek_v4')
qconfig = template.get_config(scheme='mxfp4')
ModelQuantizer(qconfig).direct_quantize_checkpoint(
pretrained_model_path='<DSV4_Pro_src_path>',
save_path='<output_dir>',
keep_excluded_layers_as_original_model_state=True,
)
Deployment
Use with vLLM
This model can be deployed efficiently using the
vLLM
backend based on the Docker image
rocm/vllm-dev:nightly_main_20260714
. vLLM and lm_eval are both installed from source.
Evaluation
The model was evaluated on gsm8k (8-shot) benchmark using the vLLM framework.
Accuracy
Benchmark
deepseek-ai/DeepSeek-V4-Pro
amd/DeepSeek-V4-Pro-MXFP4
Recovery
GSM8K (strict-match)
94.90
93.6
99.0%
Reproduction
The GSM8K results were obtained using the lm-eval framework, based on the Docker image
rocm/vllm-dev:nightly_main_20260714
.
This model is a quantized derivative of
deepseek-ai/DeepSeek-V4-Pro
and is distributed under the same license as the source model: the
MIT License
.
A copy of the upstream LICENSE is included in this repository.
Modifications Copyright (c) 2026 Advanced Micro Devices, Inc. All rights reserved.
AMD has modified the model weights of the MoE expert layers by quantizing them to
MXFP4 with AMD Quark; the modifications are provided under the same MIT License and
are not subject to any separate or different license.
Runs of amd DeepSeek-V4-Pro-MXFP4 on huggingface.co
2.8K
Total runs
-88
24-hour runs
-92
3-day runs
-1.1K
7-day runs
1.7K
30-day runs
More Information About DeepSeek-V4-Pro-MXFP4 huggingface.co Model
DeepSeek-V4-Pro-MXFP4 huggingface.co is an AI model on huggingface.co that provides DeepSeek-V4-Pro-MXFP4's model effect (), which can be used instantly with this amd DeepSeek-V4-Pro-MXFP4 model. huggingface.co supports a free trial of the DeepSeek-V4-Pro-MXFP4 model, and also provides paid use of the DeepSeek-V4-Pro-MXFP4. Support call DeepSeek-V4-Pro-MXFP4 model through api, including Node.js, Python, http.
DeepSeek-V4-Pro-MXFP4 huggingface.co is an online trial and call api platform, which integrates DeepSeek-V4-Pro-MXFP4's modeling effects, including api services, and provides a free online trial of DeepSeek-V4-Pro-MXFP4, you can try DeepSeek-V4-Pro-MXFP4 online for free by clicking the link below.
amd DeepSeek-V4-Pro-MXFP4 online free url in huggingface.co:
DeepSeek-V4-Pro-MXFP4 is an open source model from GitHub that offers a free installation service, and any user can find DeepSeek-V4-Pro-MXFP4 on GitHub to install. At the same time, huggingface.co provides the effect of DeepSeek-V4-Pro-MXFP4 install, users can directly use DeepSeek-V4-Pro-MXFP4 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
DeepSeek-V4-Pro-MXFP4 install url in huggingface.co: