amd / DeepSeek-V4-Pro-MXFP4

huggingface.co
Total runs: 2.8K
24-hour runs: -88
7-day runs: -1.1K
30-day runs: 1.7K
Model's Last Updated: September 18 2026
text-generation

Introduction of DeepSeek-V4-Pro-MXFP4

Model Details of DeepSeek-V4-Pro-MXFP4

DeepSeek-V4-Pro-MXFP4

Model Overview
  • Model Architecture: DeepseekV4ForCausalLM
    • Input: Text
    • Output: Text
  • Supported Hardware Microarchitecture: AMD MI355 / MI350 (gfx950)
  • ROCm: 7.2.0
  • PyTorch: 2.9.1
  • Transformers: 5.13.1
  • Operating System(s): Linux
  • Inference Engine: vLLM
  • Model Optimizer: AMD-Quark (v0.12.0)
    • Quantized layers: All routed + shared MoE expert projections. All other modules (attention, the MoE router gate, norms, embeddings, the output head, and the MTP block) are excluded and kept in original precision.
    • Weight quantization: OCP MXFP4, Static
    • Activation quantization: OCP MXFP4, Dynamic
Model Quantization

Quantized from deepseek-ai/DeepSeek-V4-Pro with AMD Quark . The pipeline re-quantizes only the MoE expert weights and activations to MXFP4. All non-expert modules are kept as-is via the exclude list.

Quantization script
from quark.torch import ModelQuantizer
from quark.torch.quantization.config.template import LLMTemplate

template = LLMTemplate.get('deepseek_v4')
qconfig = template.get_config(scheme='mxfp4')
ModelQuantizer(qconfig).direct_quantize_checkpoint(
    pretrained_model_path='<DSV4_Pro_src_path>',
    save_path='<output_dir>',
    keep_excluded_layers_as_original_model_state=True,
)
Deployment
Use with vLLM

This model can be deployed efficiently using the vLLM backend based on the Docker image rocm/vllm-dev:nightly_main_20260714 . vLLM and lm_eval are both installed from source.

Evaluation

The model was evaluated on gsm8k (8-shot) benchmark using the vLLM framework.

Accuracy
Benchmark deepseek-ai/DeepSeek-V4-Pro amd/DeepSeek-V4-Pro-MXFP4 Recovery
GSM8K (strict-match) 94.90 93.6 99.0%
Reproduction

The GSM8K results were obtained using the lm-eval framework, based on the Docker image rocm/vllm-dev:nightly_main_20260714 .

Launching server
export VLLM_ROCM_USE_AITER=1
export VLLM_ROCM_USE_AITER_FUSION_SHARED_EXPERTS=1
vllm serve amd/DeepSeek-V4-Pro-MXFP4 --tensor-parallel-size 4 --kv-cache-dtype fp8 \
  --trust-remote-code --tokenizer-mode deepseek_v4 --reasoning-parser deepseek_v4 \
  --tool-call-parser deepseek_v4 --enable-auto-tool-choice \
  --compilation-config '{"mode": 3, "cudagraph_mode": "FULL_DECODE_ONLY"}'
Evaluating model in a new terminal
lm_eval --model local-completions \
    --model_args model=amd/DeepSeek-V4-Pro-MXFP4,base_url=http://localhost:30000/v1/completions,tokenized_requests=False,num_concurrent=32 \
    --tasks gsm8k --batch_size auto --num_fewshot 8
License

This model is a quantized derivative of deepseek-ai/DeepSeek-V4-Pro and is distributed under the same license as the source model: the MIT License . A copy of the upstream LICENSE is included in this repository.

Modifications Copyright (c) 2026 Advanced Micro Devices, Inc. All rights reserved. AMD has modified the model weights of the MoE expert layers by quantizing them to MXFP4 with AMD Quark; the modifications are provided under the same MIT License and are not subject to any separate or different license.

Runs of amd DeepSeek-V4-Pro-MXFP4 on huggingface.co

2.8K
Total runs
-88
24-hour runs
-92
3-day runs
-1.1K
7-day runs
1.7K
30-day runs

More Information About DeepSeek-V4-Pro-MXFP4 huggingface.co Model

More DeepSeek-V4-Pro-MXFP4 license Visit here:

https://choosealicense.com/licenses/mit

DeepSeek-V4-Pro-MXFP4 huggingface.co

DeepSeek-V4-Pro-MXFP4 huggingface.co is an AI model on huggingface.co that provides DeepSeek-V4-Pro-MXFP4's model effect (), which can be used instantly with this amd DeepSeek-V4-Pro-MXFP4 model. huggingface.co supports a free trial of the DeepSeek-V4-Pro-MXFP4 model, and also provides paid use of the DeepSeek-V4-Pro-MXFP4. Support call DeepSeek-V4-Pro-MXFP4 model through api, including Node.js, Python, http.

DeepSeek-V4-Pro-MXFP4 huggingface.co Url

https://huggingface.co/amd/DeepSeek-V4-Pro-MXFP4

amd DeepSeek-V4-Pro-MXFP4 online free

DeepSeek-V4-Pro-MXFP4 huggingface.co is an online trial and call api platform, which integrates DeepSeek-V4-Pro-MXFP4's modeling effects, including api services, and provides a free online trial of DeepSeek-V4-Pro-MXFP4, you can try DeepSeek-V4-Pro-MXFP4 online for free by clicking the link below.

amd DeepSeek-V4-Pro-MXFP4 online free url in huggingface.co:

https://huggingface.co/amd/DeepSeek-V4-Pro-MXFP4

DeepSeek-V4-Pro-MXFP4 install

DeepSeek-V4-Pro-MXFP4 is an open source model from GitHub that offers a free installation service, and any user can find DeepSeek-V4-Pro-MXFP4 on GitHub to install. At the same time, huggingface.co provides the effect of DeepSeek-V4-Pro-MXFP4 install, users can directly use DeepSeek-V4-Pro-MXFP4 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

DeepSeek-V4-Pro-MXFP4 install url in huggingface.co:

https://huggingface.co/amd/DeepSeek-V4-Pro-MXFP4

Url of DeepSeek-V4-Pro-MXFP4

DeepSeek-V4-Pro-MXFP4 huggingface.co Url

Provider of DeepSeek-V4-Pro-MXFP4 huggingface.co

amd
ORGANIZATIONS

Other API from amd

huggingface.co

Total runs: 234.7K
Run Growth: 145.1K
Growth Rate: 64.60%
Updated:April 14 2026
huggingface.co

Total runs: 116.3K
Run Growth: 6.5K
Growth Rate: 5.43%
Updated:July 17 2026
huggingface.co

Total runs: 91.5K
Run Growth: 9.3K
Growth Rate: 9.42%
Updated:June 19 2026
huggingface.co

Total runs: 52.7K
Run Growth: -12.9K
Growth Rate: -24.62%
Updated:July 01 2026
huggingface.co

Total runs: 21.1K
Run Growth: 8.7K
Growth Rate: 40.11%
Updated:July 27 2026
huggingface.co

Total runs: 11.0K
Run Growth: 1.3K
Growth Rate: 11.79%
Updated:August 12 2025
huggingface.co

Total runs: 6.8K
Run Growth: -7.0K
Growth Rate: -100.32%
Updated:October 09 2024
huggingface.co

Total runs: 2.8K
Run Growth: 850
Growth Rate: 30.36%
Updated:June 19 2026
huggingface.co

Total runs: 1.3K
Run Growth: -1.9K
Growth Rate: -146.41%
Updated:May 27 2026
huggingface.co

Total runs: 1.1K
Run Growth: -31.8K
Growth Rate: -2802.64%
Updated:June 19 2026
huggingface.co

Total runs: 608
Run Growth: 245
Growth Rate: 44.14%
Updated:November 03 2025
huggingface.co

Total runs: 544
Run Growth: -3.7K
Growth Rate: -692.02%
Updated:June 19 2026
huggingface.co

Total runs: 542
Run Growth: -384
Growth Rate: -71.24%
Updated:November 15 2025