This model is a quantized version of
Qwen/Qwen3.5-4B
. Evaluation results and reproduction steps are provided below.
Model Optimizations
This model was obtained by quantizing the weights and activations of
Qwen/Qwen3.5-4B
to FP8 data type, ready for inference with vLLM.
This optimization reduces the model weights from 9.3 GB to 8.0 GB on disk (~14% reduction). Activations are quantized dynamically at inference time using per-tensor scaling, requiring no calibration data.
Only the weights and activations of the linear operators within transformer blocks are quantized using
LLM Compressor
.
This model was evaluated on GSM8k-Platinum, MMLU-Pro, IFEval, Math 500, AIME 2025, and GPQA Diamond using
lm-evaluation-harness
and
lighteval
, with inference served via vLLM.
Accuracy
Category
Benchmark
Qwen/Qwen3.5-4B
RedHatAI/Qwen3.5-4B-FP8-dynamic
Recovery
Instruction Following
GSM8k-Platinum (0-shot)
94.2%
94.5%
100.3%
MMLU-Pro (0-shot)
79.3%
79.1%
99.9%
IFEval — prompt strict (0-shot)
88.0%
88.4%
100.5%
IFEval — instruction strict (0-shot)
91.2%
91.6%
100.4%
Reasoning
Math 500 (0-shot)
84.6%
84.7%
100.2%
AIME 2025 (0-shot)
85.0%
85.0%
100.0%
GPQA Diamond (0-shot)
76.8%
76.3%
99.3%
Reproduction
The results were obtained using the following commands. GSM8k-Platinum, MMLU-Pro, IFEval, Math 500, and GPQA Diamond were each run 3 times with different seeds and results averaged. AIME 2025 was run 8 times. The vLLM server was started with
--language-model-only
for all evaluations.
Qwen3.5-4B-FP8-dynamic huggingface.co is an AI model on huggingface.co that provides Qwen3.5-4B-FP8-dynamic's model effect (), which can be used instantly with this RedHatAI Qwen3.5-4B-FP8-dynamic model. huggingface.co supports a free trial of the Qwen3.5-4B-FP8-dynamic model, and also provides paid use of the Qwen3.5-4B-FP8-dynamic. Support call Qwen3.5-4B-FP8-dynamic model through api, including Node.js, Python, http.
Qwen3.5-4B-FP8-dynamic huggingface.co is an online trial and call api platform, which integrates Qwen3.5-4B-FP8-dynamic's modeling effects, including api services, and provides a free online trial of Qwen3.5-4B-FP8-dynamic, you can try Qwen3.5-4B-FP8-dynamic online for free by clicking the link below.
RedHatAI Qwen3.5-4B-FP8-dynamic online free url in huggingface.co:
Qwen3.5-4B-FP8-dynamic is an open source model from GitHub that offers a free installation service, and any user can find Qwen3.5-4B-FP8-dynamic on GitHub to install. At the same time, huggingface.co provides the effect of Qwen3.5-4B-FP8-dynamic install, users can directly use Qwen3.5-4B-FP8-dynamic installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Qwen3.5-4B-FP8-dynamic install url in huggingface.co: