Model size:
14.0 GB (reduced from 19.3 GB in BF16)
Release Date:
2026-05-11
Version:
1.0
Model Developers:
RedHatAI
This model is a quantized version of
Qwen/Qwen3.5-9B
. Evaluation results and reproduction steps are provided below.
Model Optimizations
This model was obtained by quantizing the weights and activations of
Qwen/Qwen3.5-9B
to FP8 data type, ready for inference with vLLM.
This optimization reduces the model weights from 19.3 GB to 14.0 GB on disk (~27% reduction). Activations are quantized dynamically at inference time using per-tensor scaling, requiring no calibration data.
Only the weights and activations of the linear operators within transformer blocks are quantized using
LLM Compressor
.
This model was evaluated on GSM8k-Platinum, MMLU-Pro, IFEval, Math 500, AIME 2025, and GPQA Diamond using
lm-evaluation-harness
and
lighteval
, with inference served via vLLM.
Accuracy
Category
Benchmark
Qwen/Qwen3.5-9B
RedHatAI/Qwen3.5-9B-FP8-dynamic
Recovery
Instruction Following
GSM8k-Platinum (0-shot)
94.7%
94.5%
99.8%
MMLU-Pro (0-shot)
82.5%
82.4%
99.9%
IFEval — prompt strict (0-shot)
90.3%
88.9%
98.4%
IFEval — instruction strict (0-shot)
92.9%
92.0%
99.0%
Reasoning
Math 500 (0-shot)
85.0%
84.7%
99.7%
AIME 2025 (0-shot)
88.3%
87.9%
99.5%
GPQA Diamond (0-shot)
84.0%
83.8%
99.8%
Reproduction
The results were obtained using the following commands. GSM8k-Platinum, MMLU-Pro, IFEval, Math 500, and GPQA Diamond were each run 3 times with different seeds and results averaged. AIME 2025 was run 8 times. The vLLM server was started with
--language-model-only
for all evaluations.
Qwen3.5-9B-FP8-dynamic huggingface.co is an AI model on huggingface.co that provides Qwen3.5-9B-FP8-dynamic's model effect (), which can be used instantly with this RedHatAI Qwen3.5-9B-FP8-dynamic model. huggingface.co supports a free trial of the Qwen3.5-9B-FP8-dynamic model, and also provides paid use of the Qwen3.5-9B-FP8-dynamic. Support call Qwen3.5-9B-FP8-dynamic model through api, including Node.js, Python, http.
Qwen3.5-9B-FP8-dynamic huggingface.co is an online trial and call api platform, which integrates Qwen3.5-9B-FP8-dynamic's modeling effects, including api services, and provides a free online trial of Qwen3.5-9B-FP8-dynamic, you can try Qwen3.5-9B-FP8-dynamic online for free by clicking the link below.
RedHatAI Qwen3.5-9B-FP8-dynamic online free url in huggingface.co:
Qwen3.5-9B-FP8-dynamic is an open source model from GitHub that offers a free installation service, and any user can find Qwen3.5-9B-FP8-dynamic on GitHub to install. At the same time, huggingface.co provides the effect of Qwen3.5-9B-FP8-dynamic install, users can directly use Qwen3.5-9B-FP8-dynamic installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Qwen3.5-9B-FP8-dynamic install url in huggingface.co: