inference-optimization / Qwen3-8B-DFlash-FP8-DYNAMIC

huggingface.co
Total runs: 32
24-hour runs: 4
7-day runs: 32
30-day runs: 32
Model's Last Updated: September 24 2026
text-generation

Introduction of Qwen3-8B-DFlash-FP8-DYNAMIC

Model Details of Qwen3-8B-DFlash-FP8-DYNAMIC

Drift8-FP8-DYNAMIC — Data-free FP8_DYNAMIC quantization

Drift8 is a quantized DFlash drafter for the Qwen3-8B target , derived from RedHatAI/Qwen3-8B-speculator.dflash . This repository contains the drafter component; it is not a standalone chat model.

Variant
  • Quantization: FP8_DYNAMIC (per-channel weights with token-dynamic activations).
  • Calibration: Data-free quantization: no calibration data or calibration samples were used; manifest seed 0. This is the single FP8_DYNAMIC checkpoint used in the evaluation.
  • Calibration seed: 0.
  • Quantization settings and source revisions: quant_run_manifest.json .
Use with vLLM

Pair this drafter with the Qwen3-8B target and a DFlash-capable vLLM build:

vllm serve Qwen/Qwen3-8B \
  --spec-model inference-optimization/Qwen3-8B-DFlash-Drift8-FP8-DYNAMIC \
  --spec-tokens 7 \
  --spec-method dflash

config.py provides the custom drafter configuration. The experiment's serving command and runtime patch are in provenance/evaluation/ .

Reproducibility

The manifests are included at the repository root. provenance/ contains the source drafter's captured train_command.txt , a quantization command explicitly marked as reconstructed, the quantizer and calibration source snapshot, the vLLM command and patch, both target and drafter checkpoint hashes, and the nine per-subset evaluation commands. The selected seed-0 checkpoint is the same checkpoint used in the 2026-09-24 all-subset evaluation.

The PerfectBlend preparation, prompts, and hidden-state cache remain local because the prepared prompts are not redistributable. The cache sample counts and content hashes are recorded in calibration_manifest.json ; no prompts or hidden-state tensors are uploaded.

The source drafter lists Apache-2.0 licensing on its Hugging Face model card .

Runs of inference-optimization Qwen3-8B-DFlash-FP8-DYNAMIC on huggingface.co

32
Total runs
4
24-hour runs
8
3-day runs
32
7-day runs
32
30-day runs

More Information About Qwen3-8B-DFlash-FP8-DYNAMIC huggingface.co Model

More Qwen3-8B-DFlash-FP8-DYNAMIC license Visit here:

https://choosealicense.com/licenses/apache-2.0

Qwen3-8B-DFlash-FP8-DYNAMIC huggingface.co

Qwen3-8B-DFlash-FP8-DYNAMIC huggingface.co is an AI model on huggingface.co that provides Qwen3-8B-DFlash-FP8-DYNAMIC's model effect (), which can be used instantly with this inference-optimization Qwen3-8B-DFlash-FP8-DYNAMIC model. huggingface.co supports a free trial of the Qwen3-8B-DFlash-FP8-DYNAMIC model, and also provides paid use of the Qwen3-8B-DFlash-FP8-DYNAMIC. Support call Qwen3-8B-DFlash-FP8-DYNAMIC model through api, including Node.js, Python, http.

inference-optimization Qwen3-8B-DFlash-FP8-DYNAMIC online free

Qwen3-8B-DFlash-FP8-DYNAMIC huggingface.co is an online trial and call api platform, which integrates Qwen3-8B-DFlash-FP8-DYNAMIC's modeling effects, including api services, and provides a free online trial of Qwen3-8B-DFlash-FP8-DYNAMIC, you can try Qwen3-8B-DFlash-FP8-DYNAMIC online for free by clicking the link below.

inference-optimization Qwen3-8B-DFlash-FP8-DYNAMIC online free url in huggingface.co:

https://huggingface.co/inference-optimization/Qwen3-8B-DFlash-FP8-DYNAMIC

Qwen3-8B-DFlash-FP8-DYNAMIC install

Qwen3-8B-DFlash-FP8-DYNAMIC is an open source model from GitHub that offers a free installation service, and any user can find Qwen3-8B-DFlash-FP8-DYNAMIC on GitHub to install. At the same time, huggingface.co provides the effect of Qwen3-8B-DFlash-FP8-DYNAMIC install, users can directly use Qwen3-8B-DFlash-FP8-DYNAMIC installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Qwen3-8B-DFlash-FP8-DYNAMIC install url in huggingface.co:

https://huggingface.co/inference-optimization/Qwen3-8B-DFlash-FP8-DYNAMIC

Url of Qwen3-8B-DFlash-FP8-DYNAMIC

Provider of Qwen3-8B-DFlash-FP8-DYNAMIC huggingface.co

inference-optimization
ORGANIZATIONS

Other API from inference-optimization