chutesai / Qwen3-32B-FP8-DFLASH

huggingface.co
Total runs: 522
24-hour runs: -1
7-day runs: -87
30-day runs: 85
Model's Last Updated: June 27 2026

Introduction of Qwen3-32B-FP8-DFLASH

Model Details of Qwen3-32B-FP8-DFLASH

Qwen3-32B-FP8 DFlash Draft Model

DFlash speculative decoding draft model for Qwen/Qwen3-32B-FP8 . Trained using the DFlash (Block Diffusion for Flash Speculative Decoding) method from Z-Lab.

Architecture
Parameter Value
Draft layers 5
Hidden size 5120
Attention heads 32 (8 KV heads, GQA)
Head dim 128
Intermediate size 9728
Block size 16
Target layers captured [1, 16, 31, 46, 61]
Parameters (draft-only) ~3.2B (bf16)
Tied embeddings Yes (shared with target)

The draft model takes concatenated hidden states from 5 target model layers as input and predicts a block of 16 tokens in parallel via iterative denoising. Attention is non-causal: queries attend to both target hidden states (context) and noise embeddings (draft tokens).

Training
Detail Value
Target model Qwen/Qwen3-32B-FP8
Dataset ~50k multi-turn conversations (ShareGPT, OpenHermes, WildChat)
Max sequence length 2048
Effective batch size 8 sequences/step (DDP across 8 GPUs)
Training steps ~103k (1 epoch)
Hardware 8x NVIDIA H200 141GB
Optimizer AdamW, lr=4.8e-3, cosine schedule
Loss Focal cross-entropy (gamma=7.0)
Precision bf16 (draft), FP8 (target, frozen)
Benchmarks

All benchmarks on a single NVIDIA RTX PRO 6000 Blackwell (98GB VRAM) using SGLang v0.5.13.post1.

ShareGPT (200 prompts, concurrency 8, max 1024 output tokens)
Metric DFlash Vanilla Speedup
Output throughput (tok/s) 423.0 229.1 1.85x
Median TTFT (ms) 88.0 313.2 3.56x
Median ITL (ms) 14.7 34.6 2.35x
Median TPOT (ms) 19.5 34.5 1.77x
Accept length 2.46 — —
Synthetic (random tokens, 100 prompts per config)
Input/Output Concurrency DFlash tok/s Vanilla tok/s Speedup Accept len
128/128 1 54.5 20.7 2.64x 2.15
128/128 8 299.0 205.4 1.46x 2.13
128/128 32 578.3 605.8 0.96x 2.13
512/512 1 71.0 20.9 3.40x 2.47
512/512 8 438.6 215.5 2.04x 2.58
512/512 32 814.3 631.2 1.29x 2.63
1024/1024 1 77.8 — — 2.77
1024/1024 8 478.3 — — 2.83
1024/1024 32 869.0 — — 2.87
2048/256 1 100.9 — — 2.92
2048/256 8 555.0 — — 2.96
2048/256 32 945.5 — — 2.99

Accept length increases with context length (2.13 at 128 tokens to 2.99 at 2048 tokens).

Usage (SGLang)
python -m sglang.launch_server \
    --model-path Qwen/Qwen3-32B-FP8 \
    --speculative-algorithm DFLASH \
    --speculative-draft-model-path chutesai/Qwen3-32B-FP8-DFLASH \
    --speculative-num-draft-tokens 16 \
    --speculative-draft-attention-backend triton \
    --trust-remote-code \
    --mem-fraction-static 0.85 \
    --host 0.0.0.0 --port 30000

Requires SGLang >= v0.5.13 with DFlash support.

Runs of chutesai Qwen3-32B-FP8-DFLASH on huggingface.co

522
Total runs
-1
24-hour runs
-53
3-day runs
-87
7-day runs
85
30-day runs

More Information About Qwen3-32B-FP8-DFLASH huggingface.co Model

More Qwen3-32B-FP8-DFLASH license Visit here:

https://choosealicense.com/licenses/apache-2.0

Qwen3-32B-FP8-DFLASH huggingface.co

Qwen3-32B-FP8-DFLASH huggingface.co is an AI model on huggingface.co that provides Qwen3-32B-FP8-DFLASH's model effect (), which can be used instantly with this chutesai Qwen3-32B-FP8-DFLASH model. huggingface.co supports a free trial of the Qwen3-32B-FP8-DFLASH model, and also provides paid use of the Qwen3-32B-FP8-DFLASH. Support call Qwen3-32B-FP8-DFLASH model through api, including Node.js, Python, http.

Qwen3-32B-FP8-DFLASH huggingface.co Url

https://huggingface.co/chutesai/Qwen3-32B-FP8-DFLASH

chutesai Qwen3-32B-FP8-DFLASH online free

Qwen3-32B-FP8-DFLASH huggingface.co is an online trial and call api platform, which integrates Qwen3-32B-FP8-DFLASH's modeling effects, including api services, and provides a free online trial of Qwen3-32B-FP8-DFLASH, you can try Qwen3-32B-FP8-DFLASH online for free by clicking the link below.

chutesai Qwen3-32B-FP8-DFLASH online free url in huggingface.co:

https://huggingface.co/chutesai/Qwen3-32B-FP8-DFLASH

Qwen3-32B-FP8-DFLASH install

Qwen3-32B-FP8-DFLASH is an open source model from GitHub that offers a free installation service, and any user can find Qwen3-32B-FP8-DFLASH on GitHub to install. At the same time, huggingface.co provides the effect of Qwen3-32B-FP8-DFLASH install, users can directly use Qwen3-32B-FP8-DFLASH installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Qwen3-32B-FP8-DFLASH install url in huggingface.co:

https://huggingface.co/chutesai/Qwen3-32B-FP8-DFLASH

Url of Qwen3-32B-FP8-DFLASH

Qwen3-32B-FP8-DFLASH huggingface.co Url

Provider of Qwen3-32B-FP8-DFLASH huggingface.co

chutesai
ORGANIZATIONS

Other API from chutesai

huggingface.co

Total runs: 324
Run Growth: 284
Growth Rate: 87.65%
Updated:October 15 2025
huggingface.co

Total runs: 24
Run Growth: 5
Growth Rate: 20.83%
Updated:March 18 2025