DFlash speculative decoding draft model for
Qwen/Qwen3-32B-FP8
. Trained using the
DFlash
(Block Diffusion for Flash Speculative Decoding) method from Z-Lab.
Architecture
Parameter
Value
Draft layers
5
Hidden size
5120
Attention heads
32 (8 KV heads, GQA)
Head dim
128
Intermediate size
9728
Block size
16
Target layers captured
[1, 16, 31, 46, 61]
Parameters (draft-only)
~3.2B (bf16)
Tied embeddings
Yes (shared with target)
The draft model takes concatenated hidden states from 5 target model layers as input and predicts a block of 16 tokens in parallel via iterative denoising. Attention is non-causal: queries attend to both target hidden states (context) and noise embeddings (draft tokens).
Qwen3-32B-FP8-DFLASH huggingface.co is an AI model on huggingface.co that provides Qwen3-32B-FP8-DFLASH's model effect (), which can be used instantly with this chutesai Qwen3-32B-FP8-DFLASH model. huggingface.co supports a free trial of the Qwen3-32B-FP8-DFLASH model, and also provides paid use of the Qwen3-32B-FP8-DFLASH. Support call Qwen3-32B-FP8-DFLASH model through api, including Node.js, Python, http.
Qwen3-32B-FP8-DFLASH huggingface.co is an online trial and call api platform, which integrates Qwen3-32B-FP8-DFLASH's modeling effects, including api services, and provides a free online trial of Qwen3-32B-FP8-DFLASH, you can try Qwen3-32B-FP8-DFLASH online for free by clicking the link below.
chutesai Qwen3-32B-FP8-DFLASH online free url in huggingface.co:
Qwen3-32B-FP8-DFLASH is an open source model from GitHub that offers a free installation service, and any user can find Qwen3-32B-FP8-DFLASH on GitHub to install. At the same time, huggingface.co provides the effect of Qwen3-32B-FP8-DFLASH install, users can directly use Qwen3-32B-FP8-DFLASH installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Qwen3-32B-FP8-DFLASH install url in huggingface.co: