This repository contains the DFlash 2 draft model for
Qwen/Qwen3.8-27B
.
It is not a standalone language model: it runs inside a speculative
decoding server and drafts tokens for the target model to verify. This repository is a mirror
of
incoai/Qwen3.8-27B-DFlash2
.
DFlash 2 is a block-diffusion drafter for speculative decoding. It predicts
a whole block of tokens in a single pass and keeps the top candidates at
every position. A lightweight selector then traces one coherent path through them.
Two-tap dynamic convolutions in the backbone keep the draft from decaying
toward the end of the block. Decoding is lossless: greedy output
matches the target model exactly, and sampling preserves its distribution.
We compare autoregressive decoding, Qwen3.8's built-in seven-token MTP,
a community DSpark drafter
(
RadixArk/Qwen3.8-27B-DSpark
),
and DFlash 2. All speculative methods propose seven draft tokens per
verification step.
Acceptance Length
Acceptance length is the per-request mean of completion tokens divided by verification steps.
Higher is better.
Task
MTP
DSpark
DFlash 2
GSM8K
5.02
4.36
5.46
MATH-500
4.72
3.92
5.28
HumanEval
3.91
3.30
4.39
MBPP
3.99
3.51
4.79
MT-Bench
3.74
3.01
4.10
Throughput
Throughput is total output tokens divided by end-to-end wall time.
Each cell shows
output tok/s (speedup vs. autoregressive)
.
Qwen3.8-27B-DFlash2 huggingface.co is an AI model on huggingface.co that provides Qwen3.8-27B-DFlash2's model effect (), which can be used instantly with this z-lab Qwen3.8-27B-DFlash2 model. huggingface.co supports a free trial of the Qwen3.8-27B-DFlash2 model, and also provides paid use of the Qwen3.8-27B-DFlash2. Support call Qwen3.8-27B-DFlash2 model through api, including Node.js, Python, http.
Qwen3.8-27B-DFlash2 huggingface.co is an online trial and call api platform, which integrates Qwen3.8-27B-DFlash2's modeling effects, including api services, and provides a free online trial of Qwen3.8-27B-DFlash2, you can try Qwen3.8-27B-DFlash2 online for free by clicking the link below.
z-lab Qwen3.8-27B-DFlash2 online free url in huggingface.co:
Qwen3.8-27B-DFlash2 is an open source model from GitHub that offers a free installation service, and any user can find Qwen3.8-27B-DFlash2 on GitHub to install. At the same time, huggingface.co provides the effect of Qwen3.8-27B-DFlash2 install, users can directly use Qwen3.8-27B-DFlash2 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Qwen3.8-27B-DFlash2 install url in huggingface.co: