Qwen3.5 27B DFlash
, the DFlash speculative-decoding
draft
converted to GGUF by
Atomic Chat
. Built straight from
z-lab
's original weights. Runs fully offline.
What this is
DFlash
is a speculative-decoding method that drafts a whole
block
of candidate tokens in a single forward pass using a lightweight block-diffusion model, instead of one token at a time. This repo is the
draft component only
— it does nothing on its own. You run it alongside the target model
Qwen/Qwen3.5-27B
, which verifies the drafted block and keeps the longest correct prefix. Output is identical to running the target alone, just faster.
These GGUFs are
converted from z-lab's original weights
, not a repack of someone else's GGUF. The draft attaches to any GGUF of the target model (Atomic, unsloth, bartowski, ...).
Run in llama.cpp
Needs a build of
llama.cpp
with DFlash speculative decoding (PR #22105). You supply the target as
-m
and this draft as
-md
:
DFlash is trained for
non-thinking
generation — pass
enable_thinking=false
in the chat template for best acceptance.
Choosing a quant
Quant
Size
Notes
Q8_0
2.27 GB
Recommended. Near-lossless draft head, small and fast to draft with.
Performance
z-lab report up to
6.17x
lossless acceleration on their reference stack (vLLM / SGLang / Transformers). In
llama.cpp
today the DFlash port is newer: in our tests
dense
targets get roughly
1.8x-2.8x
end-to-end on code generation, and acceptance climbs on larger targets and structured/code output. Acceptance and speedup depend on the target and the content, not on the quantization. Speedups shrink on free-form prose and on small-active MoE targets.
License
Released by z-lab under the MIT license. Converted to GGUF by Atomic Chat. See the
DFlash paper
and
project page
.
Runs of AtomicChat Qwen3.5-27B-DFlash-GGUF on huggingface.co
124
Total runs
4
24-hour runs
10
3-day runs
11
7-day runs
-228
30-day runs
More Information About Qwen3.5-27B-DFlash-GGUF huggingface.co Model
Qwen3.5-27B-DFlash-GGUF huggingface.co is an AI model on huggingface.co that provides Qwen3.5-27B-DFlash-GGUF's model effect (), which can be used instantly with this AtomicChat Qwen3.5-27B-DFlash-GGUF model. huggingface.co supports a free trial of the Qwen3.5-27B-DFlash-GGUF model, and also provides paid use of the Qwen3.5-27B-DFlash-GGUF. Support call Qwen3.5-27B-DFlash-GGUF model through api, including Node.js, Python, http.
Qwen3.5-27B-DFlash-GGUF huggingface.co is an online trial and call api platform, which integrates Qwen3.5-27B-DFlash-GGUF's modeling effects, including api services, and provides a free online trial of Qwen3.5-27B-DFlash-GGUF, you can try Qwen3.5-27B-DFlash-GGUF online for free by clicking the link below.
AtomicChat Qwen3.5-27B-DFlash-GGUF online free url in huggingface.co:
Qwen3.5-27B-DFlash-GGUF is an open source model from GitHub that offers a free installation service, and any user can find Qwen3.5-27B-DFlash-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of Qwen3.5-27B-DFlash-GGUF install, users can directly use Qwen3.5-27B-DFlash-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Qwen3.5-27B-DFlash-GGUF install url in huggingface.co: