AtomicChat / Qwen3.5-27B-DFlash-GGUF

huggingface.co
Total runs: 124
24-hour runs: 4
7-day runs: 11
30-day runs: -228
Model's Last Updated: July 23 2026
text-generation

Introduction of Qwen3.5-27B-DFlash-GGUF

Model Details of Qwen3.5-27B-DFlash-GGUF


DFlash

Qwen3.5 27B DFlash , the DFlash speculative-decoding draft converted to GGUF by Atomic Chat . Built straight from z-lab 's original weights. Runs fully offline.

What this is

DFlash is a speculative-decoding method that drafts a whole block of candidate tokens in a single forward pass using a lightweight block-diffusion model, instead of one token at a time. This repo is the draft component only — it does nothing on its own. You run it alongside the target model Qwen/Qwen3.5-27B , which verifies the drafted block and keeps the longest correct prefix. Output is identical to running the target alone, just faster.

These GGUFs are converted from z-lab's original weights , not a repack of someone else's GGUF. The draft attaches to any GGUF of the target model (Atomic, unsloth, bartowski, ...).

Run in llama.cpp

Needs a build of llama.cpp with DFlash speculative decoding (PR #22105). You supply the target as -m and this draft as -md :

./llama-server \
    -m   Qwen3.5-27B.gguf \
    -md  Qwen3.5-27B-DFlash.Q8_0.gguf \
    --spec-type draft-dflash --spec-draft-n-max 15 \
    -ngl 99 -fa on --jinja -c 8192

DFlash is trained for non-thinking generation — pass enable_thinking=false in the chat template for best acceptance.

Choosing a quant
Quant Size Notes
Q8_0 2.27 GB Recommended. Near-lossless draft head, small and fast to draft with.
Performance
Qwen3.5-27B-DFlash DFlash speedup

z-lab report up to 6.17x lossless acceleration on their reference stack (vLLM / SGLang / Transformers). In llama.cpp today the DFlash port is newer: in our tests dense targets get roughly 1.8x-2.8x end-to-end on code generation, and acceptance climbs on larger targets and structured/code output. Acceptance and speedup depend on the target and the content, not on the quantization. Speedups shrink on free-form prose and on small-active MoE targets.

License

Released by z-lab under the MIT license. Converted to GGUF by Atomic Chat. See the DFlash paper and project page .

Runs of AtomicChat Qwen3.5-27B-DFlash-GGUF on huggingface.co

124
Total runs
4
24-hour runs
10
3-day runs
11
7-day runs
-228
30-day runs

More Information About Qwen3.5-27B-DFlash-GGUF huggingface.co Model

More Qwen3.5-27B-DFlash-GGUF license Visit here:

https://choosealicense.com/licenses/apache-2.0

Qwen3.5-27B-DFlash-GGUF huggingface.co

Qwen3.5-27B-DFlash-GGUF huggingface.co is an AI model on huggingface.co that provides Qwen3.5-27B-DFlash-GGUF's model effect (), which can be used instantly with this AtomicChat Qwen3.5-27B-DFlash-GGUF model. huggingface.co supports a free trial of the Qwen3.5-27B-DFlash-GGUF model, and also provides paid use of the Qwen3.5-27B-DFlash-GGUF. Support call Qwen3.5-27B-DFlash-GGUF model through api, including Node.js, Python, http.

Qwen3.5-27B-DFlash-GGUF huggingface.co Url

https://huggingface.co/AtomicChat/Qwen3.5-27B-DFlash-GGUF

AtomicChat Qwen3.5-27B-DFlash-GGUF online free

Qwen3.5-27B-DFlash-GGUF huggingface.co is an online trial and call api platform, which integrates Qwen3.5-27B-DFlash-GGUF's modeling effects, including api services, and provides a free online trial of Qwen3.5-27B-DFlash-GGUF, you can try Qwen3.5-27B-DFlash-GGUF online for free by clicking the link below.

AtomicChat Qwen3.5-27B-DFlash-GGUF online free url in huggingface.co:

https://huggingface.co/AtomicChat/Qwen3.5-27B-DFlash-GGUF

Qwen3.5-27B-DFlash-GGUF install

Qwen3.5-27B-DFlash-GGUF is an open source model from GitHub that offers a free installation service, and any user can find Qwen3.5-27B-DFlash-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of Qwen3.5-27B-DFlash-GGUF install, users can directly use Qwen3.5-27B-DFlash-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Qwen3.5-27B-DFlash-GGUF install url in huggingface.co:

https://huggingface.co/AtomicChat/Qwen3.5-27B-DFlash-GGUF

Url of Qwen3.5-27B-DFlash-GGUF

Qwen3.5-27B-DFlash-GGUF huggingface.co Url

Provider of Qwen3.5-27B-DFlash-GGUF huggingface.co

AtomicChat
ORGANIZATIONS

Other API from AtomicChat

huggingface.co

Total runs: 395
Run Growth: -576
Growth Rate: -145.82%
Updated:July 23 2026
huggingface.co

Total runs: 31
Run Growth: -40
Growth Rate: -125.00%
Updated:July 28 2026