This model is still under training, and inference engine support may not be fully available yet due to architectural changes, including causal SWA layers.
DFlash
is a novel speculative decoding method that utilizes a lightweight
block diffusion
model for drafting. It enables efficient, high-quality parallel drafting that pushes the limits of inference speed.
This model is the
drafter
component. It must be used in conjunction with the target model
Qwen/Qwen3.6-27B
.
from openai import OpenAI
client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="Qwen/Qwen3.6-27B",
messages=[{"role": "user", "content": "Write a quicksort in Python."}],
max_tokens=4096,
temperature=0.0
)
print(response.choices[0].message.content)
Benchmark Results
N/A
Acknowledgements
Special thanks to
David Wang
for his outstanding engineering support on this project. We are also grateful to
Modal
,
InnoMatrix
, and
Yotta Labs
for providing the compute resources used to train this draft model.
Citation
If you find DFlash useful, please cite our work. To share feedback on DFlash or request new model support, please fill out this form:
DFlash Feedback
.
@article{chen2026dflash,
title = {{DFlash: Block Diffusion for Flash Speculative Decoding}},
author = {Chen, Jian and Liang, Yesheng and Liu, Zhijian},
journal = {arXiv preprint arXiv:2602.06036},
year = {2026}
}
Runs of z-lab Qwen3.6-27B-DFlash on huggingface.co
59.8K
Total runs
-4.4K
24-hour runs
-13.5K
3-day runs
-26.5K
7-day runs
-90.5K
30-day runs
More Information About Qwen3.6-27B-DFlash huggingface.co Model
Qwen3.6-27B-DFlash huggingface.co is an AI model on huggingface.co that provides Qwen3.6-27B-DFlash's model effect (), which can be used instantly with this z-lab Qwen3.6-27B-DFlash model. huggingface.co supports a free trial of the Qwen3.6-27B-DFlash model, and also provides paid use of the Qwen3.6-27B-DFlash. Support call Qwen3.6-27B-DFlash model through api, including Node.js, Python, http.
Qwen3.6-27B-DFlash huggingface.co is an online trial and call api platform, which integrates Qwen3.6-27B-DFlash's modeling effects, including api services, and provides a free online trial of Qwen3.6-27B-DFlash, you can try Qwen3.6-27B-DFlash online for free by clicking the link below.
z-lab Qwen3.6-27B-DFlash online free url in huggingface.co:
Qwen3.6-27B-DFlash is an open source model from GitHub that offers a free installation service, and any user can find Qwen3.6-27B-DFlash on GitHub to install. At the same time, huggingface.co provides the effect of Qwen3.6-27B-DFlash install, users can directly use Qwen3.6-27B-DFlash installed effect in huggingface.co for debugging and trial. It also supports api for free installation.