DFlash
is a novel speculative decoding method that utilizes a lightweight
block diffusion
model for drafting. It enables efficient, high-quality parallel drafting that pushes the limits of inference speed.
This model is the
drafter
component. It must be used in conjunction with the target model
moonshotai/Kimi-K2.6
.
Tip:
For long-context or agentic workloads, add
--speculative-dflash-draft-window-size WINDOW_SIZE
to enable sliding-window attention for the drafter.
Usage
from openai import OpenAI
client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="moonshotai/Kimi-K2.6",
messages=[{"role": "user", "content": "Write a quicksort in Python."}],
max_tokens=4096,
)
print(response.choices[0].message.content)
Benchmark Results
Acceptance Length
Thinking: enabled
Max new tokens: 4096
Block size: 8
SGLang results.
Dataset
Accept Length
GSM8K
4.8
Math500
4.8
HumanEval
4.7
MBPP
4.2
MT-Bench
3.5
Throughput
Dataset
C=32
GSM8K
2256
Math500
2879
HumanEval
2759
MBPP
2949
MT-Bench
1765
Acknowledgements
Special thanks to
David Wang
for his outstanding engineering support on this project. We are also grateful to
Modal
,
InnoMatrix
, and
Yotta Labs
for providing the compute resources used to train this draft model.
Citation
If you find DFlash useful, please cite our work. To share feedback on DFlash or request new model support, please fill out this form:
DFlash Feedback
.
@article{chen2026dflash,
title = {{DFlash: Block Diffusion for Flash Speculative Decoding}},
author = {Chen, Jian and Liang, Yesheng and Liu, Zhijian},
journal = {arXiv preprint arXiv:2602.06036},
year = {2026}
}
Runs of z-lab Kimi-K2.6-DFlash on huggingface.co
0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
-4
30-day runs
More Information About Kimi-K2.6-DFlash huggingface.co Model
Kimi-K2.6-DFlash huggingface.co is an AI model on huggingface.co that provides Kimi-K2.6-DFlash's model effect (), which can be used instantly with this z-lab Kimi-K2.6-DFlash model. huggingface.co supports a free trial of the Kimi-K2.6-DFlash model, and also provides paid use of the Kimi-K2.6-DFlash. Support call Kimi-K2.6-DFlash model through api, including Node.js, Python, http.
Kimi-K2.6-DFlash huggingface.co is an online trial and call api platform, which integrates Kimi-K2.6-DFlash's modeling effects, including api services, and provides a free online trial of Kimi-K2.6-DFlash, you can try Kimi-K2.6-DFlash online for free by clicking the link below.
z-lab Kimi-K2.6-DFlash online free url in huggingface.co:
Kimi-K2.6-DFlash is an open source model from GitHub that offers a free installation service, and any user can find Kimi-K2.6-DFlash on GitHub to install. At the same time, huggingface.co provides the effect of Kimi-K2.6-DFlash install, users can directly use Kimi-K2.6-DFlash installed effect in huggingface.co for debugging and trial. It also supports api for free installation.