DFlash
is a novel speculative decoding method that utilizes a lightweight
block diffusion
model for drafting. It enables efficient, high-quality parallel drafting that pushes the limits of inference speed.
This model is the
drafter
component. It must be used in conjunction with the target model
moonshotai/Kimi-K2.5
. It was trained with a context length of 4096 tokens.
Note:
For long-context or agentic usage, consider adding
--speculative-dflash-draft-window-size WINDOW_SIZE
to enable sliding-window attention for the draft model.
Kimi-K2.5-DFlash huggingface.co is an AI model on huggingface.co that provides Kimi-K2.5-DFlash's model effect (), which can be used instantly with this z-lab Kimi-K2.5-DFlash model. huggingface.co supports a free trial of the Kimi-K2.5-DFlash model, and also provides paid use of the Kimi-K2.5-DFlash. Support call Kimi-K2.5-DFlash model through api, including Node.js, Python, http.
Kimi-K2.5-DFlash huggingface.co is an online trial and call api platform, which integrates Kimi-K2.5-DFlash's modeling effects, including api services, and provides a free online trial of Kimi-K2.5-DFlash, you can try Kimi-K2.5-DFlash online for free by clicking the link below.
z-lab Kimi-K2.5-DFlash online free url in huggingface.co:
Kimi-K2.5-DFlash is an open source model from GitHub that offers a free installation service, and any user can find Kimi-K2.5-DFlash on GitHub to install. At the same time, huggingface.co provides the effect of Kimi-K2.5-DFlash install, users can directly use Kimi-K2.5-DFlash installed effect in huggingface.co for debugging and trial. It also supports api for free installation.