A GGUF conversion of the
DFlash block-diffusion draft
for
MiMo-V2.5-Pro
, built for
speculative decoding in
ik_llama.cpp
. This is the
validation artifact for ik_llama.cpp
PR #2048
(MiMo DFlash draft support).
⚠️
Not a standalone model.
This is a 5-layer draft that shares the target's token embedding
and output head. It only produces meaningful output when launched as
--model-draft
alongside a
full
MiMo-V2.5-Pro
target GGUF. Loading it on its own will not generate coherent text.
Use an
unfused
target GGUF. A fused-MoE MiMo GGUF produced garbled output with this draft in
testing; the unfused tensor layout is required for the draft's captured-layer features to line up.
--spec-type dflash:...
selects the DFlash speculator.
n_max=1
was optimal in testing (the
draft runs autoregressively, so per-step overhead grows with
n_max
).
cross_ctx=16
is a small
ring buffer of recent target positions visible to the draft; larger values collapsed acceptance.
-sm graph --max-gpu 4
(tensor + graph split) was the best multi-GPU profile on 4×H200 NVL.
Pinning the draft to one GPU with
-devd
added overhead and is not recommended under graph split.
Conversion metadata
GGUF key
Value
Note
dflash-draft.block_count
5
5-layer draft
dflash-draft.attention.head_count
/
head_count_kv
128 / 8
GQA
dflash-draft.rope.dimension_count
64
head_dim 128 × partial_rotary_factor 0.5
(upper half is NoPE)
dflash-draft.rope.freq_base
5,000,000
backbone rotary base
dflash-draft.attention.value_scale
0.612
dflash-draft.attention.sliding_window
1024
all-SWA pattern across the 5 layers
IO contract
io=shared-target
token_embd
/
output
reuse the target's tensors
Validation
4×H200 NVL (NVLink), ik_llama.cpp PR #2048. Draft accepted greedily at
n_max=1, cross_ctx=16
,
q4_0 KV on both target and draft.
MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF huggingface.co is an AI model on huggingface.co that provides MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF's model effect (), which can be used instantly with this ji-farthing MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF model. huggingface.co supports a free trial of the MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF model, and also provides paid use of the MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF. Support call MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF model through api, including Node.js, Python, http.
MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF huggingface.co is an online trial and call api platform, which integrates MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF's modeling effects, including api services, and provides a free online trial of MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF, you can try MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF online for free by clicking the link below.
ji-farthing MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF online free url in huggingface.co:
MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF is an open source model from GitHub that offers a free installation service, and any user can find MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF install, users can directly use MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF install url in huggingface.co: