ji-farthing / MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF

huggingface.co
Total runs: 1.3K
24-hour runs: 0
7-day runs: 993
30-day runs: 906
Model's Last Updated: July 31 2026
text-generation

Introduction of MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF

Model Details of MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF

MiMo-V2.5-Pro DFlash Draft — BF16 GGUF (ik_llama.cpp)

A GGUF conversion of the DFlash block-diffusion draft for MiMo-V2.5-Pro , built for speculative decoding in ik_llama.cpp . This is the validation artifact for ik_llama.cpp PR #2048 (MiMo DFlash draft support).

⚠️ Not a standalone model. This is a 5-layer draft that shares the target's token embedding and output head. It only produces meaningful output when launched as --model-draft alongside a full MiMo-V2.5-Pro target GGUF. Loading it on its own will not generate coherent text.

Files
File Size Precision SHA256
mimo-v25-pro-dflash-draft-bf16-rope5m-vscale-rope64.gguf 5.5 GB BF16 75294724ae22df8ff2d10a390e631b86bcafb34479a4b9fc8866a2a07f69a183
Matching target

Pair this draft with a MiMo-V2.5-Pro target GGUF. Validated against:

Use an unfused target GGUF. A fused-MoE MiMo GGUF produced garbled output with this draft in testing; the unfused tensor layout is required for the draft's captured-layer features to line up.

Usage (ik_llama.cpp)
llama-server \
  -m   MiMo-V2.5-Pro-IQ3_S-unfused.gguf \
  --model-draft mimo-v25-pro-dflash-draft-bf16-rope5m-vscale-rope64.gguf \
  -ngl 999 -ngld 999 \
  -sm graph --max-gpu 4 \
  -fa on -ctk q4_0 -ctv q4_0 -ctkd q4_0 -ctvd q4_0 \
  --spec-type dflash:n_max=1,p_min=0.0,cross_ctx=16 \
  -c 8192 --host 127.0.0.1 --port 8080
  • --spec-type dflash:... selects the DFlash speculator. n_max=1 was optimal in testing (the draft runs autoregressively, so per-step overhead grows with n_max ). cross_ctx=16 is a small ring buffer of recent target positions visible to the draft; larger values collapsed acceptance.
  • -sm graph --max-gpu 4 (tensor + graph split) was the best multi-GPU profile on 4×H200 NVL. Pinning the draft to one GPU with -devd added overhead and is not recommended under graph split.
Conversion metadata
GGUF key Value Note
dflash-draft.block_count 5 5-layer draft
dflash-draft.attention.head_count / head_count_kv 128 / 8 GQA
dflash-draft.rope.dimension_count 64 head_dim 128 × partial_rotary_factor 0.5 (upper half is NoPE)
dflash-draft.rope.freq_base 5,000,000 backbone rotary base
dflash-draft.attention.value_scale 0.612
dflash-draft.attention.sliding_window 1024 all-SWA pattern across the 5 layers
IO contract io=shared-target token_embd / output reuse the target's tensors
Validation

4×H200 NVL (NVLink), ik_llama.cpp PR #2048. Draft accepted greedily at n_max=1, cross_ctx=16 , q4_0 KV on both target and draft.

Prompt Draft acceptance DFlash decode (tok/s) No-spec decode (tok/s)
double-link-list fixture 54.6% 55.6 59.9
quick-sort fixture 60.4% 59.4 60.8
Conversion notes
  • Source draft weights: XiaomiMiMo/MiMo-V2.5-Pro-FP4-DFlash ( dflash/ subdir, BF16 drafter).
  • Vocab/tokenizer: shared from the MiMo-V2.5-Pro target (the draft carries no embeddings of its own).
  • Converter: convert_hf_to_gguf.py on the filament/mimo-dflash-draft-convert branch (ik_llama.cpp PR #2048), DFlashDraftModel path, --outtype bf16 .
python convert_hf_to_gguf.py dflash \
  --target-model-dir . \
  --outtype bf16 \
  --outfile mimo-v25-pro-dflash-draft-bf16-rope5m-vscale-rope64.gguf

Runs of ji-farthing MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF on huggingface.co

1.3K
Total runs
0
24-hour runs
11
3-day runs
993
7-day runs
906
30-day runs

More Information About MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF huggingface.co Model

More MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF license Visit here:

https://choosealicense.com/licenses/mit

MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF huggingface.co

MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF huggingface.co is an AI model on huggingface.co that provides MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF's model effect (), which can be used instantly with this ji-farthing MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF model. huggingface.co supports a free trial of the MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF model, and also provides paid use of the MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF. Support call MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF model through api, including Node.js, Python, http.

MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF huggingface.co Url

https://huggingface.co/ji-farthing/MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF

ji-farthing MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF online free

MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF huggingface.co is an online trial and call api platform, which integrates MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF's modeling effects, including api services, and provides a free online trial of MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF, you can try MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF online for free by clicking the link below.

ji-farthing MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF online free url in huggingface.co:

https://huggingface.co/ji-farthing/MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF

MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF install

MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF is an open source model from GitHub that offers a free installation service, and any user can find MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF install, users can directly use MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF install url in huggingface.co:

https://huggingface.co/ji-farthing/MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF

Url of MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF

MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF huggingface.co Url

Provider of MiMo-V2.5-Pro-DFlash-draft-ik-llama-GGUF huggingface.co

ji-farthing
ORGANIZATIONS

Other API from ji-farthing