AxionML / MiMo-V2.6-Flash-MOPD-NVFP4

huggingface.co
Total runs: 198
24-hour runs: 43
7-day runs: 198
30-day runs: 198
Model's Last Updated: September 30 2026
image-text-to-text

Introduction of MiMo-V2.6-Flash-MOPD-NVFP4

Model Details of MiMo-V2.6-Flash-MOPD-NVFP4

AxionML MiMo-V2.6-Flash-MOPD-NVFP4

Developed by AxionML for open-source serving and deployment use cases. Part of AxionML's effort to provide ready-to-serve quantized models for the community.

NVFP4 version of XiaomiMiMo/MiMo-V2.6-Flash-MOPD for Blackwell. The routed experts run as W4A4 NVFP4 on native FP4 Tensor Cores, and their weights are a bit-exact transcode of Xiaomi's released MXFP4 experts: every E2M1 code is kept byte-for-byte and each E8M0 block scale is re-expressed exactly as two E4M3 NVFP4 block scales. All 9,462,349,824 expert scale blocks convert exactly, and none are re-rounded. Every other tensor is copied byte-for-byte from the source checkpoint. The checkpoint is 175 GB (source: 178 GB).

It serves on the stock lmsysorg/sglang:v0.5.20-cu130 image on B300. On that image we could not serve the released MXFP4 checkpoint on B300: the default MoE backend routes MXFP4 experts into an FP8-only kernel, the flashinfer_mxfp4 backend rejects MiMo's router output, and the Marlin MXFP4 path is SM90/SM120-only, and the Triton path fails a hidden-size check.

About NVFP4 quantization: NVFP4 on Blackwell couples a compact E2M1 FP4 codebook with blockwise FP8 (E4M3) scaling over 16-element micro-blocks, so that 4-bit stored values remain numerically useful for neural-network computation. The E2M1 codebook provides a small, nonuniform set of representable magnitudes up to ±6 and relies on saturating behavior rather than IEEE NaN/Inf encodings to maximize usable range per bit. Using an FP8 block scale (rather than power-of-two-only E8M0) enables fractional scales and error-minimizing scale selection. On Blackwell Tensor Cores, native FP4 multipliers exploit E2M1 simplicity to reduce multiplier area while higher-precision FP32 accumulation protects dot-product accuracy.

Ready for commercial and non-commercial use under the MIT License (inherited from the base model).

Quantization Details
Routed experts (47 MoE layers × 256 experts × gate/up/down) NVFP4 W4A4. Weights: exact transcode of the released MXFP4 (E2M1 codes unchanged; per-tensor weight_scale_2 = 2^m , E4M3 block scale 2^(k-m) per 16 elements). Activations: NVFP4 with a static per-tensor global scale of 1.0 ( input_scale = 1.0 , amax 6 × 448 = 2688)
Fused attention qkv_proj , dense layer-0 MLP, MTP layers FP8 block (128 × 128), unchanged from the source ( FP8_PB_WO in the ModelOpt config)
o_proj , embeddings, lm_head , router, vision and audio encoders, DFlash draft BF16, unchanged from the source
KV-cache BF16 (not quantized)
Config ModelOpt MIXED_PRECISION ( quantized_layers in config.json / hf_quant_config.json )
Checkpoint size 175 GB (source MXFP4/FP8 checkpoint: 178 GB)
Target hardware Blackwell (verified on 4× B300, sm_103)

Why a unit activation scale. (Measured on MiMo-V2.6-Flash-RL, which has the same architecture; this checkpoint uses the same recipe.) Calibrating expert activations on the dequantized model (agentic-coding, diverse and long-reasoning text plus VQA; ModelOpt max calibration) shows benign expert inputs (amax ≤ 183) but extreme outliers at the down-projection input of the last layers (amax 2,960 at layer 44, 18,560 at layer 46, 311,296 at layer 47). SGLang's NVFP4 MoE kernels use one activation global scale per layer, so a calibrated scale that covers those outliers wastes E4M3 range for every other token. We measured three choices against a BF16-activation reference (same bit-exact weights) on 26K sampled reasoning tokens: raw calibrated scales (+0.0127 NLL/token), calibrated with the down-projection input capped at 2688 (+0.0121), and a unit scale everywhere (+0.0120). They are within noise of each other; we ship the unit scale because it needs no calibration data and measured best.

Usage
Deploy with SGLang

Verified on lmsysorg/sglang:v0.5.20-cu130 , 4× B300:

docker run --gpus all --shm-size=64g --network=host --ipc=host \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  lmsysorg/sglang:v0.5.20-cu130 \
  sglang serve \
    --trust-remote-code \
    --model-path AxionML/MiMo-V2.6-Flash-MOPD-NVFP4 \
    --tp 4 \
    --attention-backend fa4 \
    --mm-attention-backend fa4 \
    --moe-runner-backend flashinfer_cutlass \
    --mem-fraction-static 0.85 \
    --reasoning-parser mimo \
    --tool-call-parser mimo \
    --host 0.0.0.0 --port 30000
  • Attention TP must be 4. The fused qkv_proj is TP=4-interleaved, as in the source checkpoint. For 8 GPUs the SGLang MiMo cookbook uses --tp 8 --dp 2 --enable-dp-attention ; we have only verified TP4 with this checkpoint.
  • --attention-backend fa4 is required on Blackwell (asymmetric 192/128 K/V head dims); --mm-attention-backend fa4 for the vision encoder.
  • --moe-runner-backend flashinfer_cutlass is required: on v0.5.20 the flashinfer_trtllm NVFP4 MoE path rejects this model.
  • Sampling: Xiaomi recommends temperature=1.0, top_p=0.95 . Thinking is on by default; pass chat_template_kwargs={"enable_thinking": false} to turn it off.
  • BF16-activation mode. The same checkpoint also runs with BF16 activations (weight-only FP4, the reference column below): add -e SGLANG_FLASHINFER_CUTEDSL_NVFP4_W4A16=1 to docker run and replace the MoE flags with --quantization modelopt_mixed --moe-runner-backend flashinfer_cutedsl .
from openai import OpenAI

client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
r = client.chat.completions.create(
    model="AxionML/MiMo-V2.6-Flash-MOPD-NVFP4",
    messages=[{"role": "user", "content": "What is 15% of 240?"}],
    temperature=1.0, top_p=0.95, max_tokens=4096,
)
print(r.choices[0].message.reasoning_content)
print(r.choices[0].message.content)
Accuracy

Both columns use this checkpoint's bit-exact expert weights on the same stack ( lmsysorg/sglang:v0.5.20-cu130 , 4× B300, sgl-eval 0.1.2, thinking on, temperature=1.0, top_p=0.95 ). The reference runs the experts with BF16 activations (CuTe DSL W4A16 path); the NVFP4 column is the default W4A4 deployment above.

Benchmark Budget BF16 activations (reference) NVFP4 W4A4 (this repo)
GSM8K (1319, pass@1) 16K 96.89 96.44
GPQA-Diamond (avg of 4) 32K 77.78 77.78
AIME 2025 (avg of 8) 32K 75.83 78.75
MMMU-Pro (standard, 10 options) 16K 41.10 39.65
Tool call (get_weather, parsed args) 4K pass pass
Image (shape + color identification) 4K pass pass

Share of generations hitting the budget, reference / NVFP4: GPQA 13.9% / 15.3%, AIME 23.8% / 20.4%, MMMU-Pro 17.2% / 16.8%. Differences between the columns are within run-to-run noise (standard error ≈ 1–1.5 points for GPQA, AIME and MMMU-Pro).

Teacher-forced log-likelihood of 24,905 sampled reasoning tokens (48 MMLU-Pro prompts, T=1.0, generated by the BF16-activation reference): mean NLL 0.3245 (reference) vs 0.3344 (NVFP4), top-1 agreement with the reference tokens 88.0% vs 87.7%.

Reproduce (server from Deploy with SGLang on port 30000):

pip install sgl-eval==0.1.2
COMMON="--base-url http://localhost:30000/v1 --temperature 1.0 --top-p 0.95 --chat-template-kwarg enable_thinking=true --num-threads 256"
sgl-eval run gsm8k    $COMMON --max-tokens 16384
sgl-eval run gpqa     $COMMON --max-tokens 32768 --n-repeats 4
sgl-eval run aime25   $COMMON --max-tokens 32768 --n-repeats 8
sgl-eval run mmmu_pro $COMMON --max-tokens 16384
Reproduce the checkpoint

The conversion needs no GPU and no calibration data. tools/build_flash_nvfp4.py reads the released checkpoint, transcodes the MXFP4 experts, copies everything else unchanged, and writes the ModelOpt config:

hf download XiaomiMiMo/MiMo-V2.6-Flash-MOPD --local-dir ./src
python tools/build_flash_nvfp4.py --src ./src --unit-input-scale --out ./MiMo-V2.6-Flash-MOPD-NVFP4 --workers 24
# -> expert scale blocks: 9,462,349,824, all-zero: 0, out-of-E4M3-range (re-rounded): 0

The activation-calibration study above used local-inference-lab/quant-toolkit @ 8bdb101 with its MiMo-V2 streaming calibrator; tools/quant-toolkit-8bdb101-mimo-v26-flash.patch adds MXFP4 expert dequantization and ports it to ModelOpt 0.46 / transformers 5.12. It is not needed to rebuild this checkpoint.

Base model

MiMo-V2.6-Flash-MOPD is the MOPD2 upgrade of MiMo-V2.6-Flash-RL : multi-teacher on-policy distillation that fuses several domain-specialized teachers into the student and mitigates tool-call repetition. Same architecture as Flash-RL: a 309B-total / 15B-active sparse MoE with native text, image, video and audio input and a 1M-token context. See the original model card and the technical report for architecture, training and evaluation details. This repository changes only the storage format of the routed-expert weights and adds NVFP4 activation quantization for them; the tokenizer, chat template, processor configs, remote code, audio tokenizer and DFlash draft model are copied unchanged.

Limitations

The base model may generate inaccurate, biased or offensive content, and the quantized model inherits these limitations. W4A4 activation quantization changes numerics relative to the released checkpoint even though the expert weights are identical. Please refer to the original model card for full details.

Citation
@misc{mimo2026v26,
  title={MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement},
  author={{Xiaomi MiMo Team}},
  year={2026},
  howpublished={\url{https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL}},
}

Runs of AxionML MiMo-V2.6-Flash-MOPD-NVFP4 on huggingface.co

198
Total runs
43
24-hour runs
198
3-day runs
198
7-day runs
198
30-day runs

More Information About MiMo-V2.6-Flash-MOPD-NVFP4 huggingface.co Model

More MiMo-V2.6-Flash-MOPD-NVFP4 license Visit here:

https://choosealicense.com/licenses/mit

MiMo-V2.6-Flash-MOPD-NVFP4 huggingface.co

MiMo-V2.6-Flash-MOPD-NVFP4 huggingface.co is an AI model on huggingface.co that provides MiMo-V2.6-Flash-MOPD-NVFP4's model effect (), which can be used instantly with this AxionML MiMo-V2.6-Flash-MOPD-NVFP4 model. huggingface.co supports a free trial of the MiMo-V2.6-Flash-MOPD-NVFP4 model, and also provides paid use of the MiMo-V2.6-Flash-MOPD-NVFP4. Support call MiMo-V2.6-Flash-MOPD-NVFP4 model through api, including Node.js, Python, http.

MiMo-V2.6-Flash-MOPD-NVFP4 huggingface.co Url

https://huggingface.co/AxionML/MiMo-V2.6-Flash-MOPD-NVFP4

AxionML MiMo-V2.6-Flash-MOPD-NVFP4 online free

MiMo-V2.6-Flash-MOPD-NVFP4 huggingface.co is an online trial and call api platform, which integrates MiMo-V2.6-Flash-MOPD-NVFP4's modeling effects, including api services, and provides a free online trial of MiMo-V2.6-Flash-MOPD-NVFP4, you can try MiMo-V2.6-Flash-MOPD-NVFP4 online for free by clicking the link below.

AxionML MiMo-V2.6-Flash-MOPD-NVFP4 online free url in huggingface.co:

https://huggingface.co/AxionML/MiMo-V2.6-Flash-MOPD-NVFP4

MiMo-V2.6-Flash-MOPD-NVFP4 install

MiMo-V2.6-Flash-MOPD-NVFP4 is an open source model from GitHub that offers a free installation service, and any user can find MiMo-V2.6-Flash-MOPD-NVFP4 on GitHub to install. At the same time, huggingface.co provides the effect of MiMo-V2.6-Flash-MOPD-NVFP4 install, users can directly use MiMo-V2.6-Flash-MOPD-NVFP4 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

MiMo-V2.6-Flash-MOPD-NVFP4 install url in huggingface.co:

https://huggingface.co/AxionML/MiMo-V2.6-Flash-MOPD-NVFP4

Url of MiMo-V2.6-Flash-MOPD-NVFP4

MiMo-V2.6-Flash-MOPD-NVFP4 huggingface.co Url

Provider of MiMo-V2.6-Flash-MOPD-NVFP4 huggingface.co

AxionML
ORGANIZATIONS

Other API from AxionML

huggingface.co

Total runs: 189
Run Growth: 189
Growth Rate: 100.00%
Updated:September 29 2026