Inferact / MiniMax-M3-EAGLE3-GQA

huggingface.co
Total runs: 8.1K
24-hour runs: 0
7-day runs: -4.3K
30-day runs: -12.3K
Model's Last Updated: July 15 2026
text-generation

Introduction of MiniMax-M3-EAGLE3-GQA

Model Details of MiniMax-M3-EAGLE3-GQA

Model Overview

Inferact/MiniMax-M3-EAGLE3-GQA is a grouped-query-attention (GQA) EAGLE3 draft model for accelerating inference of MiniMax-M3 , served with vLLM and trained with TorchSpec .

It is retrained on the same datasets as the multi-head-attention version Inferact/MiniMax-M3-EAGLE3 kimi-mtp, OpenCodeInstruct, SWE-bench, and SWE-bench-Pro — with the draft's attention changed from MHA to GQA ( num_key_value_heads: 64 → 4 ) for inference efficiency (16× smaller draft KV cache) and compatibility with the target model .

The draft is a 1-layer dense Llama ( LlamaForCausalLMEagle3 ) on MiniMax-M3's hidden_size=6144 / vocab_size=200064 ; at serve time it shares the target's embedding and LM head (EAGLE3). See config.json for the full architecture.


Performance

Mean accepted length and draft accept rate measured end-to-end against MiniMaxAI/MiniMax-M3-MXFP8 served with vLLM at tensor-parallel-size=4 , num_speculative_tokens=3 , greedy sampling ( temperature=0 , top_p=1.0 ), max-concurrency=16 .

Dataset n Mean accepted length Draft accept rate Per-position accept rate (pos 1 / 2 / 3)
MT-Bench 64 2.668 55.62% 0.745 / 0.537 / 0.387
SPEED-Bench (qualitative) 64 2.561 52.04% 0.719 / 0.500 / 0.342

Runs of Inferact MiniMax-M3-EAGLE3-GQA on huggingface.co

8.1K
Total runs
0
24-hour runs
-186
3-day runs
-4.3K
7-day runs
-12.3K
30-day runs

More Information About MiniMax-M3-EAGLE3-GQA huggingface.co Model

More MiniMax-M3-EAGLE3-GQA license Visit here:

https://choosealicense.com/licenses/mit

MiniMax-M3-EAGLE3-GQA huggingface.co

MiniMax-M3-EAGLE3-GQA huggingface.co is an AI model on huggingface.co that provides MiniMax-M3-EAGLE3-GQA's model effect (), which can be used instantly with this Inferact MiniMax-M3-EAGLE3-GQA model. huggingface.co supports a free trial of the MiniMax-M3-EAGLE3-GQA model, and also provides paid use of the MiniMax-M3-EAGLE3-GQA. Support call MiniMax-M3-EAGLE3-GQA model through api, including Node.js, Python, http.

MiniMax-M3-EAGLE3-GQA huggingface.co Url

https://huggingface.co/Inferact/MiniMax-M3-EAGLE3-GQA

Inferact MiniMax-M3-EAGLE3-GQA online free

MiniMax-M3-EAGLE3-GQA huggingface.co is an online trial and call api platform, which integrates MiniMax-M3-EAGLE3-GQA's modeling effects, including api services, and provides a free online trial of MiniMax-M3-EAGLE3-GQA, you can try MiniMax-M3-EAGLE3-GQA online for free by clicking the link below.

Inferact MiniMax-M3-EAGLE3-GQA online free url in huggingface.co:

https://huggingface.co/Inferact/MiniMax-M3-EAGLE3-GQA

MiniMax-M3-EAGLE3-GQA install

MiniMax-M3-EAGLE3-GQA is an open source model from GitHub that offers a free installation service, and any user can find MiniMax-M3-EAGLE3-GQA on GitHub to install. At the same time, huggingface.co provides the effect of MiniMax-M3-EAGLE3-GQA install, users can directly use MiniMax-M3-EAGLE3-GQA installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

MiniMax-M3-EAGLE3-GQA install url in huggingface.co:

https://huggingface.co/Inferact/MiniMax-M3-EAGLE3-GQA

Url of MiniMax-M3-EAGLE3-GQA

MiniMax-M3-EAGLE3-GQA huggingface.co Url

Provider of MiniMax-M3-EAGLE3-GQA huggingface.co

Inferact
ORGANIZATIONS

Other API from Inferact