Inferact/MiniMax-M3-EAGLE3
is an EAGLE3 draft model for accelerating inference of
MiniMax-M3
. It is served end-to-end with
vLLM
and was trained using
TorchSpec
— a torch-native online speculative-decoding training framework that runs FSDP training and vLLM-based target inference concurrently, learning from
MiniMax-M3-regenerated responses and live vLLM-generated hidden states
to match the base model's exact token distribution.
The draft is a
1-layer
dense Llama (
LlamaForCausalLMEagle3
, ~3.3 B params) operating on MiniMax-M3's
hidden_size=6144
/
vocab_size=200064
; at serve time it shares the target's embedding and LM head (EAGLE3). See
config.json
for the full architecture.
Performance
All numbers are measured end-to-end against
MiniMaxAI/MiniMax-M3-MXFP8
served with vLLM at
tensor-parallel-size=4
,
num_speculative_tokens=3
, and
--enforce-eager
. Greedy draft sampling (
topk=1
).
Data:
~456,881 training conversations (the
mix2
dataset: SWE-bench-Pro, SWE-bench, OpenCodeInstruct, kimi-mtp), with
all responses regenerated by MiniMax-M3
— preserving the target's reasoning traces and MiniMax-M3 chat formatting.
Method:
EAGLE3 TTT,
ttt_length=7
,
max_seq_length=32 768
, AdamW at
lr=1 × 10⁻⁴
(cosine decay to 0, 2 % warmup,
max_grad_norm=1.0
), bf16 + gradient checkpointing, FlexAttention, 1 epoch (~14,277 steps). Trained on
5 × GB300 nodes
(2 nodes FSDP2 draft training, dp=8, global batch 32 + 3 nodes vLLM TP=4 target inference). EAGLE3 aux hidden states from target layers (2, 30, 57) + the final layer. Embedding / LM head / final norm are shared from the target (M3 is a VL model, so these live under the
language_model.*
prefix).
Core training command
—
torchspec.train_entry
spawns the FSDP2 trainer and vLLM inference engines as decoupled Ray actors, streaming hidden states through Mooncake:
Draft architecture, TTT depth, sequence length, cluster layout, and optimizer are all YAML-configurable — retargeting or scaling is a config change. See the
TorchSpec repo
for full customization instructions.
MiniMax-M3-EAGLE3 huggingface.co is an AI model on huggingface.co that provides MiniMax-M3-EAGLE3's model effect (), which can be used instantly with this Inferact MiniMax-M3-EAGLE3 model. huggingface.co supports a free trial of the MiniMax-M3-EAGLE3 model, and also provides paid use of the MiniMax-M3-EAGLE3. Support call MiniMax-M3-EAGLE3 model through api, including Node.js, Python, http.
MiniMax-M3-EAGLE3 huggingface.co is an online trial and call api platform, which integrates MiniMax-M3-EAGLE3's modeling effects, including api services, and provides a free online trial of MiniMax-M3-EAGLE3, you can try MiniMax-M3-EAGLE3 online for free by clicking the link below.
Inferact MiniMax-M3-EAGLE3 online free url in huggingface.co:
MiniMax-M3-EAGLE3 is an open source model from GitHub that offers a free installation service, and any user can find MiniMax-M3-EAGLE3 on GitHub to install. At the same time, huggingface.co provides the effect of MiniMax-M3-EAGLE3 install, users can directly use MiniMax-M3-EAGLE3 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.