Mean accepted length and draft accept rate measured end-to-end against
MiniMaxAI/MiniMax-M3-MXFP8
served with vLLM at
tensor-parallel-size=4
,
num_speculative_tokens=3
, greedy sampling (
temperature=0
,
top_p=1.0
),
max-concurrency=16
.
Dataset
n
Mean accepted length
Draft accept rate
Per-position accept rate (pos 1 / 2 / 3)
MT-Bench
64
2.663
55.42%
0.742 / 0.534 / 0.386
SPEED-Bench (qualitative)
64
2.633
54.43%
0.736 / 0.526 / 0.371
Runs of Inferact MiniMax-M3-EAGLE3-GQA-NVFP4 on huggingface.co
146
Total runs
3
24-hour runs
-54
3-day runs
-52
7-day runs
23
30-day runs
More Information About MiniMax-M3-EAGLE3-GQA-NVFP4 huggingface.co Model
More MiniMax-M3-EAGLE3-GQA-NVFP4 license Visit here:
MiniMax-M3-EAGLE3-GQA-NVFP4 huggingface.co is an AI model on huggingface.co that provides MiniMax-M3-EAGLE3-GQA-NVFP4's model effect (), which can be used instantly with this Inferact MiniMax-M3-EAGLE3-GQA-NVFP4 model. huggingface.co supports a free trial of the MiniMax-M3-EAGLE3-GQA-NVFP4 model, and also provides paid use of the MiniMax-M3-EAGLE3-GQA-NVFP4. Support call MiniMax-M3-EAGLE3-GQA-NVFP4 model through api, including Node.js, Python, http.
MiniMax-M3-EAGLE3-GQA-NVFP4 huggingface.co is an online trial and call api platform, which integrates MiniMax-M3-EAGLE3-GQA-NVFP4's modeling effects, including api services, and provides a free online trial of MiniMax-M3-EAGLE3-GQA-NVFP4, you can try MiniMax-M3-EAGLE3-GQA-NVFP4 online for free by clicking the link below.
Inferact MiniMax-M3-EAGLE3-GQA-NVFP4 online free url in huggingface.co:
MiniMax-M3-EAGLE3-GQA-NVFP4 is an open source model from GitHub that offers a free installation service, and any user can find MiniMax-M3-EAGLE3-GQA-NVFP4 on GitHub to install. At the same time, huggingface.co provides the effect of MiniMax-M3-EAGLE3-GQA-NVFP4 install, users can directly use MiniMax-M3-EAGLE3-GQA-NVFP4 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
MiniMax-M3-EAGLE3-GQA-NVFP4 install url in huggingface.co: