This repository contains an
EAGLE-3 style draft model
specifically trained to accelerate the inference of the
Qwen3-4B-Thinking-2507
large language model.
This is
not a standalone model
. It must be used in conjunction with its corresponding base model (
Qwen3-4B-Thinking-2507
) within a speculative decoding framework to achieve significant speedups in text generation.
Base Model:
Qwen3-4B-Thinking-2507
Model Architecture:
EAGLE-3 (Speculative Decoding Draft Model)
Primary Benefit:
Accelerates text generation throughput by 1.5x to 2.5x without compromising the generation quality of the base model.
What is EAGLE?
EAGLE (Extrapolative A* Generative Language Engine) is an advanced speculative decoding method. It uses a small draft model to generate a sequence of draft tokens in parallel. These tokens are then verified by the larger, more powerful base model in a single forward pass. If the draft is accepted, the generation process advances multiple steps at once, leading to a substantial increase in speed.
Performance
This model was evaluated on a diverse set of benchmarks. The
acc_length
(average number of accepted draft tokens) indicates the efficiency of the acceleration. A higher value is better.
Benchmark
acc_length
(num_draft_tokens=4)
acc_length
(num_draft_tokens=8)
gsm8k
2.07
2.07
humaneval
1.99
1.98
math500
1.98
1.98
ceval
1.82
1.82
cmmlu
1.76
1.76
mtbench
1.71
1.71
Average
~1.89
~1.89
These results demonstrate consistent and effective acceleration across various tasks, including coding, math, and general conversation.
Training Details
Training Framework:
This model was trained using
SpecForge
, an open-source framework for speculative decoding research.
Training Data:
The model was trained on the
EagleChat
dataset. Available on
Hugging Face
and
ModelScope
.
Qwen3-4B-Thinking-2507-Eagle huggingface.co is an AI model on huggingface.co that provides Qwen3-4B-Thinking-2507-Eagle's model effect (), which can be used instantly with this taobao-mnn Qwen3-4B-Thinking-2507-Eagle model. huggingface.co supports a free trial of the Qwen3-4B-Thinking-2507-Eagle model, and also provides paid use of the Qwen3-4B-Thinking-2507-Eagle. Support call Qwen3-4B-Thinking-2507-Eagle model through api, including Node.js, Python, http.
Qwen3-4B-Thinking-2507-Eagle huggingface.co is an online trial and call api platform, which integrates Qwen3-4B-Thinking-2507-Eagle's modeling effects, including api services, and provides a free online trial of Qwen3-4B-Thinking-2507-Eagle, you can try Qwen3-4B-Thinking-2507-Eagle online for free by clicking the link below.
taobao-mnn Qwen3-4B-Thinking-2507-Eagle online free url in huggingface.co:
Qwen3-4B-Thinking-2507-Eagle is an open source model from GitHub that offers a free installation service, and any user can find Qwen3-4B-Thinking-2507-Eagle on GitHub to install. At the same time, huggingface.co provides the effect of Qwen3-4B-Thinking-2507-Eagle install, users can directly use Qwen3-4B-Thinking-2507-Eagle installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Qwen3-4B-Thinking-2507-Eagle install url in huggingface.co: