mgoin / Qwen3-4B-speculator.dflash2

huggingface.co
Total runs: 404
24-hour runs: 6
7-day runs: 112
30-day runs: 404
Model's Last Updated: August 20 2026

Introduction of Qwen3-4B-speculator.dflash2

Model Details of Qwen3-4B-speculator.dflash2

Qwen3-4B DFlash2 speculator — epoch 3

This is the final checkpoint from a three-epoch experimental DFlash2 draft-model run for Qwen/Qwen3-4B . It was trained with the implementation proposed in vllm-project/speculators#1006 at commit 0a1b3e0a15d67d551041933529c2c41032f5b28d .

Earlier checkpoints remain available under the epoch-1 and epoch-2 tags. The epoch-3 model weights have SHA-256 459f75b6da6a70b7d5630408196e2203798af0ca33db23bb1d628d8c9212e805 .

Architecture
  • DFlash2DraftModel
  • 5 draft layers
  • block size 8, producing 7 speculative tokens
  • verifier hidden-state taps 1, 9, 17, 25, and 33
  • full Qwen3-4B vocabulary
  • block-local grouped dynamic convolution: kernel 2, group size 16
  • predecessor-conditioned candidate selector: rank 256, unary top-16
  • sample_from_anchor=false

The selector loss weight is 0.1. The base objective is fused CE 0.1 + TV 0.9 with fixed exponential positional decay.

Epoch-3 validation

These are teacher-forced/offline validation metrics on the held-out 1% split. Values are reproduced at the precision stored in val_metrics.json . They are not substitutes for end-to-end vLLM acceptance or downstream accuracy.

Metric Value
loss_epoch 0.32483251770930505
ce_loss_epoch 0.9619933171058769
tv_loss_epoch 0.21214886565706623
accept_rate_epoch 0.51322829890726
accept_len_epoch 3.8197594204219256
full_acc_epoch 0.5349658797316249
position_1_acc_epoch 0.8465721966555794
position_2_acc_epoch 0.6959040977255117
position_3_acc_epoch 0.5847381120430708
position_4_acc_epoch 0.4986562940718212
position_5_acc_epoch 0.4279298022768999
position_6_acc_epoch 0.36984750871261424
position_7_acc_epoch 0.3203285645934762
unary_loss_epoch 0.2871333164955253
selector_loss_epoch 0.37699215447724754
unary_candidate_recall_at_16_epoch 0.8608250633524366
unary_candidate_target_mass_at_16_epoch 0.8473111800272566
teacher_forced_selector_acc_epoch 0.6711308100432931
self_conditioned_path_position_1_conditional_acc_epoch 0.8432806306101062
self_conditioned_path_position_2_conditional_acc_epoch 0.7595550318374937
self_conditioned_path_position_3_conditional_acc_epoch 0.7653011828624474
self_conditioned_path_position_4_conditional_acc_epoch 0.7772607818467425
self_conditioned_path_position_5_conditional_acc_epoch 0.787237711986203
self_conditioned_path_position_6_conditional_acc_epoch 0.7955226060915931
self_conditioned_path_position_7_conditional_acc_epoch 0.8021858592661758
self_conditioned_path_accepted_length_epoch 4.082682826810043
unary_top_16_oracle_accepted_length_epoch 6.427947327652982

Training completed at global step 78,646. Optimizer and scheduler state remain in the local experiment bundle and are intentionally excluded from this serving repository.

Training data and run

Training uses exactly 300,000 Magpie and 207,864 UltraChat rows from inference-optimization/Qwen3-8B-Regenerated-Collection at revision 65d219d6b40bb27c45afe16665147a1d3fa21069 .

The responses are Qwen3-8B regenerated trajectories, as stated by the reference DSpark model card. Qwen3-4B renders/tokenizes and teacher-forces those trajectories; they are not Qwen3-4B on-policy samples. The prepared artifact contains 507,864 rows, 1,618,498,917 total tokens, and 1,547,655,657 supervised tokens, with a 16,384-token preparation window. The pinned regenerated dataset has no declared license or card metadata.

The run used four DDP trainer ranks, AdamW at 6e-4 , a cosine schedule, a 99/1 train/validation split, and online verifier hidden-state generation. After the second epoch, the verifier topology changed from four data-parallel servers to one because the original servers were arrival-starved; the same four trainer ranks and all data, optimizer, schedule, and model settings were retained. Training and validation completed cleanly, and the checkpoint manifest was fully revalidated.

vLLM evaluation

The following end-to-end results use the epoch-2 checkpoint. Epoch 3 has not yet been evaluated, so these numbers are included as serving-validation evidence rather than attributed to the weights on this revision.

The vLLM V2 runner used PR #52816 at 19c9351904 , Ben Chislett's safety fix at 31840cf3ea , and the local Speculators config adapter at 9c6917525f . Baseline and DFlash2 ran concurrently on separate otherwise-free B300 GPUs; throughput figures are therefore preliminary cross-GPU measurements.

Evaluation Qwen3-4B baseline DFlash2 epoch 2
GSM8K 5-shot accuracy 85.82% 86.05%
GSM8K questions/s 151.05 195.29
SPEED-Bench qualitative output tok/s 2,898.25 6,210.74
SPEED-Bench throughput_2k output tok/s 5,890.10 14,728.61
throughput_2k completed / failed 1,536 / 0 1,536 / 0
throughput_2k draft-token acceptance 40.95%
throughput_2k mean accepted length 3.866

The native vLLM qualitative loader consumes only messages[0].content , so that result is a single-turn projection rather than a complete multi-turn SPEED-Bench evaluation. In throughput_2k, both variants emitted exactly 6,291,456 output tokens with max_model_len=32768 , concurrency 32, a 4,096-token output length, and ignore_eos .

Serving

This is a draft model and cannot generate independently. Use it with the unquantized Qwen/Qwen3-4B verifier and seven speculative tokens.

The checkpoint tensors match the public DFlash2/vLLM contract, but the Speculators config is flat while the current serving PR consumes nested dflash_config . Until that conversion lands upstream, serving requires vllm-project/vllm#52816 , Ben's safety fix, and the corresponding config adapter.

vllm serve Qwen/Qwen3-4B \
  --speculative-config '{"model":"mgoin/Qwen3-4B-speculator.dflash2","num_speculative_tokens":7}'
Experimental scope and attribution

The architecture and checkpoint tensor contract are adapted from Z Lab's MIT-licensed implementation pinned at 07ebd93 . The corresponding source attribution and license are retained in the Speculators implementation.

Public DFlash2 materials describe inference but do not publish the training loss or full recipe. The split unary + selector objective used here is an experimental Speculators-native baseline; this checkpoint does not claim to reproduce Z Lab's unpublished objective, parameterization, or reported quality.

Included files
  • portable config.json and custom config.py
  • model.safetensors
  • exact val_metrics.json
  • SHA256SUMS

AI assistance was used for implementation, validation, experiment orchestration, evaluation, and documentation. The human publisher remains responsible for the checkpoint and its claims.

Runs of mgoin Qwen3-4B-speculator.dflash2 on huggingface.co

404
Total runs
6
24-hour runs
60
3-day runs
112
7-day runs
404
30-day runs

More Information About Qwen3-4B-speculator.dflash2 huggingface.co Model

Qwen3-4B-speculator.dflash2 huggingface.co

Qwen3-4B-speculator.dflash2 huggingface.co is an AI model on huggingface.co that provides Qwen3-4B-speculator.dflash2's model effect (), which can be used instantly with this mgoin Qwen3-4B-speculator.dflash2 model. huggingface.co supports a free trial of the Qwen3-4B-speculator.dflash2 model, and also provides paid use of the Qwen3-4B-speculator.dflash2. Support call Qwen3-4B-speculator.dflash2 model through api, including Node.js, Python, http.

Qwen3-4B-speculator.dflash2 huggingface.co Url

https://huggingface.co/mgoin/Qwen3-4B-speculator.dflash2

mgoin Qwen3-4B-speculator.dflash2 online free

Qwen3-4B-speculator.dflash2 huggingface.co is an online trial and call api platform, which integrates Qwen3-4B-speculator.dflash2's modeling effects, including api services, and provides a free online trial of Qwen3-4B-speculator.dflash2, you can try Qwen3-4B-speculator.dflash2 online for free by clicking the link below.

mgoin Qwen3-4B-speculator.dflash2 online free url in huggingface.co:

https://huggingface.co/mgoin/Qwen3-4B-speculator.dflash2

Qwen3-4B-speculator.dflash2 install

Qwen3-4B-speculator.dflash2 is an open source model from GitHub that offers a free installation service, and any user can find Qwen3-4B-speculator.dflash2 on GitHub to install. At the same time, huggingface.co provides the effect of Qwen3-4B-speculator.dflash2 install, users can directly use Qwen3-4B-speculator.dflash2 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Qwen3-4B-speculator.dflash2 install url in huggingface.co:

https://huggingface.co/mgoin/Qwen3-4B-speculator.dflash2

Url of Qwen3-4B-speculator.dflash2

Qwen3-4B-speculator.dflash2 huggingface.co Url

Provider of Qwen3-4B-speculator.dflash2 huggingface.co

mgoin
ORGANIZATIONS

Other API from mgoin

huggingface.co

Total runs: 351
Run Growth: 67
Growth Rate: 19.09%
Updated:February 07 2025
huggingface.co

Total runs: 11
Run Growth: 3
Growth Rate: 27.27%
Updated:September 30 2023