This is the final checkpoint from a three-epoch experimental DFlash2 draft-model
run for
Qwen/Qwen3-4B
. It was trained
with the implementation proposed in
vllm-project/speculators#1006
at commit
0a1b3e0a15d67d551041933529c2c41032f5b28d
.
Earlier checkpoints remain available under the
epoch-1
and
epoch-2
tags.
The epoch-3 model weights have SHA-256
459f75b6da6a70b7d5630408196e2203798af0ca33db23bb1d628d8c9212e805
.
Architecture
DFlash2DraftModel
5 draft layers
block size 8, producing 7 speculative tokens
verifier hidden-state taps 1, 9, 17, 25, and 33
full Qwen3-4B vocabulary
block-local grouped dynamic convolution: kernel 2, group size 16
The selector loss weight is 0.1. The base objective is fused CE 0.1 + TV 0.9
with fixed exponential positional decay.
Epoch-3 validation
These are teacher-forced/offline validation metrics on the held-out 1% split.
Values are reproduced at the precision stored in
val_metrics.json
. They are
not substitutes for end-to-end vLLM acceptance or downstream accuracy.
Training completed at global step 78,646. Optimizer and scheduler state remain
in the local experiment bundle and are intentionally excluded from this
serving repository.
The responses are Qwen3-8B regenerated trajectories, as stated by the
reference DSpark model card. Qwen3-4B renders/tokenizes and teacher-forces
those trajectories; they are not Qwen3-4B on-policy samples. The prepared
artifact contains 507,864 rows, 1,618,498,917 total tokens, and
1,547,655,657 supervised tokens, with a 16,384-token preparation window.
The pinned regenerated dataset has no declared license or card metadata.
The run used four DDP trainer ranks, AdamW at
6e-4
, a cosine schedule, a 99/1
train/validation split, and online verifier hidden-state generation. After the
second epoch, the verifier topology changed from four data-parallel servers to
one because the original servers were arrival-starved; the same four trainer
ranks and all data, optimizer, schedule, and model settings were retained.
Training and validation completed cleanly, and the checkpoint manifest was
fully revalidated.
The following end-to-end results use the
epoch-2
checkpoint. Epoch 3 has
not yet been evaluated, so these numbers are included as serving-validation
evidence rather than attributed to the weights on this revision.
The vLLM V2 runner used PR #52816 at
19c9351904
, Ben Chislett's safety fix
at
31840cf3ea
, and the local Speculators config adapter at
9c6917525f
.
Baseline and DFlash2 ran concurrently on separate otherwise-free B300 GPUs;
throughput figures are therefore preliminary cross-GPU measurements.
Evaluation
Qwen3-4B baseline
DFlash2 epoch 2
GSM8K 5-shot accuracy
85.82%
86.05%
GSM8K questions/s
151.05
195.29
SPEED-Bench qualitative output tok/s
2,898.25
6,210.74
SPEED-Bench throughput_2k output tok/s
5,890.10
14,728.61
throughput_2k completed / failed
1,536 / 0
1,536 / 0
throughput_2k draft-token acceptance
—
40.95%
throughput_2k mean accepted length
—
3.866
The native vLLM qualitative loader consumes only
messages[0].content
, so
that result is a single-turn projection rather than a complete multi-turn
SPEED-Bench evaluation. In throughput_2k, both variants emitted exactly
6,291,456 output tokens with
max_model_len=32768
, concurrency 32, a
4,096-token output length, and
ignore_eos
.
Serving
This is a draft model and cannot generate independently. Use it with the
unquantized
Qwen/Qwen3-4B
verifier and seven speculative tokens.
The checkpoint tensors match the public DFlash2/vLLM contract, but the
Speculators config is flat while the current serving PR consumes nested
dflash_config
. Until that conversion lands upstream, serving requires
vllm-project/vllm#52816
,
Ben's safety fix, and the corresponding config adapter.
The architecture and checkpoint tensor contract are adapted from Z Lab's
MIT-licensed implementation pinned at
07ebd93
.
The corresponding source attribution and license are retained in the
Speculators implementation.
Public DFlash2 materials describe inference but do not publish the training
loss or full recipe. The split unary + selector objective used here is an
experimental Speculators-native baseline; this checkpoint does not claim to
reproduce Z Lab's unpublished objective, parameterization, or reported
quality.
Included files
portable
config.json
and custom
config.py
model.safetensors
exact
val_metrics.json
SHA256SUMS
AI assistance was used for implementation, validation, experiment
orchestration, evaluation, and documentation. The human publisher remains
responsible for the checkpoint and its claims.
Runs of mgoin Qwen3-4B-speculator.dflash2 on huggingface.co
404
Total runs
6
24-hour runs
60
3-day runs
112
7-day runs
404
30-day runs
More Information About Qwen3-4B-speculator.dflash2 huggingface.co Model
Qwen3-4B-speculator.dflash2 huggingface.co
Qwen3-4B-speculator.dflash2 huggingface.co is an AI model on huggingface.co that provides Qwen3-4B-speculator.dflash2's model effect (), which can be used instantly with this mgoin Qwen3-4B-speculator.dflash2 model. huggingface.co supports a free trial of the Qwen3-4B-speculator.dflash2 model, and also provides paid use of the Qwen3-4B-speculator.dflash2. Support call Qwen3-4B-speculator.dflash2 model through api, including Node.js, Python, http.
Qwen3-4B-speculator.dflash2 huggingface.co is an online trial and call api platform, which integrates Qwen3-4B-speculator.dflash2's modeling effects, including api services, and provides a free online trial of Qwen3-4B-speculator.dflash2, you can try Qwen3-4B-speculator.dflash2 online for free by clicking the link below.
mgoin Qwen3-4B-speculator.dflash2 online free url in huggingface.co:
Qwen3-4B-speculator.dflash2 is an open source model from GitHub that offers a free installation service, and any user can find Qwen3-4B-speculator.dflash2 on GitHub to install. At the same time, huggingface.co provides the effect of Qwen3-4B-speculator.dflash2 install, users can directly use Qwen3-4B-speculator.dflash2 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Qwen3-4B-speculator.dflash2 install url in huggingface.co: