A DSpark speculator model for the
zai-org/GLM-5.2-FP8
base model, enabling faster
inference through speculative decoding. DSpark extends the DFlash parallel draft
backbone with two lightweight heads: a
Markov logit-bias head
(low-rank
intra-block token dependency) and a
per-position confidence head
(accept-rate
prediction). Trained with the
speculators
library.
Model Specifications
Base Model
: zai-org/GLM-5.2-FP8
Chat Template
: GLM-5.2 (compatible with
/chat/completions
)
Format
: Safetensors
License
: MIT
Draft
: 5 layers,
block_size=8
, full vocabulary (154,880), aux layers
[8, 23, 39, 55, 70]
Validation Hardware
: NVIDIA B300
Checkpoint series
This repo publishes
per-epoch checkpoints
of a single 3-epoch run.
main
tracks
the latest available epoch; each epoch is also a permanent revision.
revision
epoch
status
epoch-1
1 / 3
✅ available
epoch-2
2 / 3
✅ this checkpoint (=
main
)
epoch-3
3 / 3
training
from transformers import AutoModel
model = AutoModel.from_pretrained(
"mgoin/GLM-5.2-speculator.dspark", revision="epoch-2", trust_remote_code=True
)
Training Details
The model was trained using the Speculators library on prompts from
Magpie-Align/Magpie-Llama-3.1-Pro-300K-Filtered
and
HuggingFaceH4/ultrachat_200k
,
with responses regenerated by GLM-5.2-FP8 itself (published as
mgoin/GLM-5.2-FP8-magpie-ultrachat
).
Training is
online
: the draft consumes hidden states streamed on-the-fly from a
live GLM-5.2-FP8 vLLM server, with the trainer running FSDP data-parallel on separate
GPUs. The three commands below (data prep → server → trainer) reproduce the run.
Install
speculators
and vLLM from main.
GPU indices/parallelism are examples — adjust to your hardware.
--assistant-pattern
is currently needed for GLM-5.2's inline-reasoning chat
format (the
<think>...</think>
trace is kept inside the assistant turn); it may be
auto-detected by future speculators versions.
Train-set metrics at the end of epoch 2 (per-epoch validation passes did not
complete due to server restarts; per-dataset acceptance evaluation will accompany
the final checkpoint):
Still improving epoch-over-epoch (epoch 3 in training).
Acceptance length in vLLM (revision
epoch-2
)
Measured end-to-end in vLLM speculative decoding (nightly
0.23.1rc1.dev709+g2b753ad20
), serving
zai-org/GLM-5.2-FP8
on 4xB300 with
num_speculative_tokens=7
,
draft_sample_method="probabilistic"
, greedy
sampling, batch size 1, 64 single-turn chat prompts (32
HumanEval
+ 32
math_reasoning
from
RedHatAI/speculator_benchmarks
),
1024 output tokens each. Acceptance is computed from vLLM's
spec_decode_num_accepted_tokens_per_pos
/
num_drafts
counter deltas.
The earlier
GLM-5.2-speculator.dspark-preview
is included for reference (measured on a 16-prompt subset of the same set).
This table is updated as later epochs are published;
main
currently
points to
epoch-2
.
Checkpoint
Pos 1
Pos 2
Pos 3
Pos 4
Pos 5
Pos 6
Pos 7
Accept Len
Decode tok/s
epoch-2
(=
main
)
74.7%
56.0%
41.5%
31.1%
23.7%
18.0%
13.2%
3.58
225
epoch-1
74.7%
55.4%
40.2%
29.5%
21.5%
15.4%
10.9%
3.48
219
dspark-preview
57.5%
31.7%
15.4%
7.4%
3.3%
1.5%
0.7%
2.18
139
The epoch-2 gain comes from deeper draft positions (position-1 acceptance is
unchanged), consistent with continued training. For reference, the same
server without speculative decoding decodes at 102 tok/s (2.2x speedup for
epoch-2
).
References
DFlash
: Block Diffusion for Flash Speculative Decoding (arXiv:2602.06036) — the
parallel draft backbone DSpark builds on.
DSpark
(DeepSeek) — the Markov + confidence-head additions replicated here.
GLM-5.2-speculator.dspark huggingface.co is an AI model on huggingface.co that provides GLM-5.2-speculator.dspark's model effect (), which can be used instantly with this RedHatAI GLM-5.2-speculator.dspark model. huggingface.co supports a free trial of the GLM-5.2-speculator.dspark model, and also provides paid use of the GLM-5.2-speculator.dspark. Support call GLM-5.2-speculator.dspark model through api, including Node.js, Python, http.
GLM-5.2-speculator.dspark huggingface.co is an online trial and call api platform, which integrates GLM-5.2-speculator.dspark's modeling effects, including api services, and provides a free online trial of GLM-5.2-speculator.dspark, you can try GLM-5.2-speculator.dspark online for free by clicking the link below.
RedHatAI GLM-5.2-speculator.dspark online free url in huggingface.co:
GLM-5.2-speculator.dspark is an open source model from GitHub that offers a free installation service, and any user can find GLM-5.2-speculator.dspark on GitHub to install. At the same time, huggingface.co provides the effect of GLM-5.2-speculator.dspark install, users can directly use GLM-5.2-speculator.dspark installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
GLM-5.2-speculator.dspark install url in huggingface.co: