Model Details of GLM-5.2-speculator.dspark-block16
GLM-5.2 DSpark (block16) Speculator Model Card
Overview
This is a DSpark speculator model for the
zai-org/GLM-5.2-FP8
base model, enabling
faster inference through speculative decoding. The architecture extends DFlash with
dual lightweight components: a Markov logit-bias head addressing token dependencies
and a per-position confidence head for acceptance-rate forecasting. Development
utilized the speculators library.
This variant changes one knob versus
RedHatAI/GLM-5.2-speculator.dspark
:
block_size = 16
(versus
8
).
Everything else matches: full attention, full vocabulary, aux layers
[8, 23, 39, 55, 70]
,
lr 6e-4
cosine, 3 epochs,
--max-anchors 1024
, loss
{"ce": 0.1, "tv": 0.9}
, Markov + confidence heads.
Checkpoint series
This repository distributes per-epoch checkpoints from a single 3-epoch training run.
The
main
branch tracks the most recent epoch; each epoch remains permanently
accessible as a separate revision.
revision
epoch
status
epoch-1
1 / 3
✅ this checkpoint
epoch-2
2 / 3
training
epoch-3
3 / 3
training
Training Details
Training employed the Speculators library on prompts from Magpie-Align and
HuggingFaceH4 collections, with GLM-5.2-FP8 generating responses (available as
mgoin/GLM-5.2-FP8-magpie-ultrachat
).
The approach is "online": the draft receives hidden states from a live GLM-5.2-FP8
vLLM server, while the trainer runs FSDP data-parallel on separate hardware. Three
commands reproduce this process — install speculators and vLLM from main branches
first.
The
--assistant-pattern
flag addresses GLM-5.2's inline-reasoning format where
reasoning traces remain within assistant turns; future versions may auto-detect this.
The published checkpoint was trained at scale on
60× GB300
(9 vLLM producer nodes
serving the FP8 verifier + a 6-node, DP=24 FSDP trainer) with hidden states streamed
over a
Mooncake RDMA store
; the command above is the equivalent single-node recipe.
The only difference from the reference
dspark
is
--block-size 16
(versus
8
).
Excluding
--draft-vocab-size
trains on full vocabulary; pass
--draft-vocab-size 32000
to reduce it.
DSpark-specific parameters:
--markov-rank
,
--enable-confidence-head
,
--confidence-head-with-markov
,
--confidence-head-alpha
. Removing these (with
--speculator-type dflash
) produces standard DFlash.
Sub-epoch checkpoints enable resumability with
--checkpoint-freq 0.2
.
GLM-5.2-speculator.dspark-block16 huggingface.co is an AI model on huggingface.co that provides GLM-5.2-speculator.dspark-block16's model effect (), which can be used instantly with this mgoin GLM-5.2-speculator.dspark-block16 model. huggingface.co supports a free trial of the GLM-5.2-speculator.dspark-block16 model, and also provides paid use of the GLM-5.2-speculator.dspark-block16. Support call GLM-5.2-speculator.dspark-block16 model through api, including Node.js, Python, http.
GLM-5.2-speculator.dspark-block16 huggingface.co is an online trial and call api platform, which integrates GLM-5.2-speculator.dspark-block16's modeling effects, including api services, and provides a free online trial of GLM-5.2-speculator.dspark-block16, you can try GLM-5.2-speculator.dspark-block16 online for free by clicking the link below.
mgoin GLM-5.2-speculator.dspark-block16 online free url in huggingface.co:
GLM-5.2-speculator.dspark-block16 is an open source model from GitHub that offers a free installation service, and any user can find GLM-5.2-speculator.dspark-block16 on GitHub to install. At the same time, huggingface.co provides the effect of GLM-5.2-speculator.dspark-block16 install, users can directly use GLM-5.2-speculator.dspark-block16 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
GLM-5.2-speculator.dspark-block16 install url in huggingface.co: