This repository contains an MXFP8-quantized DFlash 2 draft model for
local-inference-lab/GLM-5.3-Flash-NVFP4
.
It is not a standalone language model. A compatible speculative-decoding
server loads it beside the target model and verifies every drafted token
against the target.
Scale values: biased E8M0 exponents stored as
uint8
Quantization block: 1×32 values
Scale layout: row-major and unswizzled
Excluded module:
lm_head
Draft KV cache quantization: not encoded in the checkpoint
conversion_manifest.json
records the immutable source revision, source and
output checksums, tensor coverage, aggregate quantization error, and
per-weight validation statistics.
Validation status
Status:
qualified
for checkpoint structure, exact format reproduction,
loading, and smoke inference under the following conditions:
Hardware: four NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPUs
Parallelism: tensor parallel size 4 and decode-context parallel size 1
Target attention, MoE, linear, and tensor-parallel all-reduce: B12X
DFlash attention: FlashAttention 2
DFlash linear: B12X MXFP8
DFlash proposal length: seven tokens
DFlash KV cache:
auto
(BF16)
CUDA graph mode:
FULL
requested; target and DFlash2 decode are captured,
while target GDN prefill remains eager
The runtime detected ModelOpt MXFP8, selected
B12xMxfp8LinearKernel
for
draft GEMMs and the fused DFlash context K/V projection, and loaded 1.20 GB of
draft weights. With seven draft tokens, the qualified runtime measured a
2.1157-second median time to first token for a 32,320-token prompt and
185.5 output tokens per second at concurrency one. Speculative throughput
depends on prompt content and acceptance length.
The checkpoint is unsupported in vLLM builds that do not contain the DFlash 2
and ModelOpt MXFP8 integration used by the qualified runtime.
An empty
DFLASH_MODEL_REVISION
makes the launcher resolve the repository's
main
branch. For reproducible deployments, replace the empty value with an
immutable Hugging Face commit hash. The OpenAI-compatible endpoint is
available at
http://127.0.0.1:8000/v1
.
License and attribution
The source DFlash 2 model is distributed under
CC BY-NC-ND 4.0
.
See the
source model card
for its use restrictions and attribution information.
GLM-5.3-Flash-DFlash2 huggingface.co is an AI model on huggingface.co that provides GLM-5.3-Flash-DFlash2's model effect (), which can be used instantly with this local-inference-lab GLM-5.3-Flash-DFlash2 model. huggingface.co supports a free trial of the GLM-5.3-Flash-DFlash2 model, and also provides paid use of the GLM-5.3-Flash-DFlash2. Support call GLM-5.3-Flash-DFlash2 model through api, including Node.js, Python, http.
GLM-5.3-Flash-DFlash2 huggingface.co is an online trial and call api platform, which integrates GLM-5.3-Flash-DFlash2's modeling effects, including api services, and provides a free online trial of GLM-5.3-Flash-DFlash2, you can try GLM-5.3-Flash-DFlash2 online for free by clicking the link below.
local-inference-lab GLM-5.3-Flash-DFlash2 online free url in huggingface.co:
GLM-5.3-Flash-DFlash2 is an open source model from GitHub that offers a free installation service, and any user can find GLM-5.3-Flash-DFlash2 on GitHub to install. At the same time, huggingface.co provides the effect of GLM-5.3-Flash-DFlash2 install, users can directly use GLM-5.3-Flash-DFlash2 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
GLM-5.3-Flash-DFlash2 install url in huggingface.co: