local-inference-lab / GLM-5.3-Flash-DFlash2

huggingface.co
Total runs: 1.3K
24-hour runs: 0
7-day runs: 275
30-day runs: 1.3K
Model's Last Updated: September 17 2026
text-generation

Introduction of GLM-5.3-Flash-DFlash2

Model Details of GLM-5.3-Flash-DFlash2

GLM-5.3-Flash-DFlash2

This repository contains an MXFP8-quantized DFlash 2 draft model for local-inference-lab/GLM-5.3-Flash-NVFP4 . It is not a standalone language model. A compatible speculative-decoding server loads it beside the target model and verifies every drafted token against the target.

The source checkpoint is incoai/GLM-5.3-Flash-DFlash2 at immutable revision dc77ff1c99eeb2df044ee3d4f0094eb033fee410 .

Format
  • Linear weights: float8_e4m3fn
  • Scale values: biased E8M0 exponents stored as uint8
  • Quantization block: 1×32 values
  • Scale layout: row-major and unswizzled
  • Excluded module: lm_head
  • Draft KV cache quantization: not encoded in the checkpoint

conversion_manifest.json records the immutable source revision, source and output checksums, tensor coverage, aggregate quantization error, and per-weight validation statistics.

Validation status

Status: qualified for checkpoint structure, exact format reproduction, loading, and smoke inference under the following conditions:

  • Target: local-inference-lab/GLM-5.3-Flash-NVFP4 revision 520de24eabf507659eaef7c70f14fd584527facc
  • Runtime: voipmonitor/vllm@sha256:ef53437759e3a41d5ee1c4e9045ffdd7df2972faad50d1dc687e3ab479c5867a
  • Hardware: four NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPUs
  • Parallelism: tensor parallel size 4 and decode-context parallel size 1
  • Target attention, MoE, linear, and tensor-parallel all-reduce: B12X
  • DFlash attention: FlashAttention 2
  • DFlash linear: B12X MXFP8
  • DFlash proposal length: seven tokens
  • DFlash KV cache: auto (BF16)
  • CUDA graph mode: FULL requested; target and DFlash2 decode are captured, while target GDN prefill remains eager

The runtime detected ModelOpt MXFP8, selected B12xMxfp8LinearKernel for draft GEMMs and the fused DFlash context K/V projection, and loaded 1.20 GB of draft weights. With seven draft tokens, the qualified runtime measured a 2.1157-second median time to first token for a 32,320-token prompt and 185.5 output tokens per second at concurrency one. Speculative throughput depends on prompt content and acceptance length.

The checkpoint is unsupported in vLLM builds that do not contain the DFlash 2 and ModelOpt MXFP8 integration used by the qualified runtime.

Serving
docker run --rm \
  --gpus '"device=0,1,2,3"' \
  --network host \
  --ipc host \
  --shm-size 32g \
  -e MODEL=local-inference-lab/GLM-5.3-Flash-NVFP4 \
  -e SERVED_MODEL_NAME=GLM-5.3-Flash-NVFP4 \
  -e PORT=8000 \
  -e TP=4 \
  -e DCP=1 \
  -e MAX_NUM_SEQS=16 \
  -e MAX_MODEL_LEN=262144 \
  -e MAX_NUM_BATCHED_TOKENS=4096 \
  -e SPECULATOR=dflash \
  -e NUM_SPECULATIVE_TOKENS=7 \
  -e DFLASH_MODEL=local-inference-lab/GLM-5.3-Flash-DFlash2 \
  -e DFLASH_MODEL_REVISION= \
  -e DFLASH_KV_CACHE_DTYPE=auto \
  -e DFLASH_ATTENTION_BACKEND=FLASH_ATTN \
  -e ATTENTION_BACKEND=B12X \
  -e MOE_BACKEND=b12x \
  -e LINEAR_BACKEND=b12x \
  -e B12X_PCIE_ALLREDUCE=1 \
  -e CUDAGRAPH_MODE=FULL \
  -e VLLM_B12X_MOE_FP4_FORCE_A16=0 \
  voipmonitor/vllm:jovian-judgement-community-dflash2-20260830-r7

An empty DFLASH_MODEL_REVISION makes the launcher resolve the repository's main branch. For reproducible deployments, replace the empty value with an immutable Hugging Face commit hash. The OpenAI-compatible endpoint is available at http://127.0.0.1:8000/v1 .

License and attribution

The source DFlash 2 model is distributed under CC BY-NC-ND 4.0 . See the source model card for its use restrictions and attribution information.

@misc{inco2026dflash2,
  title  = {{DFlash 2: Keep Drafting Parallel}},
  author = {{Inco AI}},
  year   = {2026},
  month  = {August},
  url    = {https://inco.ai/blog/dflash2/}
}

Runs of local-inference-lab GLM-5.3-Flash-DFlash2 on huggingface.co

1.3K
Total runs
0
24-hour runs
-162
3-day runs
275
7-day runs
1.3K
30-day runs

More Information About GLM-5.3-Flash-DFlash2 huggingface.co Model

More GLM-5.3-Flash-DFlash2 license Visit here:

https://choosealicense.com/licenses/cc-by-nc-nd-4.0

GLM-5.3-Flash-DFlash2 huggingface.co

GLM-5.3-Flash-DFlash2 huggingface.co is an AI model on huggingface.co that provides GLM-5.3-Flash-DFlash2's model effect (), which can be used instantly with this local-inference-lab GLM-5.3-Flash-DFlash2 model. huggingface.co supports a free trial of the GLM-5.3-Flash-DFlash2 model, and also provides paid use of the GLM-5.3-Flash-DFlash2. Support call GLM-5.3-Flash-DFlash2 model through api, including Node.js, Python, http.

local-inference-lab GLM-5.3-Flash-DFlash2 online free

GLM-5.3-Flash-DFlash2 huggingface.co is an online trial and call api platform, which integrates GLM-5.3-Flash-DFlash2's modeling effects, including api services, and provides a free online trial of GLM-5.3-Flash-DFlash2, you can try GLM-5.3-Flash-DFlash2 online for free by clicking the link below.

local-inference-lab GLM-5.3-Flash-DFlash2 online free url in huggingface.co:

https://huggingface.co/local-inference-lab/GLM-5.3-Flash-DFlash2

GLM-5.3-Flash-DFlash2 install

GLM-5.3-Flash-DFlash2 is an open source model from GitHub that offers a free installation service, and any user can find GLM-5.3-Flash-DFlash2 on GitHub to install. At the same time, huggingface.co provides the effect of GLM-5.3-Flash-DFlash2 install, users can directly use GLM-5.3-Flash-DFlash2 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

GLM-5.3-Flash-DFlash2 install url in huggingface.co:

https://huggingface.co/local-inference-lab/GLM-5.3-Flash-DFlash2

Url of GLM-5.3-Flash-DFlash2

Provider of GLM-5.3-Flash-DFlash2 huggingface.co

local-inference-lab
ORGANIZATIONS

Other API from local-inference-lab