modal-labs / GLM-5.3-Flash-DFlash

huggingface.co
Total runs: 928
24-hour runs: 9
7-day runs: 105
30-day runs: 497
Model's Last Updated: September 21 2026
text-generation

Introduction of GLM-5.3-Flash-DFlash

Model Details of GLM-5.3-Flash-DFlash

GLM-5.3-Flash-DFlash

Paper | Github | Blog

This repository contains a DFlash draft model for zai-org/GLM-5.3-Flash . It is not a standalone language model. It is intended to be paired with the target model in a speculative decoding server.

DFlash uses a lightweight block diffusion draft model to propose multiple tokens in parallel. The target model verifies those proposals, improving serving throughput while preserving the target model's output distribution.

Quick Start

GLM-5.3-Flash needs the DFlash capture hooks in the glm5_next model, available on SGLang main. An example deployment is:

python -m sglang.launch_server \
  --model-path zai-org/GLM-5.3-Flash \
  --tp-size 4 \
  --trust-remote-code \
  --speculative-algorithm DFLASH \
  --speculative-draft-model-path modal-labs/GLM-5.3-Flash-DFlash \
  --speculative-dflash-block-size 8 \
  --speculative-draft-model-quantization unquant \
  --speculative-draft-attention-backend trtllm_mha \
  --speculative-draft-kv-cache-dtype fp8_e4m3 \
  --host 0.0.0.0 \
  --port 30000

Keep the draft model unquantized. Quantizing it lowers the accept length.

License

Distributed under the MIT License , inherited from the target model.

Citation

If you find DFlash useful, please cite the original paper:

@article{chen2026dflash,
  title   = {{DFlash: Block Diffusion for Flash Speculative Decoding}},
  author  = {Chen, Jian and Liang, Yesheng and Liu, Zhijian},
  journal = {arXiv preprint arXiv:2602.06036},
  year    = {2026}
}

Runs of modal-labs GLM-5.3-Flash-DFlash on huggingface.co

928
Total runs
9
24-hour runs
32
3-day runs
105
7-day runs
497
30-day runs

More Information About GLM-5.3-Flash-DFlash huggingface.co Model

More GLM-5.3-Flash-DFlash license Visit here:

https://choosealicense.com/licenses/mit

GLM-5.3-Flash-DFlash huggingface.co

GLM-5.3-Flash-DFlash huggingface.co is an AI model on huggingface.co that provides GLM-5.3-Flash-DFlash's model effect (), which can be used instantly with this modal-labs GLM-5.3-Flash-DFlash model. huggingface.co supports a free trial of the GLM-5.3-Flash-DFlash model, and also provides paid use of the GLM-5.3-Flash-DFlash. Support call GLM-5.3-Flash-DFlash model through api, including Node.js, Python, http.

GLM-5.3-Flash-DFlash huggingface.co Url

https://huggingface.co/modal-labs/GLM-5.3-Flash-DFlash

modal-labs GLM-5.3-Flash-DFlash online free

GLM-5.3-Flash-DFlash huggingface.co is an online trial and call api platform, which integrates GLM-5.3-Flash-DFlash's modeling effects, including api services, and provides a free online trial of GLM-5.3-Flash-DFlash, you can try GLM-5.3-Flash-DFlash online for free by clicking the link below.

modal-labs GLM-5.3-Flash-DFlash online free url in huggingface.co:

https://huggingface.co/modal-labs/GLM-5.3-Flash-DFlash

GLM-5.3-Flash-DFlash install

GLM-5.3-Flash-DFlash is an open source model from GitHub that offers a free installation service, and any user can find GLM-5.3-Flash-DFlash on GitHub to install. At the same time, huggingface.co provides the effect of GLM-5.3-Flash-DFlash install, users can directly use GLM-5.3-Flash-DFlash installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

GLM-5.3-Flash-DFlash install url in huggingface.co:

https://huggingface.co/modal-labs/GLM-5.3-Flash-DFlash

Url of GLM-5.3-Flash-DFlash

GLM-5.3-Flash-DFlash huggingface.co Url

Provider of GLM-5.3-Flash-DFlash huggingface.co

modal-labs
ORGANIZATIONS

Other API from modal-labs