z-lab / Qwen3.5-4B-DFlash

huggingface.co
Total runs: 6.1K
24-hour runs: -3
7-day runs: 330
30-day runs: -874
Model's Last Updated: June 20 2026
text-generation

Introduction of Qwen3.5-4B-DFlash

Model Details of Qwen3.5-4B-DFlash

Qwen3.5-4B-DFlash

Paper | GitHub | Blog

This model is still under training.

DFlash is a novel speculative decoding method that utilizes a lightweight block diffusion model for drafting. It enables efficient, high-quality parallel drafting that pushes the limits of inference speed.

This model is the drafter component. It must be used in conjunction with the target model Qwen/Qwen3.5-4B . It was trained with a context length of 4096 tokens.

DFlash Architecture
🚀 Quick Start
SGLang
Installation
uv pip install "git+https://github.com/sgl-project/sglang.git@refs/pull/16818/head#subdirectory=python"
Inference
python -m sglang.launch_server \
    --model-path Qwen/Qwen3.5-4B \
    --speculative-algorithm DFLASH \
    --speculative-draft-model-path z-lab/Qwen3.5-4B-DFlash \
    --speculative-num-draft-tokens 16 \
    --tp-size 1 \
    --dtype bfloat16 \
    --attention-backend fa3 \
    --mem-fraction-static 0.75 \
    --trust-remote-code \
    --mamba-scheduler-strategy extra_buffer \
    --reasoning-parser qwen3 \
    --tool-call-parser qwen3_coder
Early Results
  • Thinking: enabled
  • Max new tokens: 4096
  • Block size: 16
    Dataset Accept Length
    GSM8K 6.174
    Math500 6.841
    HumanEval 6.909
    MBPP 6.018
    MT-Bench 5.326
    Alpaca 5.012

Runs of z-lab Qwen3.5-4B-DFlash on huggingface.co

6.1K
Total runs
-3
24-hour runs
19
3-day runs
330
7-day runs
-874
30-day runs

More Information About Qwen3.5-4B-DFlash huggingface.co Model

More Qwen3.5-4B-DFlash license Visit here:

https://choosealicense.com/licenses/apache-2.0

Qwen3.5-4B-DFlash huggingface.co

Qwen3.5-4B-DFlash huggingface.co is an AI model on huggingface.co that provides Qwen3.5-4B-DFlash's model effect (), which can be used instantly with this z-lab Qwen3.5-4B-DFlash model. huggingface.co supports a free trial of the Qwen3.5-4B-DFlash model, and also provides paid use of the Qwen3.5-4B-DFlash. Support call Qwen3.5-4B-DFlash model through api, including Node.js, Python, http.

Qwen3.5-4B-DFlash huggingface.co Url

https://huggingface.co/z-lab/Qwen3.5-4B-DFlash

z-lab Qwen3.5-4B-DFlash online free

Qwen3.5-4B-DFlash huggingface.co is an online trial and call api platform, which integrates Qwen3.5-4B-DFlash's modeling effects, including api services, and provides a free online trial of Qwen3.5-4B-DFlash, you can try Qwen3.5-4B-DFlash online for free by clicking the link below.

z-lab Qwen3.5-4B-DFlash online free url in huggingface.co:

https://huggingface.co/z-lab/Qwen3.5-4B-DFlash

Qwen3.5-4B-DFlash install

Qwen3.5-4B-DFlash is an open source model from GitHub that offers a free installation service, and any user can find Qwen3.5-4B-DFlash on GitHub to install. At the same time, huggingface.co provides the effect of Qwen3.5-4B-DFlash install, users can directly use Qwen3.5-4B-DFlash installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Qwen3.5-4B-DFlash install url in huggingface.co:

https://huggingface.co/z-lab/Qwen3.5-4B-DFlash

Url of Qwen3.5-4B-DFlash

Qwen3.5-4B-DFlash huggingface.co Url

Provider of Qwen3.5-4B-DFlash huggingface.co

z-lab
ORGANIZATIONS

Other API from z-lab

huggingface.co

Total runs: 1.9K
Run Growth: -15.0K
Growth Rate: -800.00%
Updated:May 08 2026
huggingface.co

Total runs: 339
Run Growth: -20
Growth Rate: -5.90%
Updated:May 08 2026
huggingface.co

Total runs: 311
Run Growth: -42
Growth Rate: -13.50%
Updated:May 08 2026
huggingface.co

Total runs: 242
Run Growth: -111
Growth Rate: -45.87%
Updated:May 08 2026
huggingface.co

Total runs: 9
Run Growth: 0
Growth Rate: 0.00%
Updated:June 23 2025