diffuse-cpp / LLaDA-8B-Instruct-GGUF

huggingface.co
Total runs: 307
24-hour runs: 0
7-day runs: -35
30-day runs: -9
Model's Last Updated: March 31 2026
text-generation

Introduction of LLaDA-8B-Instruct-GGUF

Model Details of LLaDA-8B-Instruct-GGUF

LLaDA-8B-Instruct GGUF

GGUF quantized versions of GSAI-ML/LLaDA-8B-Instruct for use with diffuse-cpp .

LLaDA is a diffusion language model that generates text by iterative unmasking rather than autoregressive token-by-token prediction.

Paper: Diffusion Language Models are Faster than Autoregressive on CPU -- C. Esteban, 2026

Available Quantizations
File Quant Size Description
llada-8b-q4km.gguf Q4_K_M 5.1 GB Recommended best throughput
llada-8b-q8_0.gguf Q8_0 8.4 GB High quality, good throughput
llada-8b-f16.gguf F16 14.9 GB Full precision reference
Benchmark (AMD EPYC 4465P 12-Core, 64 tokens, steps=16, threads=12)
Real Prompt Performance (Q4_K_M + entropy_exit)
Prompt type tok/s Steps used Speedup
Factual ("Capital of France?") 9.22 4 3.9x
Translation ("Translate to French") 10.23 3 4.6x
Arithmetic ("15 x 23?") 11.49 3 5.5x
Code (is_prime function) 2.53 15 1.1x
Creative (poem, explanation) 2.33 17 1.0x

entropy_exit adapts to prompt difficulty: 3-4 steps for easy, 16 for hard. Never slower than baseline.

Quantization Comparison (low_confidence baseline)
Model Size tok/s vs F16
F16 14.9 GB 1.64 1.00x
Q8_0 8.4 GB 1.84 1.12x
Q4_K_M 5.1 GB 2.52 1.54x
Summary
  • ~10 tok/s on easy real prompts (Q4_K_M + entropy_exit)
  • ~6x faster than F16 baseline on factual/translation tasks
  • 7.5x thread scaling from 1 to 12 threads
  • 40+ tok/s peak on synthetic benchmarks (single forward pass)

Full results: research/benchmark/RESULTS.md

Usage
git clone --recursive https://github.com/iafiscal1212/diffuse-cpp
cd diffuse-cpp
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j$(nproc)

# Generate with entropy_exit (recommended)
python tools/generate.py     --model-dir /path/to/LLaDA-8B-Instruct     --gguf llada-8b-q4km.gguf     -p "What is the capital of France?"     -s 16 -t 12 --remasking entropy_exit

Runs of diffuse-cpp LLaDA-8B-Instruct-GGUF on huggingface.co

307
Total runs
0
24-hour runs
-25
3-day runs
-35
7-day runs
-9
30-day runs

More Information About LLaDA-8B-Instruct-GGUF huggingface.co Model

More LLaDA-8B-Instruct-GGUF license Visit here:

https://choosealicense.com/licenses/apache-2.0

LLaDA-8B-Instruct-GGUF huggingface.co

LLaDA-8B-Instruct-GGUF huggingface.co is an AI model on huggingface.co that provides LLaDA-8B-Instruct-GGUF's model effect (), which can be used instantly with this diffuse-cpp LLaDA-8B-Instruct-GGUF model. huggingface.co supports a free trial of the LLaDA-8B-Instruct-GGUF model, and also provides paid use of the LLaDA-8B-Instruct-GGUF. Support call LLaDA-8B-Instruct-GGUF model through api, including Node.js, Python, http.

LLaDA-8B-Instruct-GGUF huggingface.co Url

https://huggingface.co/diffuse-cpp/LLaDA-8B-Instruct-GGUF

diffuse-cpp LLaDA-8B-Instruct-GGUF online free

LLaDA-8B-Instruct-GGUF huggingface.co is an online trial and call api platform, which integrates LLaDA-8B-Instruct-GGUF's modeling effects, including api services, and provides a free online trial of LLaDA-8B-Instruct-GGUF, you can try LLaDA-8B-Instruct-GGUF online for free by clicking the link below.

diffuse-cpp LLaDA-8B-Instruct-GGUF online free url in huggingface.co:

https://huggingface.co/diffuse-cpp/LLaDA-8B-Instruct-GGUF

LLaDA-8B-Instruct-GGUF install

LLaDA-8B-Instruct-GGUF is an open source model from GitHub that offers a free installation service, and any user can find LLaDA-8B-Instruct-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of LLaDA-8B-Instruct-GGUF install, users can directly use LLaDA-8B-Instruct-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

LLaDA-8B-Instruct-GGUF install url in huggingface.co:

https://huggingface.co/diffuse-cpp/LLaDA-8B-Instruct-GGUF

Url of LLaDA-8B-Instruct-GGUF

LLaDA-8B-Instruct-GGUF huggingface.co Url

Provider of LLaDA-8B-Instruct-GGUF huggingface.co

diffuse-cpp
ORGANIZATIONS

Other API from diffuse-cpp