The Latest AIs, every day
AIs with the most favorites on Toolify
AIs with the highest website traffic (monthly visits)
AI Tools by Apps
Discover the Discord of AI
AI Tools by browser extensions
GPTs from GPT Store
Discover The Best Model For AI
Top AI lists by month and monthly visits.
Top AI lists by category and monthly visits.
Top AI lists by region and monthly visits.
Top AI lists by source and monthly visits.
Top AI lists by revenue and real traffic.

Speculative decoding (DFlash2 block-diffusion drafter) tuned specifically for NVIDIA Tesla P40 (Pascal, compute capability 6.1, no FP16 tensor cores).
Fork of
ggml-org/llama.cpp
(PR #27342 lineage) with DFlash2 support, optimized
for the Pascal backend: int8 DP4A (not FP16), multi-stage quantization of the
draft, and an adaptive margin that shortens the verify batch where the
draft's selector is uncertain.
Mini-article: Title placeholder -> 30.26 tok/s on Pascal P40 with 27B dense
A runnable, measured build of llama.cpp + DFlash2 speculative decoding on a
single Tesla P40 (24 GB, sm_61). Everything in this repo's
DFLASH2_RING_PORT.md
is a real measurement on the P40; no other GPUs / cloud numbers are claimed.
Primary pairing tested:
Qwen3.8-27B-UD-Q4_K_XL.gguf
(qwen35 hybrid, Gated DeltaNet)
Qwen3.8-27B-DFlash2-q4mix-self.gguf
(our q4-mix quant)
See
run_best.sh
. Summary:
./build-p40-ring/bin/llama-server \
-m Qwen3.8-27B-UD-Q4_K_XL.gguf \
-md Qwen3.8-27B-DFlash2-q4mix-self.gguf \
-ngl 999 -ngld 999 -c 8192 -b 512 -ub 512 -np 1 \
--load-mode mlock --cache-ram 32768 --checkpoint-min-step 512 \
-fa 1 -ctk q8_0 -ctv q8_0 -ctkd q8_0 -ctvd q8_0 --kv-unified \
--spec-type draft-dflash \
--spec-draft-n-max 7 --spec-draft-n-min 1 --spec-draft-p-min 0.35 \
--spec-draft-ctx 0 --temp 0 --jinja --reasoning off -bs \
-lv 4 --host 0.0.0.0 --port 8080
Key flags:
--spec-type draft-dflash
- DFlash2 block-diffusion drafter
--spec-draft-n-max 7
- sweep optimum (Test 26: 7=30.26 > 8=30.18 > 5=27.4 > 4=27.8)
--spec-draft-p-min 0.35
- adaptive margin (top1-top2 on raw selector logits),
cuts the chain where the selector is a coin flip. Shortens the verify batch.
--spec-draft-n-min 1
- keep short chains alive
-bs
- backend (GPU) argmax sampling, no CPU logit readback
-fa 1
,
-ctk/ctv q8_0
,
--kv-unified
- flash attention + q8 KV
--reasoning off
- fastest mode; use
on
for real reasoning (see below)
To switch to the reasoning scenario replace
--reasoning off
with
--reasoning on
.
All numbers are our own runs, logged in
DFLASH2_RING_PORT.md
.
| Draft quant | size | tok/s | acceptance | mean len |
|---|---|---|---|---|
| q4-mix (recommended) | 1.2 GB | 30.26 | 0.83 | 5.25 |
| Q8 (lucebox) | 2.0 GB | ~30 | 0.85 | 5.27 |
| Q2-hybrid (fc Q2, head) | 0.96GB | 29.0 | 0.80 | 5.17 |
| Q2-all | 0.87GB | 27.1 | 0.76 | 4.69 |
Draft quant is q4-mix = backbone q4_0 + dflash.* heads q8_0 (protects the selector). q4-mix wins: same acceptance as Q8 with a smaller, faster draft; Q2 is slower because the 2-bit selector heads lose near-tie accuracy.
| n_max | tok/s |
|---|---|
| 4 | 27.8 |
| 5 | 27.4 |
| 7 | 30.26 |
| 8 | 30.18 |
n_max=7 is the optimum; adaptive margin already caps the useful chain at ~5.25.
ggml-cuda
selects the MMQ DP4A kernel
(
__dp4a
, cc >= 610) for quantized weights via
(!fp16_mma_hardware_available(cc) || ...)
, compile arch
61
. Draft and
target both run the int8 matmul path. No FP16 path is used for quant weights.
server/scripts/quantize_dflash_draft.py
llama-quantize
):
q4-mix
: backbone q4_0,
dflash.*
heads + conv q8_0 -> 1.2 GB, acceptance-neutral.
llama-quantize
(python-gguf has no K-quant
write
).
common/speculative.cpp
):
raw selector logits top1-top2 margin vs a threshold from
p_min
; stops the
chain when the pick is a coin flip. This is what shortens the verify batch
and makes n_max=7 optimal.
draft_dflash::draft
, one
llama_decode
) - same approach as upstream PR #27342.
-bs
): no CPU-side logit readback during draft/verify.
-fa
) +
q8_0
KV for both models, unified KV.
Our recommended draft (q4-mix, 1.2 GB):
Qwen3.8-27B-DFlash2-q4mix-self.gguf
Target:
Qwen3.8-27B-UD-Q4_K_XL.gguf
(Qwen3.8-27B-UD, qwen35 hybrid).
cmake -B build-p40-ring -DCMAKE_CUDA_ARCHITECTURES=61 -DLLAMA_CUDA=ON -DGGML_CUDA=ON -DLLAMA_CURL=ON
cmake --build build-p40-ring --config Release -j
License note: this is a private fork for experimental Pascal tuning; it wraps upstream llama.cpp (MIT) and DFlash2 draft weights (Apache-2.0 / z-lab).
llama.cpp-DFlash2-pascal6-optimized huggingface.co is an AI model on huggingface.co that provides llama.cpp-DFlash2-pascal6-optimized's model effect (), which can be used instantly with this maxwelhelp llama.cpp-DFlash2-pascal6-optimized model. huggingface.co supports a free trial of the llama.cpp-DFlash2-pascal6-optimized model, and also provides paid use of the llama.cpp-DFlash2-pascal6-optimized. Support call llama.cpp-DFlash2-pascal6-optimized model through api, including Node.js, Python, http.
llama.cpp-DFlash2-pascal6-optimized huggingface.co is an online trial and call api platform, which integrates llama.cpp-DFlash2-pascal6-optimized's modeling effects, including api services, and provides a free online trial of llama.cpp-DFlash2-pascal6-optimized, you can try llama.cpp-DFlash2-pascal6-optimized online for free by clicking the link below.
llama.cpp-DFlash2-pascal6-optimized is an open source model from GitHub that offers a free installation service, and any user can find llama.cpp-DFlash2-pascal6-optimized on GitHub to install. At the same time, huggingface.co provides the effect of llama.cpp-DFlash2-pascal6-optimized install, users can directly use llama.cpp-DFlash2-pascal6-optimized installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
