Q2-REAP-ds4
: compact DS4 profile using
IQ2_XXS
routed gate/up experts,
Q2_K
routed down experts, and
Q8_0
shared/output/attention projections.
These are DS4/DwarfStar-specific GGUF files for DeepSeek-V4 Flash REAP checkpoints. They are not generic llama.cpp files unless your runtime supports the same DeepSeek-V4 Flash tensor layout and DS4 metadata.
Validation
Validation summaries are uploaded in this repo under:
validation/20260528T160633Z/SUMMARY.md
validation/20260528T160633Z/summary.json
The Mini Q2 GGUF completed the DS4 context sweep through
200000
context on one DGX Spark:
Context
Prefill tok/s
Decode tok/s
KV bytes
2,048
348.19
12.75
52,184,460
4,096
358.51
13.50
80,373,132
8,192
352.29
13.32
136,750,476
16,384
348.25
13.24
249,505,164
32,768
322.07
12.40
475,014,540
65,536
287.26
11.49
926,033,292
131,072
241.57
9.81
1,828,070,796
200,000
194.24
9.17
2,776,775,308
API probes completed through at least the
131072
window before
spark-2822
became unreachable during the tail of the
200000
validation step:
Context
Prompt tokens
TTFT seconds
Prefill tok/s
Decode tok/s
Marker visible
65,536
59,867
176.54
339.12
13.01
true
131,072
119,696
390.59
306.45
11.70
true
This repo publishes the validated Q2 long-context profile only.
License & citation
License inherited from the base model.
@misc{lasby2025reap,
title = {REAP the Experts: Why Pruning Prevails for One-Shot MoE Compression},
author = {Mike Lasby and Ivan Lazarevich and Nish Sinnadurai and Sean Lie and Yani Ioannou and Vithursan Thangarasa},
year = {2025}, eprint = {2510.13999}, archivePrefix = {arXiv}
}
Sponsors
Made possible by
NVIDIA · TNG Technology · Lambda · Prime Intellect · Hot Aisle
.
Runs of 0xSero DeepSeek-V4-Flash-162B-GGUF on huggingface.co
1.5K
Total runs
0
24-hour runs
-31
3-day runs
-568
7-day runs
-3.7K
30-day runs
More Information About DeepSeek-V4-Flash-162B-GGUF huggingface.co Model
More DeepSeek-V4-Flash-162B-GGUF license Visit here:
DeepSeek-V4-Flash-162B-GGUF huggingface.co is an AI model on huggingface.co that provides DeepSeek-V4-Flash-162B-GGUF's model effect (), which can be used instantly with this 0xSero DeepSeek-V4-Flash-162B-GGUF model. huggingface.co supports a free trial of the DeepSeek-V4-Flash-162B-GGUF model, and also provides paid use of the DeepSeek-V4-Flash-162B-GGUF. Support call DeepSeek-V4-Flash-162B-GGUF model through api, including Node.js, Python, http.
DeepSeek-V4-Flash-162B-GGUF huggingface.co is an online trial and call api platform, which integrates DeepSeek-V4-Flash-162B-GGUF's modeling effects, including api services, and provides a free online trial of DeepSeek-V4-Flash-162B-GGUF, you can try DeepSeek-V4-Flash-162B-GGUF online for free by clicking the link below.
0xSero DeepSeek-V4-Flash-162B-GGUF online free url in huggingface.co:
DeepSeek-V4-Flash-162B-GGUF is an open source model from GitHub that offers a free installation service, and any user can find DeepSeek-V4-Flash-162B-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of DeepSeek-V4-Flash-162B-GGUF install, users can directly use DeepSeek-V4-Flash-162B-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
DeepSeek-V4-Flash-162B-GGUF install url in huggingface.co: