REAP removes 20% of MoE experts (102 of 512) while preserving the model's routing behavior and output quality. The active parameter count per token is unchanged since the router still selects 10 experts per token from the remaining pool. This yields a
~14% reduction in total disk/memory footprint
with minimal quality loss.
@misc{lasby2025reap,
title = {REAP the Experts: Why Pruning Prevails for One-Shot MoE Compression},
author = {Mike Lasby and Ivan Lazarevich and Nish Sinnadurai and Sean Lie and Yani Ioannou and Vithursan Thangarasa},
year = {2025}, eprint = {2510.13999}, archivePrefix = {arXiv}
}
Sponsors
Made possible by
NVIDIA · TNG Technology · Lambda · Prime Intellect · Hot Aisle
.
Runs of 0xSero Qwen3-Coder-64B on huggingface.co
87
Total runs
0
24-hour runs
31
3-day runs
39
7-day runs
39
30-day runs
More Information About Qwen3-Coder-64B huggingface.co Model
Qwen3-Coder-64B huggingface.co is an AI model on huggingface.co that provides Qwen3-Coder-64B's model effect (), which can be used instantly with this 0xSero Qwen3-Coder-64B model. huggingface.co supports a free trial of the Qwen3-Coder-64B model, and also provides paid use of the Qwen3-Coder-64B. Support call Qwen3-Coder-64B model through api, including Node.js, Python, http.
Qwen3-Coder-64B huggingface.co is an online trial and call api platform, which integrates Qwen3-Coder-64B's modeling effects, including api services, and provides a free online trial of Qwen3-Coder-64B, you can try Qwen3-Coder-64B online for free by clicking the link below.
0xSero Qwen3-Coder-64B online free url in huggingface.co:
Qwen3-Coder-64B is an open source model from GitHub that offers a free installation service, and any user can find Qwen3-Coder-64B on GitHub to install. At the same time, huggingface.co provides the effect of Qwen3-Coder-64B install, users can directly use Qwen3-Coder-64B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.