The source checkpoint for this quantization is the BF16 50% REAP GLM-5 artifact.
AutoRound reported total tuning time
4549.26s
.
License & citation
License inherited from the base model.
@misc{lasby2025reap,
title = {REAP the Experts: Why Pruning Prevails for One-Shot MoE Compression},
author = {Mike Lasby and Ivan Lazarevich and Nish Sinnadurai and Sean Lie and Yani Ioannou and Vithursan Thangarasa},
year = {2025}, eprint = {2510.13999}, archivePrefix = {arXiv}
}
Sponsors
Made possible by
NVIDIA · TNG Technology · Lambda · Prime Intellect · Hot Aisle
.
Runs of 0xSero GLM-5-381B-W3A16 on huggingface.co
85
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About GLM-5-381B-W3A16 huggingface.co Model
GLM-5-381B-W3A16 huggingface.co is an AI model on huggingface.co that provides GLM-5-381B-W3A16's model effect (), which can be used instantly with this 0xSero GLM-5-381B-W3A16 model. huggingface.co supports a free trial of the GLM-5-381B-W3A16 model, and also provides paid use of the GLM-5-381B-W3A16. Support call GLM-5-381B-W3A16 model through api, including Node.js, Python, http.
GLM-5-381B-W3A16 huggingface.co is an online trial and call api platform, which integrates GLM-5-381B-W3A16's modeling effects, including api services, and provides a free online trial of GLM-5-381B-W3A16, you can try GLM-5-381B-W3A16 online for free by clicking the link below.
0xSero GLM-5-381B-W3A16 online free url in huggingface.co:
GLM-5-381B-W3A16 is an open source model from GitHub that offers a free installation service, and any user can find GLM-5-381B-W3A16 on GitHub to install. At the same time, huggingface.co provides the effect of GLM-5-381B-W3A16 install, users can directly use GLM-5-381B-W3A16 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.