0xSero / GLM-5-381B-W3A16

huggingface.co
Total runs: 85
24-hour runs: 0
7-day runs: 0
30-day runs: 0
Model's Last Updated: May 30 2026
text-generation

Introduction of GLM-5-381B-W3A16

Model Details of GLM-5-381B-W3A16

Support this work → · X · GitHub · REAP paper · Cerebras REAP

GLM-5-381B-W3A16

W3A16 quantization of zai-org/GLM-5 .

At a glance
Base model zai-org/GLM-5
Format W3A16
Total params 381B
Active / token
Experts / layer 128
Layers 78
Hidden size 6144
Context 202,752
On-disk size 154 GB
Which variant should I pick?
Variant Format Link
GLM-5-381B BF16 link
GLM-5-381B-GGUF-BF16 GGUF link
GLM-5-381B-GGUF-IQ2_M GGUF link
GLM-5-381B-GGUF-IQ2_XXS GGUF link
GLM-5-381B-GGUF-Q3_K_M GGUF link
GLM-5-381B-W3A16 (this) W3A16 link
glm5-reap-observations BF16 link

This repository contains the W3A16 AutoRound quantization of the 50% REAP-pruned GLM-5 checkpoint.

Checkpoint
  • Base family: GLM-5
  • Architecture: GlmMoeDsaForCausalLM
  • Total parameters: 381,464,351,232
  • Source prune: refusal_contrast_reap , compression ratio 0.50 , seed 42 , router renormalization true
  • Quantization method: AutoRound
  • Quantization scheme: W3A16
  • Group size: 128
  • Calibration dataset: NeelNanda/pile-10k
  • Calibration samples: 128
  • Sequence length: 1024
  • Iterations per block: 50
Output
  • Saved model shards: 29
  • Quantized tensors: 29,571 / 29,659
  • Quantization config file: quantization_config.json
Intentionally Unquantized
  • lm_head
  • model.layers.[0-2].mlp.down_proj
  • model.layers.[0-2].mlp.gate_proj
  • model.layers.[0-2].mlp.up_proj
  • model.layers.[0-77].self_attn.indexer.weights_proj
Provenance
  • Quantized artifact path: /data0/external_research/glm5-autoround/full/glm5-reap-50pct-w3a16-pile10k-20260405T182123Z/output/layerwise_refusal_contrast_reap-renorm_true-seed_42-0.50-w3g128
  • Quantization log: /data0/external_research/glm5-autoround/full/glm5-reap-50pct-w3a16-pile10k-20260405T182123Z/quant.log
Notes
  • The source checkpoint for this quantization is the BF16 50% REAP GLM-5 artifact.
  • AutoRound reported total tuning time 4549.26s .
License & citation

License inherited from the base model.

@misc{lasby2025reap,
  title  = {REAP the Experts: Why Pruning Prevails for One-Shot MoE Compression},
  author = {Mike Lasby and Ivan Lazarevich and Nish Sinnadurai and Sean Lie and Yani Ioannou and Vithursan Thangarasa},
  year   = {2025}, eprint = {2510.13999}, archivePrefix = {arXiv}
}
Sponsors

Made possible by NVIDIA · TNG Technology · Lambda · Prime Intellect · Hot Aisle .

Runs of 0xSero GLM-5-381B-W3A16 on huggingface.co

85
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs

More Information About GLM-5-381B-W3A16 huggingface.co Model

More GLM-5-381B-W3A16 license Visit here:

https://choosealicense.com/licenses/apache-2.0

GLM-5-381B-W3A16 huggingface.co

GLM-5-381B-W3A16 huggingface.co is an AI model on huggingface.co that provides GLM-5-381B-W3A16's model effect (), which can be used instantly with this 0xSero GLM-5-381B-W3A16 model. huggingface.co supports a free trial of the GLM-5-381B-W3A16 model, and also provides paid use of the GLM-5-381B-W3A16. Support call GLM-5-381B-W3A16 model through api, including Node.js, Python, http.

GLM-5-381B-W3A16 huggingface.co Url

https://huggingface.co/0xSero/GLM-5-381B-W3A16

0xSero GLM-5-381B-W3A16 online free

GLM-5-381B-W3A16 huggingface.co is an online trial and call api platform, which integrates GLM-5-381B-W3A16's modeling effects, including api services, and provides a free online trial of GLM-5-381B-W3A16, you can try GLM-5-381B-W3A16 online for free by clicking the link below.

0xSero GLM-5-381B-W3A16 online free url in huggingface.co:

https://huggingface.co/0xSero/GLM-5-381B-W3A16

GLM-5-381B-W3A16 install

GLM-5-381B-W3A16 is an open source model from GitHub that offers a free installation service, and any user can find GLM-5-381B-W3A16 on GitHub to install. At the same time, huggingface.co provides the effect of GLM-5-381B-W3A16 install, users can directly use GLM-5-381B-W3A16 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

GLM-5-381B-W3A16 install url in huggingface.co:

https://huggingface.co/0xSero/GLM-5-381B-W3A16

Url of GLM-5-381B-W3A16

GLM-5-381B-W3A16 huggingface.co Url

Provider of GLM-5-381B-W3A16 huggingface.co

0xSero
ORGANIZATIONS

Other API from 0xSero

huggingface.co

Total runs: 820
Run Growth: 509
Growth Rate: 62.07%
Updated:May 30 2026
huggingface.co

Total runs: 357
Run Growth: -42
Growth Rate: -11.76%
Updated:June 26 2026
huggingface.co

Total runs: 100
Run Growth: 8
Growth Rate: 8.00%
Updated:May 30 2026
huggingface.co

Total runs: 89
Run Growth: 63
Growth Rate: 72.41%
Updated:May 30 2026
huggingface.co

Total runs: 64
Run Growth: 29
Growth Rate: 45.31%
Updated:May 30 2026
huggingface.co

Total runs: 61
Run Growth: 35
Growth Rate: 57.38%
Updated:May 30 2026
huggingface.co

Total runs: 45
Run Growth: 23
Growth Rate: 51.11%
Updated:May 30 2026
huggingface.co

Total runs: 37
Run Growth: 26
Growth Rate: 72.22%
Updated:May 30 2026