--kv-cache-dtype bf16
— mandatory; fp8_e4m3 produces garbled output on sm120
--attention-backend flashinfer
— sm120-compatible (trtllm_mha, flashmla are not)
SGLANG_ENABLE_DEEP_GEMM=0
— DeepGEMM needs WGMMA/TCGEN05 absent on sm120
Memory fit:
320 GB weights + KV cache fits on 8× 96GB (≈768 GB total VRAM). Minimum viable: 6× RTX PRO 6000 with
--tp 2 --pp 3
.
Will not run on sm_90 (H100)
: NVFP4 is Blackwell-native. Both vLLM (Marlin FP4 PTX mismatch) and sglang (
NotImplementedError: Current platform does not support w4a4 nvfp4 quantization
) explicitly block sm_90.
General web text (distribution anchor; provides long samples to compensate for short custom samples)
Multi-dataset loading used AutoRound's
:concat=true
option (patched during build; upstreamable) to pack short instruction samples into full-seqlen sequences.
Wall time
Model load + offload: ~55 min
Calibration + quant: 6h 34m
Save: 7 min
Total: ~7.5 hours on 8× H100 80GB (brev compute)
Quality characteristics
Layer-level loss (iter 0 → iter 49) trajectory:
Layer depth
iter 0 loss
iter 49 loss
Behavior
0-2
0
0
Attention-only; MLP skipped
3-9
1e-6 to 1e-5
1e-6 to 1e-5
Iterative tuning minimal effect
10-30
1e-4 to 1e-2
30-50% reduction
Sign-tuning active
31-55
1e-2 to 1e-1
20-30% reduction
Accumulating
56-77
1e-1 to 8e-1
10-20% reduction
Deep-layer drift
Expected quality impact:
benchmarks on sm120 recommended to measure MMLU/GSM8K/IFEval gap vs BF16 source. Loss magnitudes alone suggest non-trivial degradation at deep layers; whether this matters in practice depends on task.
@misc{lasby2025reap,
title = {REAP the Experts: Why Pruning Prevails for One-Shot MoE Compression},
author = {Mike Lasby and Ivan Lazarevich and Nish Sinnadurai and Sean Lie and Yani Ioannou and Vithursan Thangarasa},
year = {2025}, eprint = {2510.13999}, archivePrefix = {arXiv}
}
Sponsors
Made possible by
NVIDIA · TNG Technology · Lambda · Prime Intellect · Hot Aisle
.
Runs of 0xSero GLM-5.1-555B-NVFP4 on huggingface.co
30
Total runs
0
24-hour runs
-2
3-day runs
9
7-day runs
-47
30-day runs
More Information About GLM-5.1-555B-NVFP4 huggingface.co Model
GLM-5.1-555B-NVFP4 huggingface.co is an AI model on huggingface.co that provides GLM-5.1-555B-NVFP4's model effect (), which can be used instantly with this 0xSero GLM-5.1-555B-NVFP4 model. huggingface.co supports a free trial of the GLM-5.1-555B-NVFP4 model, and also provides paid use of the GLM-5.1-555B-NVFP4. Support call GLM-5.1-555B-NVFP4 model through api, including Node.js, Python, http.
GLM-5.1-555B-NVFP4 huggingface.co is an online trial and call api platform, which integrates GLM-5.1-555B-NVFP4's modeling effects, including api services, and provides a free online trial of GLM-5.1-555B-NVFP4, you can try GLM-5.1-555B-NVFP4 online for free by clicking the link below.
0xSero GLM-5.1-555B-NVFP4 online free url in huggingface.co:
GLM-5.1-555B-NVFP4 is an open source model from GitHub that offers a free installation service, and any user can find GLM-5.1-555B-NVFP4 on GitHub to install. At the same time, huggingface.co provides the effect of GLM-5.1-555B-NVFP4 install, users can directly use GLM-5.1-555B-NVFP4 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.