Mirrored by
AxionML
for open-source serving and deployment use cases. Part of AxionML's effort to provide ready-to-serve quantized models for the community.
Quantized by RadixArk.
The weights in this repository are an unmodified copy of
RadixArk/GLM-5.3-Flash-NVFP4
(revision
f46cf340d35a22d0d83d0c1dac8957cf2b1bcd35
). All credit for the quantization belongs to RadixArk.
About NVFP4 quantization:
NVFP4 on Blackwell couples a compact E2M1 FP4 codebook with blockwise FP8 (E4M3) scaling over 16-element micro-blocks, so that 4-bit stored values remain numerically useful for neural-network computation. The E2M1 codebook provides a small, nonuniform set of representable magnitudes up to ±6 and relies on saturating behavior rather than IEEE NaN/Inf encodings to maximize usable range per bit. Using an FP8 block scale (rather than power-of-two-only E8M0) enables fractional scales and error-minimizing scale selection. On Blackwell Tensor Cores, native FP4 multipliers exploit E2M1 simplicity while higher-precision FP32 accumulation protects dot-product accuracy.
Ready for commercial and non-commercial use under
MIT
.
Model Summary
Architecture
Natively multimodal hybrid-attention MoE: KDA linear attention, DSA sparse attention with indexer, MLA, manifold-constrained hyper-connections
Scores reported by RadixArk for this checkpoint (SGLang, 4x GB300, FP8 KV cache, NEXTN speculative decoding). Text-only evaluations.
Quantization Details
Quantization format:
NVFP4 W4A4 (group size 16, abs-max scaling) on
gate_proj
/
up_proj
/
down_proj
of all routed experts, the shared expert and the dense MLPs in layers 0–2
Unchanged:
all attention (KDA, DSA indexer, MLA), hyper-connections, norms, routers, vision tower, MTP layer, embeddings and
lm_head
; KV cache not quantized in the checkpoint (FP8 KV validated at serving time)
Use the
lmsysorg/sglang:glm-5.3-flash
image. Audit evidence (
tensor-audit-b.json
,
precision-contract-b.json
) is included in this repository.
Limitations
The base model was trained on data that may contain toxic language and societal biases. The quantized model inherits these limitations. It may generate inaccurate, biased, or offensive content. Please refer to the
original model card
and the
upstream quantized model card
for full details.
GLM-5.3-Flash-NVFP4 huggingface.co is an AI model on huggingface.co that provides GLM-5.3-Flash-NVFP4's model effect (), which can be used instantly with this AxionML GLM-5.3-Flash-NVFP4 model. huggingface.co supports a free trial of the GLM-5.3-Flash-NVFP4 model, and also provides paid use of the GLM-5.3-Flash-NVFP4. Support call GLM-5.3-Flash-NVFP4 model through api, including Node.js, Python, http.
GLM-5.3-Flash-NVFP4 huggingface.co is an online trial and call api platform, which integrates GLM-5.3-Flash-NVFP4's modeling effects, including api services, and provides a free online trial of GLM-5.3-Flash-NVFP4, you can try GLM-5.3-Flash-NVFP4 online for free by clicking the link below.
AxionML GLM-5.3-Flash-NVFP4 online free url in huggingface.co:
GLM-5.3-Flash-NVFP4 is an open source model from GitHub that offers a free installation service, and any user can find GLM-5.3-Flash-NVFP4 on GitHub to install. At the same time, huggingface.co provides the effect of GLM-5.3-Flash-NVFP4 install, users can directly use GLM-5.3-Flash-NVFP4 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
GLM-5.3-Flash-NVFP4 install url in huggingface.co: