AxionML / GLM-5.3-Flash-NVFP4

huggingface.co
Total runs: 122
24-hour runs: 4
7-day runs: 122
30-day runs: 122
Model's Last Updated: September 29 2026
image-text-to-text

Introduction of GLM-5.3-Flash-NVFP4

Model Details of GLM-5.3-Flash-NVFP4

AxionML GLM-5.3-Flash-NVFP4

Mirrored by AxionML for open-source serving and deployment use cases. Part of AxionML's effort to provide ready-to-serve quantized models for the community.

Quantized by RadixArk. The weights in this repository are an unmodified copy of RadixArk/GLM-5.3-Flash-NVFP4 (revision f46cf340d35a22d0d83d0c1dac8957cf2b1bcd35 ). All credit for the quantization belongs to RadixArk.

This is an NVFP4-quantized version of zai-org/GLM-5.3-Flash (320B total parameters, 18B activated), quantized with NVIDIA Model Optimizer .

About NVFP4 quantization: NVFP4 on Blackwell couples a compact E2M1 FP4 codebook with blockwise FP8 (E4M3) scaling over 16-element micro-blocks, so that 4-bit stored values remain numerically useful for neural-network computation. The E2M1 codebook provides a small, nonuniform set of representable magnitudes up to ±6 and relies on saturating behavior rather than IEEE NaN/Inf encodings to maximize usable range per bit. Using an FP8 block scale (rather than power-of-two-only E8M0) enables fractional scales and error-minimizing scale selection. On Blackwell Tensor Cores, native FP4 multipliers exploit E2M1 simplicity while higher-precision FP32 accumulation protects dot-product accuracy.

Ready for commercial and non-commercial use under MIT .

Model Summary
Architecture Natively multimodal hybrid-attention MoE: KDA linear attention, DSA sparse attention with indexer, MLA, manifold-constrained hyper-connections
Total Parameters 320B
Activated Parameters 18B
Layers / Experts 45 layers (3 dense + 42 MoE), 288 routed experts + shared expert, native MTP/NextN layer
Input Text, image, video
Context Length 1,048,576 tokens
Checkpoint Size ~203 GB
Evaluation Results
Benchmark Protocol NVFP4
GSM8K Full 1,319 × 4 seeds 97.14
AIME 2026 30 × 16 × 4 seeds 92.45
Terminal-Bench 2.1 89 tasks, terminus-2, pass@1 83.1

Scores reported by RadixArk for this checkpoint (SGLang, 4x GB300, FP8 KV cache, NEXTN speculative decoding). Text-only evaluations.

Quantization Details
  • Quantization format: NVFP4 W4A4 (group size 16, abs-max scaling) on gate_proj / up_proj / down_proj of all routed experts, the shared expert and the dense MLPs in layers 0–2
  • Unchanged: all attention (KDA, DSA indexer, MLA), hyper-connections, norms, routers, vision tower, MTP layer, embeddings and lm_head ; KV cache not quantized in the checkpoint (FP8 KV validated at serving time)
  • Calibration dataset: 1,024 cnn_dailymail samples, length 512
  • Tool: NVIDIA Model Optimizer 0.46.0
Usage
Deploy with SGLang
python3 -m sglang.launch_server \
    --model-path AxionML/GLM-5.3-Flash-NVFP4 \
    --quantization modelopt_fp4 \
    --tp-size 4 \
    --dsa-prefill-backend trtllm \
    --dsa-decode-backend trtllm \
    --kv-cache-dtype fp8_e4m3 \
    --moe-runner-backend flashinfer_cutlass \
    --speculative-algorithm NEXTN \
    --speculative-num-steps 5 \
    --speculative-eagle-topk 1 \
    --speculative-num-draft-tokens 6 \
    --speculative-adaptive \
    --reasoning-parser glm45 \
    --tool-call-parser glm47

Use the lmsysorg/sglang:glm-5.3-flash image. Audit evidence ( tensor-audit-b.json , precision-contract-b.json ) is included in this repository.

Limitations

The base model was trained on data that may contain toxic language and societal biases. The quantized model inherits these limitations. It may generate inaccurate, biased, or offensive content. Please refer to the original model card and the upstream quantized model card for full details.

Credits

Runs of AxionML GLM-5.3-Flash-NVFP4 on huggingface.co

122
Total runs
4
24-hour runs
18
3-day runs
122
7-day runs
122
30-day runs

More Information About GLM-5.3-Flash-NVFP4 huggingface.co Model

More GLM-5.3-Flash-NVFP4 license Visit here:

https://choosealicense.com/licenses/mit

GLM-5.3-Flash-NVFP4 huggingface.co

GLM-5.3-Flash-NVFP4 huggingface.co is an AI model on huggingface.co that provides GLM-5.3-Flash-NVFP4's model effect (), which can be used instantly with this AxionML GLM-5.3-Flash-NVFP4 model. huggingface.co supports a free trial of the GLM-5.3-Flash-NVFP4 model, and also provides paid use of the GLM-5.3-Flash-NVFP4. Support call GLM-5.3-Flash-NVFP4 model through api, including Node.js, Python, http.

GLM-5.3-Flash-NVFP4 huggingface.co Url

https://huggingface.co/AxionML/GLM-5.3-Flash-NVFP4

AxionML GLM-5.3-Flash-NVFP4 online free

GLM-5.3-Flash-NVFP4 huggingface.co is an online trial and call api platform, which integrates GLM-5.3-Flash-NVFP4's modeling effects, including api services, and provides a free online trial of GLM-5.3-Flash-NVFP4, you can try GLM-5.3-Flash-NVFP4 online for free by clicking the link below.

AxionML GLM-5.3-Flash-NVFP4 online free url in huggingface.co:

https://huggingface.co/AxionML/GLM-5.3-Flash-NVFP4

GLM-5.3-Flash-NVFP4 install

GLM-5.3-Flash-NVFP4 is an open source model from GitHub that offers a free installation service, and any user can find GLM-5.3-Flash-NVFP4 on GitHub to install. At the same time, huggingface.co provides the effect of GLM-5.3-Flash-NVFP4 install, users can directly use GLM-5.3-Flash-NVFP4 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

GLM-5.3-Flash-NVFP4 install url in huggingface.co:

https://huggingface.co/AxionML/GLM-5.3-Flash-NVFP4

Url of GLM-5.3-Flash-NVFP4

GLM-5.3-Flash-NVFP4 huggingface.co Url

Provider of GLM-5.3-Flash-NVFP4 huggingface.co

AxionML
ORGANIZATIONS

Other API from AxionML

huggingface.co

Total runs: 201
Run Growth: 201
Growth Rate: 100.00%
Updated:September 29 2026