A 4-bit (NVFP4) quantized release of
Z.ai
's
GLM-5.3-Flash
— a
320B-parameter
natively multimodal Mixture-of-Experts model with ~18B active per token.
4 × B300 → 1 × B300
598.5 GiB → 191.0 GiB (31.9%)
Full 1,048,576-token context · MTP speculative decoding preserved
Highlights
NVFP4 (4-bit float, W4A4)
—
group_size=16
, packed in the compressed-tensors
nvfp4-pack-quantized
format for direct serving in
vLLM
. Both weights and activations are quantized to
4-bit floating point.
Requires NVIDIA Blackwell.
NVFP4 relies on the FP4 tensor cores introduced in the
Blackwell architecture (e.g. B200 / B300 / GB200), so inference must run on a
Blackwell-class GPU. Earlier architectures (Hopper, Ada, Ampere) do not support NVFP4
execution.
Only the routed experts are quantized.
They hold 94.8% of the parameters, so the memory
saving is captured almost in full while every precision-critical path stays in BF16 — the same
split the reference GLM-5.3 NVFP4 release uses.
No architecture change.
Tensor names, layer count and expert count are identical to the
base checkpoint, so stock vLLM serves it as-is — no patched modeling file.
MTP and multimodal preserved.
The multi-token-prediction block and the vision tower stay
BF16, so speculative decoding and image/video inputs work as in the base model.
512 conversations of exactly 4,096 tokens, rendered through the GLM chat template and drawn from
the workloads this model is built for rather than generic web text: agentic tool use (20.5%), SWE
agent trajectories (14.3%), instruction following (8.2%), terminal agents (7.8%), code (7.0%),
STEM (5.9%), reasoning (4.7%), knowledge MCQ (2.3%), and
29.3% Korean
sources. 71.7% of the
samples carry reasoning traces inside
<think>
blocks.
GLM-5.3-Flash-Nota-NVFP4 huggingface.co is an AI model on huggingface.co that provides GLM-5.3-Flash-Nota-NVFP4's model effect (), which can be used instantly with this nota-ai GLM-5.3-Flash-Nota-NVFP4 model. huggingface.co supports a free trial of the GLM-5.3-Flash-Nota-NVFP4 model, and also provides paid use of the GLM-5.3-Flash-Nota-NVFP4. Support call GLM-5.3-Flash-Nota-NVFP4 model through api, including Node.js, Python, http.
GLM-5.3-Flash-Nota-NVFP4 huggingface.co is an online trial and call api platform, which integrates GLM-5.3-Flash-Nota-NVFP4's modeling effects, including api services, and provides a free online trial of GLM-5.3-Flash-Nota-NVFP4, you can try GLM-5.3-Flash-Nota-NVFP4 online for free by clicking the link below.
nota-ai GLM-5.3-Flash-Nota-NVFP4 online free url in huggingface.co:
GLM-5.3-Flash-Nota-NVFP4 is an open source model from GitHub that offers a free installation service, and any user can find GLM-5.3-Flash-Nota-NVFP4 on GitHub to install. At the same time, huggingface.co provides the effect of GLM-5.3-Flash-Nota-NVFP4 install, users can directly use GLM-5.3-Flash-Nota-NVFP4 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
GLM-5.3-Flash-Nota-NVFP4 install url in huggingface.co: