Mirrored by
AxionML
for open-source serving and deployment use cases. Part of AxionML's effort to provide ready-to-serve quantized models for the community.
Quantized by NVIDIA.
The weights in this repository are an unmodified copy of
nvidia/Gemma-4-31B-IT-NVFP4
(revision
4135a98a9b728a548947683219633b25682223ac
). All credit for the quantization belongs to NVIDIA.
About NVFP4 quantization:
NVFP4 on Blackwell couples a compact E2M1 FP4 codebook with blockwise FP8 (E4M3) scaling over 16-element micro-blocks, so that 4-bit stored values remain numerically useful for neural-network computation. The E2M1 codebook provides a small, nonuniform set of representable magnitudes up to ±6 and relies on saturating behavior rather than IEEE NaN/Inf encodings to maximize usable range per bit. Using an FP8 block scale (rather than power-of-two-only E8M0) enables fractional scales and error-minimizing scale selection. On Blackwell Tensor Cores, native FP4 multipliers exploit E2M1 simplicity while higher-precision FP32 accumulation protects dot-product accuracy.
Dense transformer, hybrid sliding-window + global attention, p-RoPE
Parameters
30.7B
Input
Text, image, video (as frames)
Context Length
256K tokens
Vocabulary Size
262,144
Checkpoint Size
~32.6 GB
Evaluation Results
Benchmark
BF16
NVFP4
GPQA Diamond
85.80
85.35
AIME 2025
87.92
87.60
MMLU Pro
85.25
84.94
LiveCodeBench (pass@1)
82.49
82.27
SciCode subtask acc (pass@1)
33.61
33.18
Terminal-Bench Hard (pass@1)
27.08
27.08
Scores reported by NVIDIA for this checkpoint.
temperature=1.0
,
top_p=0.95
,
max_new_tokens=131072
.
Quantization Details
Quantization format:
NVFP4 (W4A4, group size 16) on the MLP / feed-forward linear layers; attention, embeddings,
lm_head
and the vision tower kept in BF16
The base model was trained on data that may contain toxic language and societal biases. The quantized model inherits these limitations. It may generate inaccurate, biased, or offensive content. Please refer to the
original model card
and the
upstream quantized model card
for full details.
Gemma-4-31B-NVFP4 huggingface.co is an AI model on huggingface.co that provides Gemma-4-31B-NVFP4's model effect (), which can be used instantly with this AxionML Gemma-4-31B-NVFP4 model. huggingface.co supports a free trial of the Gemma-4-31B-NVFP4 model, and also provides paid use of the Gemma-4-31B-NVFP4. Support call Gemma-4-31B-NVFP4 model through api, including Node.js, Python, http.
Gemma-4-31B-NVFP4 huggingface.co is an online trial and call api platform, which integrates Gemma-4-31B-NVFP4's modeling effects, including api services, and provides a free online trial of Gemma-4-31B-NVFP4, you can try Gemma-4-31B-NVFP4 online for free by clicking the link below.
AxionML Gemma-4-31B-NVFP4 online free url in huggingface.co:
Gemma-4-31B-NVFP4 is an open source model from GitHub that offers a free installation service, and any user can find Gemma-4-31B-NVFP4 on GitHub to install. At the same time, huggingface.co provides the effect of Gemma-4-31B-NVFP4 install, users can directly use Gemma-4-31B-NVFP4 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.