AxionML / Gemma-4-31B-NVFP4

huggingface.co
Total runs: 543
24-hour runs: 92
7-day runs: 538
30-day runs: 538
Model's Last Updated: September 29 2026
image-text-to-text

Introduction of Gemma-4-31B-NVFP4

Model Details of Gemma-4-31B-NVFP4

AxionML Gemma-4-31B-NVFP4

Mirrored by AxionML for open-source serving and deployment use cases. Part of AxionML's effort to provide ready-to-serve quantized models for the community.

Quantized by NVIDIA. The weights in this repository are an unmodified copy of nvidia/Gemma-4-31B-IT-NVFP4 (revision 4135a98a9b728a548947683219633b25682223ac ). All credit for the quantization belongs to NVIDIA.

This is an NVFP4-quantized version of google/gemma-4-31B-it (30.7B parameters), quantized with NVIDIA Model Optimizer .

About NVFP4 quantization: NVFP4 on Blackwell couples a compact E2M1 FP4 codebook with blockwise FP8 (E4M3) scaling over 16-element micro-blocks, so that 4-bit stored values remain numerically useful for neural-network computation. The E2M1 codebook provides a small, nonuniform set of representable magnitudes up to ±6 and relies on saturating behavior rather than IEEE NaN/Inf encodings to maximize usable range per bit. Using an FP8 block scale (rather than power-of-two-only E8M0) enables fractional scales and error-minimizing scale selection. On Blackwell Tensor Cores, native FP4 multipliers exploit E2M1 simplicity while higher-precision FP32 accumulation protects dot-product accuracy.

Ready for commercial and non-commercial use under Apache 2.0 (Gemma) .

Model Summary
Architecture Dense transformer, hybrid sliding-window + global attention, p-RoPE
Parameters 30.7B
Input Text, image, video (as frames)
Context Length 256K tokens
Vocabulary Size 262,144
Checkpoint Size ~32.6 GB
Evaluation Results
Benchmark BF16 NVFP4
GPQA Diamond 85.80 85.35
AIME 2025 87.92 87.60
MMLU Pro 85.25 84.94
LiveCodeBench (pass@1) 82.49 82.27
SciCode subtask acc (pass@1) 33.61 33.18
Terminal-Bench Hard (pass@1) 27.08 27.08

Scores reported by NVIDIA for this checkpoint. temperature=1.0 , top_p=0.95 , max_new_tokens=131072 .

Quantization Details
  • Quantization format: NVFP4 (W4A4, group size 16) on the MLP / feed-forward linear layers; attention, embeddings, lm_head and the vision tower kept in BF16
  • KV cache: FP8
  • Calibration dataset: cnn_dailymail
  • Tool: NVIDIA Model Optimizer
Usage
Deploy with SGLang
python3 -m sglang.launch_server \
    --model-path AxionML/Gemma-4-31B-NVFP4 \
    --quantization modelopt_fp4 \
    --tp 1 \
    --reasoning-parser gemma4 \
    --tool-call-parser gemma4
Deploy with vLLM
vllm serve AxionML/Gemma-4-31B-NVFP4 \
    --quantization modelopt \
    --tensor-parallel-size 8

Sister checkpoint: AxionML/Gemma-4-12B-NVFP4 uses the same MLP-only recipe.

Limitations

The base model was trained on data that may contain toxic language and societal biases. The quantized model inherits these limitations. It may generate inaccurate, biased, or offensive content. Please refer to the original model card and the upstream quantized model card for full details.

Credits

Runs of AxionML Gemma-4-31B-NVFP4 on huggingface.co

543
Total runs
92
24-hour runs
538
3-day runs
538
7-day runs
538
30-day runs

More Information About Gemma-4-31B-NVFP4 huggingface.co Model

More Gemma-4-31B-NVFP4 license Visit here:

https://choosealicense.com/licenses/apache-2.0

Gemma-4-31B-NVFP4 huggingface.co

Gemma-4-31B-NVFP4 huggingface.co is an AI model on huggingface.co that provides Gemma-4-31B-NVFP4's model effect (), which can be used instantly with this AxionML Gemma-4-31B-NVFP4 model. huggingface.co supports a free trial of the Gemma-4-31B-NVFP4 model, and also provides paid use of the Gemma-4-31B-NVFP4. Support call Gemma-4-31B-NVFP4 model through api, including Node.js, Python, http.

Gemma-4-31B-NVFP4 huggingface.co Url

https://huggingface.co/AxionML/Gemma-4-31B-NVFP4

AxionML Gemma-4-31B-NVFP4 online free

Gemma-4-31B-NVFP4 huggingface.co is an online trial and call api platform, which integrates Gemma-4-31B-NVFP4's modeling effects, including api services, and provides a free online trial of Gemma-4-31B-NVFP4, you can try Gemma-4-31B-NVFP4 online for free by clicking the link below.

AxionML Gemma-4-31B-NVFP4 online free url in huggingface.co:

https://huggingface.co/AxionML/Gemma-4-31B-NVFP4

Gemma-4-31B-NVFP4 install

Gemma-4-31B-NVFP4 is an open source model from GitHub that offers a free installation service, and any user can find Gemma-4-31B-NVFP4 on GitHub to install. At the same time, huggingface.co provides the effect of Gemma-4-31B-NVFP4 install, users can directly use Gemma-4-31B-NVFP4 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Gemma-4-31B-NVFP4 install url in huggingface.co:

https://huggingface.co/AxionML/Gemma-4-31B-NVFP4

Url of Gemma-4-31B-NVFP4

Gemma-4-31B-NVFP4 huggingface.co Url

Provider of Gemma-4-31B-NVFP4 huggingface.co

AxionML
ORGANIZATIONS

Other API from AxionML

huggingface.co

Total runs: 195
Run Growth: 189
Growth Rate: 100.00%
Updated:September 29 2026