Gemma 4 E2B
, self-quantized to GGUF by
Atomic Chat
. Built straight from Google's original weights with a per-tensor importance matrix, so this is not a repack of somebody else's files. Runs fully offline.
Highlights
2.3B effective (5.1B with embeddings) parameters
: the weights this repo quantizes.
Context length
: 128K tokens, as published by Google.
35 layers
: Dense decoder, hybrid sliding-window (512) and global attention.
Modalities
: the base model handles Text, Image, Audio; this repo ships text-only quants, it carries no vision projector.
Full imatrix ladder
: every quant is calibrated with an importance matrix, published here alongside the quants.
Reasoning
: All models in the family are designed as highly capable reasoners, with configurable thinking modes.
Diverse & Efficient Architectures
: Offers Dense and Mixture-of-Experts (MoE) variants of different sizes for scalable deployment.
These GGUFs are
self-quantized from the original weights
, not a repack. The importance matrix keeps low-bit quants closer to the full-precision model.
Always pass
--jinja
so the
Gemma 4 E2B chat template
is applied. Without it the model can emit malformed turns.
Model Overview
Property
Value
Base model
google/gemma-4-E2B-it
Parameters
2.3B effective (5.1B with embeddings)
Layers
35
Sliding window
512 tokens
Context length
128K tokens
Vocabulary
262K
Modalities
Text, Image, Audio in the base model; text only in this repo, it ships no vision projector
Architecture
Dense decoder, hybrid sliding-window (512) and global attention, 8 attention heads over 1 KV head,
Gemma4ForConditionalGeneration
This repo
GGUF quants (imatrix); the importance matrix is published here as
imatrix-coding.gguf
. Quants:
Q2_K
,
IQ3_M
,
Q3_K_M
,
Q3_K_L
,
IQ4_XS
,
Q4_K_S
,
Q4_K_M
,
Q5_K_S
,
Q5_K_M
,
Q6_K
,
Q8_0
Benchmarks
Benchmark
Score
MMLU Pro
60.0%
AIME 2026 no tools
37.5%
LiveCodeBench v6
44.0%
Codeforces ELO
633
GPQA Diamond
43.4%
Tau2 (average over 3)
24.5%
BigBench Extra Hard
21.9%
MMMLU
67.4%
MMMU Pro
44.2%
OmniDocBench 1.5 (average edit distance, lower is better)
0.290
MATH-Vision
52.4%
MedXPertQA MM
23.5%
CoVoST
33.47
FLEURS (lower is better)
0.09
MRCR v2 8 needle 128k (average)
19.1%
Scores are Google's published results for the base
google/gemma-4-E2B-it
, not our own measurements. Quantization preserves the large majority of this;
Q4_K_M
and up stay close to full precision.
gemma-4-E2B-it-GGUF huggingface.co is an AI model on huggingface.co that provides gemma-4-E2B-it-GGUF's model effect (), which can be used instantly with this AtomicChat gemma-4-E2B-it-GGUF model. huggingface.co supports a free trial of the gemma-4-E2B-it-GGUF model, and also provides paid use of the gemma-4-E2B-it-GGUF. Support call gemma-4-E2B-it-GGUF model through api, including Node.js, Python, http.
gemma-4-E2B-it-GGUF huggingface.co is an online trial and call api platform, which integrates gemma-4-E2B-it-GGUF's modeling effects, including api services, and provides a free online trial of gemma-4-E2B-it-GGUF, you can try gemma-4-E2B-it-GGUF online for free by clicking the link below.
AtomicChat gemma-4-E2B-it-GGUF online free url in huggingface.co:
gemma-4-E2B-it-GGUF is an open source model from GitHub that offers a free installation service, and any user can find gemma-4-E2B-it-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of gemma-4-E2B-it-GGUF install, users can directly use gemma-4-E2B-it-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
gemma-4-E2B-it-GGUF install url in huggingface.co: