Gemma 4 E4B
, self-quantized to GGUF by
Atomic Chat
. Built straight from Google's original weights with a per-tensor importance matrix. Runs fully offline.
Highlights
Natively multimodal
— handles text, image, and audio input and generates text output.
4.5B effective parameters (8B with embeddings)
— the "E" stands for "effective", using Per-Layer Embeddings (PLE) for on-device efficiency.
128K-token context window
built on a hybrid local/global attention mechanism.
Built-in thinking mode
— configurable step-by-step reasoning, triggered with the
<|think|>
token.
Native function calling
for structured tool use and agentic workflows.
Multilingual
— out-of-the-box support for 35+ languages, pre-trained on 140+ languages.
These GGUFs are
self-quantized from the original weights
, not a repack. The importance matrix keeps low-bit quants closer to the full-precision model.
Always pass
--jinja
so the
Gemma 4 E4B chat template
is applied. Without it the model can emit malformed turns.
Model Overview
Property
Value
Base model
google/gemma-4-E4B-it
Parameters
4.5B effective (8B with embeddings); uses Per-Layer Embeddings (PLE)
Layers
42
Context length
128K tokens
Vocabulary
262K
Modalities
Text, Image, Audio
Architecture
Dense, hybrid local sliding-window (512) + global attention with p-RoPE
This repo
GGUF quants (imatrix) + vision mmproj
Gemma 4 E4B is multimodal. This repo ships the
mmproj-gemma4-e4b-it-f16.gguf
vision projector. With
-hf
it is pulled automatically; otherwise pass
--mmproj
. Use
llama-mtmd-cli
or
llama-server
to feed images.
Scores are Google's published results for the base
google/gemma-4-E4B-it
. Quantization preserves the large majority of this;
Q4_K_M
and up sit within a point or two of full precision.
Choosing a quant
Quant
Size
Notes
Q2_K
4.4 GB
Smallest. Minimal RAM, clear quality drop.
IQ3_M
4.7 GB
Beats Q3 at similar size thanks to imatrix. Best low-RAM pick.
Q3_K_M
4.9 GB
Low quality but usable.
Q3_K_L
5.0 GB
A step above Q3_K_M.
IQ4_XS
5.1 GB
Excellent quality for size. Recommended low-bit.
Q4_K_S
5.2 GB
Compact Q4, fast.
Q4_K_M
5.3 GB
Recommended default. Best balance of size, speed and quality.
UD-Q4_K_XL
6.2 GB
Dynamic. Embeddings and output kept at Q8_0 for higher quality at a Q4 footprint.
Q5_K_S
5.7 GB
Higher quality.
Q5_K_M
5.8 GB
Higher quality, low loss.
Q6_K
6.2 GB
Near lossless.
Q8_0
8.0 GB
Effectively lossless, reference quality.
Pick the largest file that fits your (V)RAM with room for context.
Q4_K_M
or
UD-Q4_K_XL
is the sweet spot for most setups;
Q6_K
or
Q8_0
for maximum fidelity.
Get started
Run Gemma 4 E4B locally with:
Atomic Chat
:
the easiest path. Open the app, search
AtomicChat/gemma4-e4b-it-GGUF
, pick a quant, hit
Use this model
.
gemma4-e4b-it-GGUF huggingface.co is an AI model on huggingface.co that provides gemma4-e4b-it-GGUF's model effect (), which can be used instantly with this AtomicChat gemma4-e4b-it-GGUF model. huggingface.co supports a free trial of the gemma4-e4b-it-GGUF model, and also provides paid use of the gemma4-e4b-it-GGUF. Support call gemma4-e4b-it-GGUF model through api, including Node.js, Python, http.
gemma4-e4b-it-GGUF huggingface.co is an online trial and call api platform, which integrates gemma4-e4b-it-GGUF's modeling effects, including api services, and provides a free online trial of gemma4-e4b-it-GGUF, you can try gemma4-e4b-it-GGUF online for free by clicking the link below.
AtomicChat gemma4-e4b-it-GGUF online free url in huggingface.co:
gemma4-e4b-it-GGUF is an open source model from GitHub that offers a free installation service, and any user can find gemma4-e4b-it-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of gemma4-e4b-it-GGUF install, users can directly use gemma4-e4b-it-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.