GGUF builds of
google/gemma-4-E2B-it-assistant
— the official Gemma 4
Multi-Token Prediction (MTP)
drafter for
google/gemma-4-E2B-it
. Use it as a speculative-decoding
draft model alongside the matching Gemma 4 target to get a meaningful decoding
speedup at zero quality loss.
Approximate size:
78M (assistant) / 5.1B target
.
These GGUFs use the custom
gemma4_assistant
architecture and
will not
load in stock
llama.cpp
. They require the
atomic-llama-cpp-turboquant
fork, which adds:
the
gemma4_assistant
MTP drafter arch (incl. the centroid LM head for E2B/E4B),
Loading these files in upstream
ggml-org/llama.cpp
will fail with an
unknown architecture error.
Files
File
Quant
Size
Notes
gemma-4-E2B-it-assistant.F16.gguf
F16
164.3 MB
reference (smallest quality loss vs source)
gemma-4-E2B-it-assistant.Q8_0.gguf
Q8_0
94.8 MB
near-lossless 8-bit
gemma-4-E2B-it-assistant.Q5_K_M.gguf
Q5_K_M
75.7 MB
balanced k-quant
gemma-4-E2B-it-assistant.Q4_K_M.gguf
Q4_K_M
74.5 MB
recommended default for speculative-decoding draft
gemma-4-E2B-it-assistant.Q4_K_S.gguf
Q4_K_S
74.3 MB
smallest k-quant
For
E2B
/
E4B
, the assistant uses an
ordered-embedding centroid head
(
mtp.centroids.weight
+
mtp.token_ordering.weight
) that compresses the LM head over the 262K-vocab into 2048 centroids; this structure is preserved across every quantization level in this repo.
Quick start
Build the fork:
git clone https://github.com/AtomicBot-ai/atomic-llama-cpp-turboquant
cd atomic-llama-cpp-turboquant
# Pick one of the platform-specific configurations:
cmake -B build -DGGML_METAL=ON # Apple Silicon# cmake -B build -DGGML_CUDA=ON # NVIDIA# cmake -B build # CPU-only
cmake --build build --target llama-server llama-cli llama-quantize -j
Download the assistant drafter (this repo) and the matching Gemma 4 target:
hf download AtomicChat/gemma-4-E2B-it-assistant-GGUF \
--include "*Q4_K_M.gguf" --local-dir ./models
# Any GGUF build of the matching target model works; e.g. unsloth's:
hf download unsloth/gemma-4-E2B-it-GGUF \
--include "*Q4_K_M*.gguf" --local-dir ./models
Run
llama-server
with MTP speculative decoding + TurboQuant KV cache:
A ready-made launcher lives at
scripts/run-gemma4-e2b-mtp-server.sh
in the fork (
MTP_PRESET=throughput|lift|balanced|quality
).
How MTP works here
Gemma 4 ships with a small "assistant" head that predicts several future tokens
from the target model's last hidden state. In
atomic-llama-cpp-turboquant
it
is loaded as a separate GGUF via
--mtp-head
and drives a custom speculative
decoder (block_size 2-3, draft_max 6-8 typical). The verifier runs the target model in
parallel, guaranteeing the same output distribution as plain greedy/sampled
decoding.
TurboQuant KV cache
turbo3
is the KV-cache quantization scheme used in this fork; it significantly
reduces KV memory and bandwidth at long contexts with no measurable quality
regression on Gemma 4. Apply it to both target and drafter via
-ctk turbo3 -ctv turbo3 -ctkd turbo3 -ctvd turbo3
.
gemma-4-E2B-it-assistant-GGUF huggingface.co is an AI model on huggingface.co that provides gemma-4-E2B-it-assistant-GGUF's model effect (), which can be used instantly with this AtomicChat gemma-4-E2B-it-assistant-GGUF model. huggingface.co supports a free trial of the gemma-4-E2B-it-assistant-GGUF model, and also provides paid use of the gemma-4-E2B-it-assistant-GGUF. Support call gemma-4-E2B-it-assistant-GGUF model through api, including Node.js, Python, http.
gemma-4-E2B-it-assistant-GGUF huggingface.co is an online trial and call api platform, which integrates gemma-4-E2B-it-assistant-GGUF's modeling effects, including api services, and provides a free online trial of gemma-4-E2B-it-assistant-GGUF, you can try gemma-4-E2B-it-assistant-GGUF online for free by clicking the link below.
AtomicChat gemma-4-E2B-it-assistant-GGUF online free url in huggingface.co:
gemma-4-E2B-it-assistant-GGUF is an open source model from GitHub that offers a free installation service, and any user can find gemma-4-E2B-it-assistant-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of gemma-4-E2B-it-assistant-GGUF install, users can directly use gemma-4-E2B-it-assistant-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
gemma-4-E2B-it-assistant-GGUF install url in huggingface.co: