This model card provides the Gemma 4 E4B model in a way that is ready for deployment on Android, iOS, Desktop, IoT and Web.
Gemma is a family of lightweight, state-of-the-art open models from Google, built from the same research and technology used to create the Gemini models. This particular Gemma 4 model is small so it is ideal for on-device use cases. By running this model on device, users can have private access to Generative AI technology without even requiring an internet connection.
These models are provided in the
.litertlm
format for use with the LiteRT-LM framework. LiteRT-LM is a specialized orchestration layer built directly on top of LiteRT, Google’s high-performance multi-platform runtime trusted by millions of Android and edge developers. LiteRT provides the foundational hardware acceleration via XNNPack for CPU and ML Drift for GPU. LiteRT-LM adds the specialized GenAI libraries and APIs, such as KV-cache management, prompt templating, and function calling. This integrated stack is the same technology powering the Google AI Edge Gallery showcase app.
The model file size is 3.65 GB, which includes a text decoder with 2.24 GB of weights and 0.67 GB of embedding parameters. LiteRT-LM framework always keeps main weights in memory, while the embedding parameters are memory mapped which enables significant working memory savings on some platforms as seen in the detailed data below. The vision and audio models are loaded as needed to further reduce memory consumption.
Ready to integrate this into your product? Get started
here
.
Gemma 4 E4B Performance on LiteRT-LM
All benchmarks were taken using 1024 prefill tokens and 256 decode tokens with a context length of 2048 tokens via LiteRT-LM. The model can support up to 32k context length. The inference on CPU is accelerated via the LiteRT XNNPACK delegate with 4 threads. Time-to-first-token does not include load time. Benchmarks were run with caches enabled and initialized. During the first run, the latency and memory usage may differ. Model size is the size of the file on disk.
CPU memory was measured using,
rusage::ru_maxrss
on Android, Linux and Raspberry Pi,
task_vm_info::phys_footprint
on iOS and MacBook and
process_memory_counters::PrivateUsage
on Windows.
It uses the Gemma quantization scheme that employs a mixture of 4bit and 8bit weights.
Running Gemma inference on the web is currently supported through
LLM Inference Engine
and uses the
gemma-4-E4B-it-web.task
model file. Try it out
live in your browser
(Chrome with WebGPU recommended). To start developing with it, download
the web model
and run with our
sample web page
, or follow the
guide
to add it to your own app.
Benchmarked in Chrome on a MacBook Pro 2024 (Apple M4 Max) with 1024 prefill tokens and 256 decode tokens, but the model can support context lengths up to 128K.
Device
Backend
Prefill (tokens/sec)
Decode (tokens/sec)
Initialization time (sec)
Model size (MB)
CPU Memory (GB)
GPU Memory (GB)
Web
GPU
1598
44.4
1.5
2964
1.1
3.3
GPU memory measured by "GPU Process" memory for all of Chrome while running. Was 130MB when inactive, before any model loading took place.
CPU memory measured for the entire tab while running. Was 55MB when inactive, before any model loading took place.
Runs of litert-community gemma-4-E4B-it-litert-lm on huggingface.co
335.5K
Total runs
8.7K
24-hour runs
11.9K
3-day runs
7.1K
7-day runs
-34.9K
30-day runs
More Information About gemma-4-E4B-it-litert-lm huggingface.co Model
gemma-4-E4B-it-litert-lm huggingface.co is an AI model on huggingface.co that provides gemma-4-E4B-it-litert-lm's model effect (), which can be used instantly with this litert-community gemma-4-E4B-it-litert-lm model. huggingface.co supports a free trial of the gemma-4-E4B-it-litert-lm model, and also provides paid use of the gemma-4-E4B-it-litert-lm. Support call gemma-4-E4B-it-litert-lm model through api, including Node.js, Python, http.
gemma-4-E4B-it-litert-lm huggingface.co is an online trial and call api platform, which integrates gemma-4-E4B-it-litert-lm's modeling effects, including api services, and provides a free online trial of gemma-4-E4B-it-litert-lm, you can try gemma-4-E4B-it-litert-lm online for free by clicking the link below.
litert-community gemma-4-E4B-it-litert-lm online free url in huggingface.co:
gemma-4-E4B-it-litert-lm is an open source model from GitHub that offers a free installation service, and any user can find gemma-4-E4B-it-litert-lm on GitHub to install. At the same time, huggingface.co provides the effect of gemma-4-E4B-it-litert-lm install, users can directly use gemma-4-E4B-it-litert-lm installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
gemma-4-E4B-it-litert-lm install url in huggingface.co: