aufklarer / Gemma-4-E2B-LiteRT-LM

huggingface.co
Total runs: 101
24-hour runs: 0
7-day runs: 0
30-day runs: 19
Model's Last Updated: April 04 2026
text-generation

Introduction of Gemma-4-E2B-LiteRT-LM

Model Details of Gemma-4-E2B-LiteRT-LM

Gemma 4 E2B — LiteRT-LM

Gemma 4 E2B converted to LiteRT-LM for on-device inference on Android, embedded Linux, and desktop.

Exported from the base PyTorch model via litert_torch.generative.export_hf with dynamic_wi4_afp32 quantization.

Model
Property Value
Parameters 5.1B total, 2.3B effective (PLE)
Quantization dynamic_wi4_afp32 (INT4 weights, FP32 activations)
Format .litertlm (LiteRT-LM)
File size 2.39 GB
Context length 32K tokens
Prefill lengths 128, 512
KV cache length 4096
Modalities Text (+ image/audio with multimodal backends)
Files
File Size Description
model.litertlm 2.39 GB Model weights + embedded tokenizer
config.json 0.4 KB Inference metadata
Performance

Benchmarked on macOS ARM64 (Apple Silicon), CPU backend, LiteRT-LM 0.10.1:

Prompt tokens TTFT (ms) Decode (tok/s) Peak memory
16 465 165.8 1.37 GB
64 482 167.4 1.39 GB
128 3,504 169.2 2.08 GB
256 3,528 166.9 2.09 GB

Model load time: 652ms.

Android reference (Samsung S26 Ultra, from Google):

Backend Decode (tok/s) TTFT
GPU 52.1 0.3s
CPU 46.9 1.8s
Usage
Python
import litert_lm

engine = litert_lm.Engine(
    model_path="model.litertlm",
    backend=litert_lm.Backend.CPU,
)

with engine.create_conversation() as conv:
    response = conv.send_message("Hello, how are you?")
    print(response)
CLI
pip install litert-lm-api
litert_lm_advanced_main --model_path=model.litertlm --backend=cpu --benchmark=true
Source

Converted from google/gemma-4-E2B-it using litert-torch-nightly (0.9.0.dev20260403).

Conversion took ~8 minutes on Apple Silicon (M-series, 64GB RAM).

Runs of aufklarer Gemma-4-E2B-LiteRT-LM on huggingface.co

101
Total runs
0
24-hour runs
9
3-day runs
0
7-day runs
19
30-day runs

More Information About Gemma-4-E2B-LiteRT-LM huggingface.co Model

More Gemma-4-E2B-LiteRT-LM license Visit here:

https://choosealicense.com/licenses/apache-2.0

Gemma-4-E2B-LiteRT-LM huggingface.co

Gemma-4-E2B-LiteRT-LM huggingface.co is an AI model on huggingface.co that provides Gemma-4-E2B-LiteRT-LM's model effect (), which can be used instantly with this aufklarer Gemma-4-E2B-LiteRT-LM model. huggingface.co supports a free trial of the Gemma-4-E2B-LiteRT-LM model, and also provides paid use of the Gemma-4-E2B-LiteRT-LM. Support call Gemma-4-E2B-LiteRT-LM model through api, including Node.js, Python, http.

Gemma-4-E2B-LiteRT-LM huggingface.co Url

https://huggingface.co/aufklarer/Gemma-4-E2B-LiteRT-LM

aufklarer Gemma-4-E2B-LiteRT-LM online free

Gemma-4-E2B-LiteRT-LM huggingface.co is an online trial and call api platform, which integrates Gemma-4-E2B-LiteRT-LM's modeling effects, including api services, and provides a free online trial of Gemma-4-E2B-LiteRT-LM, you can try Gemma-4-E2B-LiteRT-LM online for free by clicking the link below.

aufklarer Gemma-4-E2B-LiteRT-LM online free url in huggingface.co:

https://huggingface.co/aufklarer/Gemma-4-E2B-LiteRT-LM

Gemma-4-E2B-LiteRT-LM install

Gemma-4-E2B-LiteRT-LM is an open source model from GitHub that offers a free installation service, and any user can find Gemma-4-E2B-LiteRT-LM on GitHub to install. At the same time, huggingface.co provides the effect of Gemma-4-E2B-LiteRT-LM install, users can directly use Gemma-4-E2B-LiteRT-LM installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Gemma-4-E2B-LiteRT-LM install url in huggingface.co:

https://huggingface.co/aufklarer/Gemma-4-E2B-LiteRT-LM

Url of Gemma-4-E2B-LiteRT-LM

Gemma-4-E2B-LiteRT-LM huggingface.co Url

Provider of Gemma-4-E2B-LiteRT-LM huggingface.co

aufklarer
ORGANIZATIONS

Other API from aufklarer

huggingface.co

Total runs: 2.5K
Run Growth: 2.3K
Growth Rate: 94.66%
Updated:September 16 2025
huggingface.co

Total runs: 151
Run Growth: 93
Growth Rate: 61.59%
Updated:October 15 2025