Benchmarked on macOS ARM64 (Apple Silicon), CPU backend, LiteRT-LM 0.10.1:
Prompt tokens
TTFT (ms)
Decode (tok/s)
Peak memory
16
465
165.8
1.37 GB
64
482
167.4
1.39 GB
128
3,504
169.2
2.08 GB
256
3,528
166.9
2.09 GB
Model load time: 652ms.
Android reference (Samsung S26 Ultra, from Google):
Backend
Decode (tok/s)
TTFT
GPU
52.1
0.3s
CPU
46.9
1.8s
Usage
Python
import litert_lm
engine = litert_lm.Engine(
model_path="model.litertlm",
backend=litert_lm.Backend.CPU,
)
with engine.create_conversation() as conv:
response = conv.send_message("Hello, how are you?")
print(response)
Gemma-4-E2B-LiteRT-LM huggingface.co is an AI model on huggingface.co that provides Gemma-4-E2B-LiteRT-LM's model effect (), which can be used instantly with this aufklarer Gemma-4-E2B-LiteRT-LM model. huggingface.co supports a free trial of the Gemma-4-E2B-LiteRT-LM model, and also provides paid use of the Gemma-4-E2B-LiteRT-LM. Support call Gemma-4-E2B-LiteRT-LM model through api, including Node.js, Python, http.
Gemma-4-E2B-LiteRT-LM huggingface.co is an online trial and call api platform, which integrates Gemma-4-E2B-LiteRT-LM's modeling effects, including api services, and provides a free online trial of Gemma-4-E2B-LiteRT-LM, you can try Gemma-4-E2B-LiteRT-LM online for free by clicking the link below.
aufklarer Gemma-4-E2B-LiteRT-LM online free url in huggingface.co:
Gemma-4-E2B-LiteRT-LM is an open source model from GitHub that offers a free installation service, and any user can find Gemma-4-E2B-LiteRT-LM on GitHub to install. At the same time, huggingface.co provides the effect of Gemma-4-E2B-LiteRT-LM install, users can directly use Gemma-4-E2B-LiteRT-LM installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Gemma-4-E2B-LiteRT-LM install url in huggingface.co: