litert-community / SmolLM-135M-Instruct

huggingface.co
Total runs: 1.3K
24-hour runs: 0
7-day runs: -486
30-day runs: -486
Model's Last Updated: September 23 2025
text-generation

Introduction of SmolLM-135M-Instruct

Model Details of SmolLM-135M-Instruct

litert-community/SmolLM-135M-Instruct

This model provides a few variants of HuggingFaceTB/SmolLM-135M-Instruct that are ready for deployment on Android using the LiteRT (fka TFLite) stack and MediaPipe LLM Inference API .

Use the models
Colab

Disclaimer: The target deployment surface for the LiteRT models is Android/iOS/Web and the stack has been optimized for performance on these targets. Trying out the system in Colab is an easier way to familiarize yourself with the LiteRT stack, with the caveat that the performance (memory and latency) on Colab could be much worse than on a local device.

Open In Colab

Android
  • Download and install the apk .
  • Follow the instructions in the app.

To build the demo app from source, please follow the instructions from the GitHub repository.

Performance
Android

Note that all benchmark stats are from a Samsung S24 Ultra with 1280 KV cache size with multiple prefill signatures enabled.

Backend Context length Prefill (tokens/sec) Decode (tokens/sec) Time-to-first-token (sec) Memory (RSS in MB) Model size (MB)
fp32 (baseline) cpu

1280

498.05 tk/s

47.96 tk/s

0.78 s

931 MB

527 MB

dynamic_int8 cpu

1280

1084.75 tk/s

43.50 tk/s

0.46 s

579 MB

159 MB

  • Model Size: measured by the size of the .tflite flatbuffer (serialization format for LiteRT models)
  • Memory: indicator of peak RAM usage
  • The inference on CPU is accelerated via the LiteRT XNNPACK delegate with 4 threads
  • Benchmark is run with cache enabled and initialized. During the first run, the time to first token may differ.
  • dynamic_int4: quantized model with int4 weights and float activations.
  • dynamic_int8: quantized model with int8 weights and float activations.

Runs of litert-community SmolLM-135M-Instruct on huggingface.co

1.3K
Total runs
0
24-hour runs
-197
3-day runs
-486
7-day runs
-486
30-day runs

More Information About SmolLM-135M-Instruct huggingface.co Model

More SmolLM-135M-Instruct license Visit here:

https://choosealicense.com/licenses/apache-2.0

SmolLM-135M-Instruct huggingface.co

SmolLM-135M-Instruct huggingface.co is an AI model on huggingface.co that provides SmolLM-135M-Instruct's model effect (), which can be used instantly with this litert-community SmolLM-135M-Instruct model. huggingface.co supports a free trial of the SmolLM-135M-Instruct model, and also provides paid use of the SmolLM-135M-Instruct. Support call SmolLM-135M-Instruct model through api, including Node.js, Python, http.

litert-community SmolLM-135M-Instruct online free

SmolLM-135M-Instruct huggingface.co is an online trial and call api platform, which integrates SmolLM-135M-Instruct's modeling effects, including api services, and provides a free online trial of SmolLM-135M-Instruct, you can try SmolLM-135M-Instruct online for free by clicking the link below.

litert-community SmolLM-135M-Instruct online free url in huggingface.co:

https://huggingface.co/litert-community/SmolLM-135M-Instruct

SmolLM-135M-Instruct install

SmolLM-135M-Instruct is an open source model from GitHub that offers a free installation service, and any user can find SmolLM-135M-Instruct on GitHub to install. At the same time, huggingface.co provides the effect of SmolLM-135M-Instruct install, users can directly use SmolLM-135M-Instruct installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

SmolLM-135M-Instruct install url in huggingface.co:

https://huggingface.co/litert-community/SmolLM-135M-Instruct

Url of SmolLM-135M-Instruct

Provider of SmolLM-135M-Instruct huggingface.co

litert-community
ORGANIZATIONS

Other API from litert-community