Instructions to use litert-community/SmolLM2-135M-Instruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LiteRT-LM
How to use litert-community/SmolLM2-135M-Instruct with LiteRT-LM:
# LiteRT-LM runs on various platforms (Android, iOS, Windows, Linux, macOS, IoT, Web/WASM) # and supports many APIs (C++, Python, Kotlin, Swift, JavaScript, Flutter). # For platform-specific integration guides, please refer to the official developer website: # https://ai.google.dev/edge/litert-lm # To try LiteRT-LM, the easiest way is to use our CLI tool. # 1. Install the LiteRT-LM CLI tool: pip install -U litert-lm # 2. Download and run this model locally: # See: https://ai.google.dev/edge/litert-lm/cli litert-lm run \ --from-huggingface-repo=litert-community/SmolLM2-135M-Instruct \ --prompt="Write me a poem"
- Notebooks
- Google Colab
- Kaggle
metadata
license: apache-2.0
base_model: HuggingFaceTB/SmolLM2-135M-Instruct
pipeline_tag: text-generation
library_name: litert-lm
tags:
- chat
- litert-lm
- smollm
- on-device
litert-community/SmolLM2-135M-Instruct
This model provides a variant of HuggingFaceTB/SmolLM2-135M-Instruct that is ready for deployment on Android using the LiteRT-LM.
Use the model
Android
Edge Gallery App
Download or build the app from GitHub.
Install the app from Google Play.
Follow the instructions in the app.
To build the demo app from source, please follow the instructions from the GitHub repository.
Performance (Apple M4 Max, measured)
Measured with the LiteRT-LM CLI: litert-lm benchmark -p 256 -d 256 --runs 3 --cache no
(litert-lm 0.15.0) on an idle Apple M4 Max (macOS); 256 prefill / 256 decode tokens, 3 iterations
averaged by the tool. A desktop reference point — phone-side figures vary by SoC and backend.
| Backend | Prefill (tokens/s) | Decode (tokens/s) | Time-to-first-token (s) |
|---|---|---|---|
| CPU | 1,698 | 104.9 | 0.22 |
| GPU | 7,571 | 259.6 | 0.04 |