Instructions to use litert-community/SmolLM3-3B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LiteRT-LM
How to use litert-community/SmolLM3-3B with LiteRT-LM:
# LiteRT-LM runs on various platforms (Android, iOS, Windows, Linux, macOS, IoT, Web/WASM) # and supports many APIs (C++, Python, Kotlin, Swift, JavaScript, Flutter). # For platform-specific integration guides, please refer to the official developer website: # https://ai.google.dev/edge/litert-lm # To try LiteRT-LM, the easiest way is to use our CLI tool. # 1. Install the LiteRT-LM CLI tool: pip install -U litert-lm # 2. Download and run this model locally: # See: https://ai.google.dev/edge/litert-lm/cli litert-lm run \ --from-huggingface-repo=litert-community/SmolLM3-3B \ --prompt="Write me a poem"
- LiteRT
How to use litert-community/SmolLM3-3B with LiteRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
SmolLM3-3B β LiteRT-LM (blockwise int4)
HuggingFaceTB/SmolLM3-3B
converted to the LiteRT-LM (.litertlm) format for on-device inference with
Google's LiteRT-LM runtime (the
engine behind the official litert-community/* models).
SmolLM3 is a fully-open 3B decoder (Apache-2.0) with GQA, a NoPE attention schedule, multilingual support, and long-context training β a strong small reasoner.
| File | SmolLM3-3B_q4_block32_ekv4096.litertlm (SmolLM3-3B.litertlm ( |
| Quantization | int4 weights β blockwise (block 32) + OCTAV optimal-clipping, symmetric; embedding INT8 |
| Compute | integer |
| Context (KV cache) | 4096 |
| Base model | HuggingFaceTB/SmolLM3-3B |
Usage
Run with the LiteRT-LM runtime:
# build litert-lm from https://github.com/google-ai-edge/litert-lm, then:
litert_lm_main \
--model_path SmolLM3-3B_q4_block32_ekv4096.litertlm \
--backend gpu \
--input_prompt "Explain on-device AI in one sentence."
The .litertlm bundle carries the tokenizer and the prompt template (ChatML β
<|im_start|>role / <|im_end|>, stop token <|im_end|>), so no separate
tokenizer files are needed.
Run on Android
Update (July 2026): Google AI Edge Gallery v1.0.16+ can import litert-lm models directly from Hugging Face inside the app (tap +) β no computer or
adbneeded. The manual steps below are only required on older builds or for sideloading a local file.
The easiest way to try this model on a phone is the official
Google AI Edge Gallery app β it
runs .litertlm models fully on-device and can import your own:
- Install a recent Gallery (package
com.google.ai.edge.gallery, APK from the repo's releases β 1.0.15+ supports.litertlm). Older 1.0.x builds (packagecom.google.aiedge.gallery) only accept the legacy MediaPipe.taskformat and reject.litertlm. - Download
SmolLM3-3B_q4_block32_ekv4096.litertlmfrom this repo and push it to the device:adb push SmolLM3-3B_q4_block32_ekv4096.litertlm /sdcard/Download/ - In the app, tap the + button (bottom-right), pick the file, and choose the GPU backend (CPU also works).
- Chat. Nothing else to configure β the
.litertlmbundle already carries the tokenizer and ChatML prompt template.
See the Gallery
Importing Local Models
guide for details. To embed the model in your own Android app instead, use the
LiteRT-LM Kotlin API (Gradle artifact com.google.ai.edge.litertlm:litertlm-android,
getting started).
Measured on an 8 GB phone (added 2026-08-17): driving litert_lm_main directly on a Pixel 8a (Tensor G3, Mali-G715, 8 GB RAM), the graph runs entirely on the OpenCL delegate β 1476/1476 nodes in the 128-token prefill graph and 1308/1308 in decode, zero rejected ops β and answers correctly.
Run on desktop (LiteRT-LM CLI)
The same .litertlm bundle runs on macOS / Linux / Windows with the official
LiteRT-LM CLI β including as a
local OpenAI-compatible API server:
pip install litert-lm
litert-lm import --from-huggingface-repo litert-community/SmolLM3-3B SmolLM3-3B_q4_block32_ekv4096.litertlm smollm3-3b
litert-lm run smollm3-3b # interactive chat in the terminal
litert-lm serve # local OpenAI-compatible API server
Performance
litert-lm benchmark (litert-lm 0.15.0) on an Apple M4 Max, -p 256 -d 256 --runs 3 (the tool averages three iterations), max-num-tokens 4096, warm-up run discarded, otherwise idle machine.
| Device | Backend | Prefill (256) | Decode | TTFT | Load | Peak footprint |
|---|---|---|---|---|---|---|
| Apple M4 Max (macOS) | CPU | 141 tok/s | 24.1 tok/s | 2.14 s | β | β |
| Apple M4 Max (macOS) | GPU (Metal) | 1354 tok/s | 93.2 tok/s | 0.21 s | β | β |
| iPhone 17 Pro | GPU (Metal) | 30.8 tok/s | 22.5 tok/s | 0.63 s | 7.7 s | 1.24 GB |
Reproducibility: the GPU rows repeat to within about 1% across invocations; the CPU rows are noisier β re-running the 1B control six times spread its CPU decode over 29.0β33.3 tok/s, so treat the CPU column as accurate to roughly Β±7%.
The iPhone row is one cold run through the LiteRTDemo harness on iOS 27.0 (prompt "Explain on-device AI in one short sentence.", 512-token budget, no warm-up turn), read back from its run log. Its prefill figure is measured on that short prompt, so it reflects fixed per-turn overhead rather than prefill throughput and is not comparable to the 256-token desktop column.
Accuracy note
Measured on GSM8K (n=100, greedy, 0-shot chain-of-thought asking for #### <n>,
identical prompt and answer-extraction for both rows β only the quantization differs).
| Configuration | GSM8K |
|---|---|
| bf16 (reference) | 81.0% |
| This model β LiteRT int4 (BOCTAV4) | 81.0% |
LiteRT int4 is fully at parity β 0.0 pt vs the bf16 reference. The blockwise-32 +
OCTAV recipe with a 4096 KV cache preserves reasoning accuracy exactly at n=100. The
model produces visible step-by-step chain-of-thought in the answer body and
terminates cleanly at <|im_end|> (no rambling).
Galaxy S26 β GPU backend
Both published bundles run on the Android GPU backend and generate.
| file | GPU backend | delegation | peak |
|---|---|---|---|
SmolLM3-3B.litertlm |
runs | 2930 / 2930 ops across 2 subgraphs on LiteRT GPU |
704 MB |
SmolLM3-3B_q4_block32_ekv4096.litertlm |
runs | 2784 / 2784 ops across 2 subgraphs on LiteRT GPU |
1111 MB |
Measured on a Samsung Galaxy S26 (SM-S942Q / SM8850, Android 16) with litert_lm_advanced_main from litert-lm 0.16.0, --backend=gpu --sampler_backend=cpu, prompt What is the capital of France?. Peak is the process high-water mark (VmHWM) sampled during that same run. Gated 2026-08-24.
The op counts above are the LiteRT GPU partitions. In SmolLM3-3B_q4_block32_ekv4096.litertlm, XNNPACK additionally takes 1 of the 4 nodes in decode_embedder and 1 of the 4 nodes in prefill_embedder_128; the runtime accepts that split.
No speed rows, on purpose. On this handset the GPU backend wins prefill and does not win decode, so a GPU throughput figure only means something beside a CPU row from the same handset, and no S26 CPU row exists for this model yet.
GPU wiring, including the Gallery import toggle: GPU guide.
Conversion
Converted with litert-torch via its
generic export_hf path. SmolLM3ForCausalLM rides the existing converter with no
custom code: the NoPE attention schedule (rotary disabled on every 4th layer,
no_rope_layer_interval=4) lowers to generic ops with no custom kernel. The int4
recipe is blockwise (block 32) + OCTAV optimal-clipping with the embedding kept
at INT8; the embedding is externalized into its own bundle section so the main
weights section stays under the iOS ~2 GiB single-mmap limit. Blockwise (not
channelwise) int4 plus OCTAV is what holds reasoning accuracy at parity.
Training data & PII
This is a weights-exact format conversion of HuggingFaceTB/SmolLM3-3B; no new training was performed. SmolLM3 was trained by Hugging Face on ~11T tokens of publicly documented data β web (FineWeb-Edu, DCLM), code (StarCoder-family), math, and multilingual sources β plus public SFT/preference sets. Being web-derived it may incidentally contain PII; none was deliberately collected and this format conversion adds none. Apply your own content/PII filtering before deployment. See the base model card for the full data mixture.
2026-08-28 β start_token fix (weights unchanged)
The bundle's LlmMetadata start_token held the literal string "None". This tokenizer has no BOS, and the LiteRT-LM engine resolved that string to a real vocabulary token β so every prompt began with the word None, which the model was never trained on. The start token has been removed.
Metadata-only change: every section of the bundle except the LlmMetadata block is byte-identical to the previous file (verified by sha256 per section), so the weights, the graph and the tokenizer are unchanged and the speed and memory numbers on this card still describe exactly this file β only the file's own sha256 differs. What changed is the input: the token stream the model reads for a given conversation can differ from the previous file's, and it now matches this model's own reference chat stream. Greedy decoding can turn on a single token, so an individual answer can differ from the previous file in either direction. Unless a row says otherwise, the accuracy figures on this card were measured on the previous file and have not been re-measured on this one. If you downloaded before 2026-08-28, re-download.
2026-08-29 β default system prompt restored (weights unchanged)
The upstream chat template emits a default system turn whenever the caller sends no system message β for this model: the ## Metadata header (knowledge cutoff, today's date, Reasoning Mode: /think) followed by the instructions You are a helpful AI assistant named SmolLM, trained by Hugging Faceβ¦. The converter's template probe renders the template with a system message already present, so that block was never seen and never reached the bundle: with no system message the model was running without the default system turn it was tuned with. The chat template in SmolLM3-3B.litertlm, SmolLM3-3B_q4_block32_ekv4096.litertlm now emits the block exactly once when no system message is given. In SmolLM3-3B.litertlm, a system message you pass is wrapped in the same metadata header as upstream, and /no_think in it switches reasoning off. In SmolLM3-3B_q4_block32_ekv4096.litertlm, the block is not emitted when you pass a system message; the upstream template also wraps a caller's system message in its own preamble, and this file passes it through unchanged, exactly as it did before. The previous SmolLM3-3B.litertlm also forced an empty <think></think> before every answer (/no_think); upstream's default is /think, and the template now follows it. The restored block adds 240 prefill tokens to a conversation that sends no system message, so time-to-first-token grows by that much; per-token speed is unchanged.
Metadata-only change: every section of the bundle except the LlmMetadata block is byte-identical to the previous file (verified by sha256 per section), so the weights, the graph and the tokenizer are unchanged and the speed and memory numbers on this card still describe exactly this file per token β only the file's own sha256 differs. What changed is the input: with no system message, the prompt now renders byte-identical to the upstream chat template's output, verified on the LiteRT-LM runtime. A system message you pass yourself now renders through the upstream metadata header in SmolLM3-3B.litertlm; in SmolLM3-3B_q4_block32_ekv4096.litertlm it renders as before. Greedy decoding can turn on a single token, so an individual answer can differ from the previous file. Unless a row says otherwise, the accuracy figures on this card were measured on the previous file and have not been re-measured on this one. If you downloaded before 2026-08-29, re-download.
License
Apache-2.0, inherited from the base model HuggingFaceTB/SmolLM3-3B.
- Downloads last month
- 769
Model tree for litert-community/SmolLM3-3B
Base model
HuggingFaceTB/SmolLM3-3B-Base