SmolLM2-135M-Instruct for ExecuTorch

ExecuTorch exports of HuggingFaceTB/SmolLM2-135M-Instruct (revision 12fd25f77366) for on-device inference with the openweights Android app or any ExecuTorch 1.4.0 runtime.

Files

Windows in this repo: XNNPACK (CPU) at 2k to 32k; Vulkan (GPU) at 2k to 32k; MediaTek NeuroPilot MT6991 (Dimensity 9400) at 2k to 16k. The window is fixed inside the file: the runtime allocates the whole KV cache at load, so pick the largest window the device can hold (fits_phone_budget in each folder's config.json is the estimate against a 5 GB budget).

Backend Target File Window Size Smoke test
XNNPACK (CPU) any arm64 xnnpack/SmolLM2-135M-Instruct-8da4w-gptq-2k.pte 2,048 tokens 0.11 GB passed ("Paris")
XNNPACK (CPU) any arm64 xnnpack/SmolLM2-135M-Instruct-8da4w-gptq-4k.pte 4,096 tokens 0.11 GB passed ("Paris")
XNNPACK (CPU) any arm64 xnnpack/SmolLM2-135M-Instruct-8da4w-gptq-8k.pte 8,192 tokens 0.11 GB passed ("Paris")
XNNPACK (CPU) any arm64 xnnpack/SmolLM2-135M-Instruct-8da4w-gptq-16k.pte 16,384 tokens 0.11 GB passed ("Paris")
XNNPACK (CPU) any arm64 xnnpack/SmolLM2-135M-Instruct-8da4w-gptq-32k.pte 32,768 tokens 0.12 GB passed ("Paris")
Vulkan (GPU) any arm64 vulkan/SmolLM2-135M-Instruct-vulkan-8da4w-2k.pte 2,048 tokens 0.13 GB structure checked (no host NPU runtime)
Vulkan (GPU) any arm64 vulkan/SmolLM2-135M-Instruct-vulkan-8da4w-4k.pte 4,096 tokens 0.14 GB structure checked (no host NPU runtime)
Vulkan (GPU) any arm64 vulkan/SmolLM2-135M-Instruct-vulkan-8da4w-8k.pte 8,192 tokens 0.14 GB structure checked (no host NPU runtime)
Vulkan (GPU) any arm64 vulkan/SmolLM2-135M-Instruct-vulkan-8da4w-16k.pte 16,384 tokens 0.15 GB structure checked (no host NPU runtime)
Vulkan (GPU) any arm64 vulkan/SmolLM2-135M-Instruct-vulkan-8da4w-32k.pte 32,768 tokens 0.17 GB structure checked (no host NPU runtime)
MediaTek NeuroPilot MT6991 (Dimensity 9400) mtk/mt6991/SmolLM2-135M-Instruct-neuropilot-a16w8-2k-chunk1of3.pte 2,048 tokens 0.04 GB structure checked (no host NPU runtime)
MediaTek NeuroPilot MT6991 (Dimensity 9400) mtk/mt6991/SmolLM2-135M-Instruct-neuropilot-a16w8-2k-chunk2of3.pte 2,048 tokens 0.04 GB structure checked (no host NPU runtime)
MediaTek NeuroPilot MT6991 (Dimensity 9400) mtk/mt6991/SmolLM2-135M-Instruct-neuropilot-a16w8-2k-chunk3of3.pte 2,048 tokens 0.07 GB structure checked (no host NPU runtime)
MediaTek NeuroPilot MT6991 (Dimensity 9400) mtk/mt6991/SmolLM2-135M-Instruct-neuropilot-a16w8-4k-chunk1of3.pte 4,096 tokens 0.04 GB structure checked (no host NPU runtime)
MediaTek NeuroPilot MT6991 (Dimensity 9400) mtk/mt6991/SmolLM2-135M-Instruct-neuropilot-a16w8-4k-chunk2of3.pte 4,096 tokens 0.04 GB structure checked (no host NPU runtime)
MediaTek NeuroPilot MT6991 (Dimensity 9400) mtk/mt6991/SmolLM2-135M-Instruct-neuropilot-a16w8-4k-chunk3of3.pte 4,096 tokens 0.07 GB structure checked (no host NPU runtime)
MediaTek NeuroPilot MT6991 (Dimensity 9400) mtk/mt6991/SmolLM2-135M-Instruct-neuropilot-a16w8-8k-chunk1of3.pte 8,192 tokens 0.04 GB structure checked (no host NPU runtime)
MediaTek NeuroPilot MT6991 (Dimensity 9400) mtk/mt6991/SmolLM2-135M-Instruct-neuropilot-a16w8-8k-chunk2of3.pte 8,192 tokens 0.04 GB structure checked (no host NPU runtime)
MediaTek NeuroPilot MT6991 (Dimensity 9400) mtk/mt6991/SmolLM2-135M-Instruct-neuropilot-a16w8-8k-chunk3of3.pte 8,192 tokens 0.07 GB structure checked (no host NPU runtime)
MediaTek NeuroPilot MT6991 (Dimensity 9400) mtk/mt6991/SmolLM2-135M-Instruct-neuropilot-a16w8-16k-chunk1of3.pte 16,384 tokens 0.05 GB structure checked (no host NPU runtime)
MediaTek NeuroPilot MT6991 (Dimensity 9400) mtk/mt6991/SmolLM2-135M-Instruct-neuropilot-a16w8-16k-chunk2of3.pte 16,384 tokens 0.05 GB structure checked (no host NPU runtime)
MediaTek NeuroPilot MT6991 (Dimensity 9400) mtk/mt6991/SmolLM2-135M-Instruct-neuropilot-a16w8-16k-chunk3of3.pte 16,384 tokens 0.07 GB structure checked (no host NPU runtime)

MediaTek folders also hold the token embedding table the NeuroPilot runner reads from disk, shared by every window: mtk/mt6991/SmolLM2-135M-Instruct-neuropilot-embedding-fp32.bin.

Tokenizer: tokenizer.json, copied unchanged from the source repo. Each backend folder has a config.json listing every window as a variant with the metadata the .pte reports, and an export-report-<window>.json per file with the full export record.

Memory

  • XNNPACK (CPU) at 2,048 tokens: the KV cache costs 46,080 bytes per token (fp32), 94,371,840 bytes for the whole window, allocated in full when the model loads.
  • XNNPACK (CPU) at 4,096 tokens: the KV cache costs 46,080 bytes per token (fp32), 188,743,680 bytes for the whole window, allocated in full when the model loads.
  • XNNPACK (CPU) at 8,192 tokens: the KV cache costs 46,080 bytes per token (fp32), 377,487,360 bytes for the whole window, allocated in full when the model loads.
  • XNNPACK (CPU) at 16,384 tokens: the KV cache costs 46,080 bytes per token (fp32), 754,974,720 bytes for the whole window, allocated in full when the model loads.
  • XNNPACK (CPU) at 32,768 tokens: the KV cache costs 46,080 bytes per token (fp32), 1,509,949,440 bytes for the whole window, allocated in full when the model loads.
  • Vulkan (GPU) at 2,048 tokens: the KV cache costs 46,080 bytes per token (fp32), 94,371,840 bytes for the whole window, allocated in full when the model loads.
  • Vulkan (GPU) at 4,096 tokens: the KV cache costs 46,080 bytes per token (fp32), 188,743,680 bytes for the whole window, allocated in full when the model loads.
  • Vulkan (GPU) at 8,192 tokens: the KV cache costs 46,080 bytes per token (fp32), 377,487,360 bytes for the whole window, allocated in full when the model loads.
  • Vulkan (GPU) at 16,384 tokens: the KV cache costs 46,080 bytes per token (fp32), 754,974,720 bytes for the whole window, allocated in full when the model loads.
  • Vulkan (GPU) at 32,768 tokens: the KV cache costs 46,080 bytes per token (fp32), 1,509,949,440 bytes for the whole window, allocated in full when the model loads.

How it was made

  • XNNPACK (CPU) any arm64: ExecuTorch 1.4.0 export_llm: 8-bit dynamic activations and 4-bit weights in groups of 32, int8 per-channel embeddings, XNNPACK with extended ops, prefill chunk 2048, fp32 KV cache. Built by run 1.
  • Vulkan (GPU) any arm64: ExecuTorch 1.4.0 export_llm: 8-bit dynamic activations and 4-bit weights in groups of 32, int8 per-channel embeddings, the Vulkan delegate, prefill chunk 2048, fp32 KV cache. Built by run 1.
  • MediaTek NeuroPilot MT6991 (Dimensity 9400): ExecuTorch 1.4.0 MediaTek LLM export (examples/mediatek, llama.py): A16W8 (16-bit activations, 8-bit weights) calibrated on MediaTek's alpaca.txt prompts (9 of 9) in the qwen.json chat template, cut into 3 chunks, with a 128-token prompt graph and a one-token generation graph over a 2048-token cache, compiled with MediaTek NeuroPilot Express SDK (mtk_converter 8.13.0+public, mtk_neuron 8.2.23) for MT6991 (Dimensity 9400). Built by run 1.
  • MediaTek NeuroPilot MT6991 (Dimensity 9400): ExecuTorch 1.4.0 MediaTek LLM export (examples/mediatek, llama.py): A16W8 (16-bit activations, 8-bit weights) calibrated on MediaTek's alpaca.txt prompts (9 of 9) in the qwen.json chat template, cut into 3 chunks, with a 128-token prompt graph and a one-token generation graph over a 4096-token cache, compiled with MediaTek NeuroPilot Express SDK (mtk_converter 8.13.0+public, mtk_neuron 8.2.23) for MT6991 (Dimensity 9400). Built by run 1.
  • MediaTek NeuroPilot MT6991 (Dimensity 9400): ExecuTorch 1.4.0 MediaTek LLM export (examples/mediatek, llama.py): A16W8 (16-bit activations, 8-bit weights) calibrated on MediaTek's alpaca.txt prompts (9 of 9) in the qwen.json chat template, cut into 3 chunks, with a 128-token prompt graph and a one-token generation graph over a 8192-token cache, compiled with MediaTek NeuroPilot Express SDK (mtk_converter 8.13.0+public, mtk_neuron 8.2.23) for MT6991 (Dimensity 9400). Built by run 1.
  • MediaTek NeuroPilot MT6991 (Dimensity 9400): ExecuTorch 1.4.0 MediaTek LLM export (examples/mediatek, llama.py): A16W8 (16-bit activations, 8-bit weights) calibrated on MediaTek's alpaca.txt prompts (9 of 9) in the qwen.json chat template, cut into 3 chunks, with a 128-token prompt graph and a one-token generation graph over a 16384-token cache, compiled with MediaTek NeuroPilot Express SDK (mtk_converter 8.13.0+public, mtk_neuron 8.2.23) for MT6991 (Dimensity 9400). Built by run 1.

License

A quantized derivative of HuggingFaceTB/SmolLM2-135M-Instruct, distributed under the same terms (apache-2.0).

The mtk/ folders hold model binaries compiled with the MediaTek NeuroPilot Express SDK 8.0.8-build20250925 (mtk_converter 8.13.0+public, mtk_neuron 8.2.23) from MediaTek Inc., used under MediaTek's license terms for that SDK. No MediaTek SDK or runtime library is included. They run on MediaTek's LLM runner from ExecuTorch (examples/mediatek/executor_runner) with the device's NeuroPilot runtime; the settings it needs are in each folder's config.json, under each variant's runner.

Downloads last month
227
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for experimentalmachines/SmolLM2-135M-Instruct-ExecuTorch

Quantized
(134)
this model

Collection including experimentalmachines/SmolLM2-135M-Instruct-ExecuTorch