| --- |
| license: mit |
| library_name: pytorch |
| tags: |
| - mixture-of-experts |
| - tiny-language-model |
| - tinystories |
| - text-generation |
| pipeline_tag: text-generation |
| --- |
| |
| # tlm26 — Mixture-of-Experts scaling family |
|
|
| Sparse Mixture-of-Experts language models trained on TinyStories, part of the |
| [tlm26](https://github.com/ANRGUSC/tlm26) project (sub-1M-parameter LMs for |
| microcontroller deployment; USC ANRG). This repository holds the MoE |
| expert-scaling family: a shared backbone with the expert count swept from 8 to |
| 128, each trained to a common 30,000-step budget. |
|
|
| ## Expert scaling |
|
|
| All models share a `d=64`, 5-layer backbone (8 heads, 4 KV heads, vocab 512, |
| seq 256) with **top-2** gating. Only the number of experts varies. Bits-per-byte |
| is measured on the held-out shard. |
|
|
| | Experts | CE | bits/byte | Total params | Active params/token | |
| |--------:|-------:|----------:|-------------:|--------------------:| |
| | 8 | 1.0153 | 0.7050 | 1.57M | 466K | |
| | 16 | 0.9474 | 0.6579 | 3.05M | 469K | |
| | 32 | 0.8811 | 0.6119 | 6.00M | 474K | |
| | 64 | 0.8256 | 0.5733 | 11.9M | 484K | |
| | 128 | 0.7872 | 0.5466 | 23.7M | 505K | |
|
|
| Each doubling of experts lowers bits/byte while active parameters per token stay |
| near-constant (466K → 505K): added sparse capacity improves quality at almost |
| fixed inference compute. Total parameters grow with the expert count, which sets |
| the flash footprint for on-device deployment. |
|
|
| A top-1 variant of the 8-expert model (`moe_d64_e8_top1`) is included for |
| comparison. |
|
|
| ## Files |
|
|
| Each `moe_d64_e{N}/model.pt` is a weights-only checkpoint (optimizer state |
| stripped) holding the `model` state dict and `model_args`. |
|
|
| ## Loading |
|
|
| ```python |
| import torch |
| ck = torch.load("moe_d64_e128/model.pt", map_location="cpu") |
| args = ck["model_args"] # dim, n_layers, n_experts, moe_top_k, ... |
| state = ck["model"] # load into the tlm26 MoE model |
| ``` |
|
|
| Model definition and training/eval code: https://github.com/ANRGUSC/tlm26 |
|
|