tlm26 / README.md
BradyMatt
Upload folder using huggingface_hub
804c213 verified
|
Raw
History Blame Contribute Delete
2.1 kB
---
license: mit
library_name: pytorch
tags:
- mixture-of-experts
- tiny-language-model
- tinystories
- text-generation
pipeline_tag: text-generation
---
# tlm26 — Mixture-of-Experts scaling family
Sparse Mixture-of-Experts language models trained on TinyStories, part of the
[tlm26](https://github.com/ANRGUSC/tlm26) project (sub-1M-parameter LMs for
microcontroller deployment; USC ANRG). This repository holds the MoE
expert-scaling family: a shared backbone with the expert count swept from 8 to
128, each trained to a common 30,000-step budget.
## Expert scaling
All models share a `d=64`, 5-layer backbone (8 heads, 4 KV heads, vocab 512,
seq 256) with **top-2** gating. Only the number of experts varies. Bits-per-byte
is measured on the held-out shard.
| Experts | CE | bits/byte | Total params | Active params/token |
|--------:|-------:|----------:|-------------:|--------------------:|
| 8 | 1.0153 | 0.7050 | 1.57M | 466K |
| 16 | 0.9474 | 0.6579 | 3.05M | 469K |
| 32 | 0.8811 | 0.6119 | 6.00M | 474K |
| 64 | 0.8256 | 0.5733 | 11.9M | 484K |
| 128 | 0.7872 | 0.5466 | 23.7M | 505K |
Each doubling of experts lowers bits/byte while active parameters per token stay
near-constant (466K → 505K): added sparse capacity improves quality at almost
fixed inference compute. Total parameters grow with the expert count, which sets
the flash footprint for on-device deployment.
A top-1 variant of the 8-expert model (`moe_d64_e8_top1`) is included for
comparison.
## Files
Each `moe_d64_e{N}/model.pt` is a weights-only checkpoint (optimizer state
stripped) holding the `model` state dict and `model_args`.
## Loading
```python
import torch
ck = torch.load("moe_d64_e128/model.pt", map_location="cpu")
args = ck["model_args"] # dim, n_layers, n_experts, moe_top_k, ...
state = ck["model"] # load into the tlm26 MoE model
```
Model definition and training/eval code: https://github.com/ANRGUSC/tlm26