MoE-Study / README.md
OliverSundaram's picture
Add dense and top-2-of-4 MoE checkpoints, tokenizer, and benchmark assets
d83d5d1 verified
|
Raw
History Blame Contribute Delete
9.13 kB
---
license: mit
language:
- en
library_name: transformers
pipeline_tag: text-generation
inference: false
datasets:
- nampdn-ai/tiny-textbooks
tags:
- mixture-of-experts
- moe
- from-scratch
- ablation
- research
---
# MoE-Study β€” Dense vs. Mixture-of-Experts, matched active parameters
Two decoder-only language models trained **from scratch** under identical conditions, differing in
exactly one thing: whether the feed-forward block is a **dense MLP** or a **sparse top-2-of-4 MoE**.
Both checkpoints live in this one repo:
| Subfolder | Model | Total params | Active params/token |
|---------------------|----------------|--------------|---------------------|
| [`dense/`](./dense) | Dense FFN | 150.1M | 150.1M |
| [`moe/`](./moe) | Top-2-of-4 MoE | 206.8M | ~150.1M |
The MoE's active-parameter count matches Dense **by construction** β€” 2 of 4 experts at half the hidden
size means identical compute per token. The MoE only spends more *memory* for extra capacity.
Full write-up, training code, and evaluation harness:
**[github.com/OliverSundaram/MoE-Study](https://github.com/OliverSundaram/MoE-Study)**
---
## ⚠️ These are research artifacts, not usable models
Read this before downloading.
- Trained for **one epoch** on ~40.7M tokens β€” neither model is close to converged.
- **WikiText word perplexity is 551 (Dense) and 1,378 (MoE).** Generations are largely incoherent.
- **0.0% on LAMBADA** for both β€” at the task floor.
- No instruction tuning, no RLHF, no safety filtering of any kind.
They exist to answer one narrow question: *at matched active compute and matched budget, does sparsity
help?* They are not fit for any downstream use.
---
## Getting the weights
These are a custom architecture, not a variant of an existing one. The modeling code is not included
here, so `from_pretrained` on this repo alone will not build the model.
Clone [the GitHub repo](https://github.com/OliverSundaram/MoE-Study) β€” it carries the model definition
and loading instructions, and points back at these subfolders for the weights.
---
## Model details
### Shared architecture
Both models are the same custom decoder-only transformer:
| | |
|---------------------|--------------------------------------------------------------------|
| Layers | 12 |
| Attention heads | 12 |
| Embedding dim | 768 |
| Context length | 1024 |
| Vocabulary | 50,257 (GPT-2 tokenizer) |
| Attention | **Multi-Query** β€” one shared K/V projection across all query heads |
| Normalization | Custom pre-norm (learned scale + shift) |
| Position embeddings | Learned absolute |
| Weight tying | None β€” separate input embedding and output head |
### The one difference
| | `dense/` | `moe/` |
|--------------|------------------|------------------------------------------------|
| FFN block | 2-layer GELU MLP | 4 experts, top-2 routed |
| `hidden_dim` | 3072 | 1536 (per expert) |
| Router | β€” | linear β†’ softmax β†’ top-2, renormalized |
| Aux loss | β€” | load-balancing term, summed over all 12 layers |
Both models share the **same** unmodified GPT-2 tokenizer, stored once at the repo root.
---
## Training
Identical for both models. Single consumer GPU, no cloud.
| Setting | Value |
|---------------|----------------------------------------------------------------------------------------|
| Data | [`nampdn-ai/tiny-textbooks`](https://huggingface.co/datasets/nampdn-ai/tiny-textbooks) |
| Tokens | 39,717 chunks Γ— 1024 = **~40.67M** |
| Epochs | **1** (19,858 steps) |
| Batch size | 2 Γ— grad accum 4 = effective **8** |
| Optimizer | AdamW, lr `3e-4`, weight decay `0.1` (no decay on 1-D params) |
| Schedule | `OneCycleLR`, cosine, 3% warmup |
| Grad clipping | max-norm `1.0` |
| Precision | AMP autocast + `GradScaler` |
| Seed | 42 |
| Hardware | 1Γ— NVIDIA RTX 4060, 8 GB VRAM |
| Wall-clock | ~44.6 min (Dense) Β· ~59.8 min (MoE) |
### Final losses
| | Dense | MoE |
|----------------------------|-----------|-----------|
| Train loss (final step) | 5.166 | 5.936 |
| **Test loss (pure LM)** | **5.063** | **5.911** |
| Test loss (+ unscaled aux) | n/a | 17.91 |
Dense has the lower loss at **every** checkpoint.
---
## Evaluation
All benchmarks via [lm-evaluation-harness](https://github.com/EleutherAI/lm-evaluation-harness) on the
final checkpoints.
| Benchmark | Shots | Metric | Dense | MoE | abs(Ξ”) | Winner |
|------------------|-------|-------------------|-----------|-----------|--------|----------|
| ARC-Easy | 0 | `acc` | **29.2%** | 27.4% | 1.8 | πŸ”΅ Dense |
| PIQA | 0 | `acc` | **55.0%** | 54.1% | 0.9 | πŸ”΅ Dense |
| WikiText | 0 | `word_perplexity` | **551.0** | 1,377.8 | 826.8 | πŸ”΅ Dense |
| LAMBADA (OpenAI) | 0 | `acc` | 0.0% | 0.0% | 0.0 | βšͺ Tie |
| WinoGrande | 5 | `acc` | 50.2% | **50.7%** | 0.5 | 🟠 MoE |
| HellaSwag | 10 | `acc_norm` | 24.9% | **25.1%** | 0.2 | 🟠 MoE |
| ARC-Challenge | 25 | `acc_norm` | 22.9% | **23.0%** | 0.1 | 🟠 MoE |
**How to read this:**
- Dense wins on everything sensitive to raw LLM quality β€” perplexity, ARC-Easy, PIQA.
- WinoGrande, HellaSwag, and ARC-Challenge are won by MoE, but with such a negligible difference, that they could be considered to have an equal accuracy
### Inference speed
Greedy decoding, 32-token prompt β†’ 64 new tokens, 5 trials, 2 warmup, no KV cache.
| Model | Tokens/sec | Total params | Active params/token |
|-------|-------------------|--------------|---------------------|
| Dense | **106.49 Β± 0.30** | 150.1M | 150.1M |
| MoE | 34.40 Β± 0.08 | 206.8M | ~150.1M |
MoE is **~3.1Γ— slower** despite matched active compute β€” an artifact of unoptimized expert dispatch, not
a property of the architecture.
<details>
<summary><b>Benchmark charts</b></summary>
![ARC-Easy](./assets/arc_easy.png)
![PIQA](./assets/piqa.png)
![WikiText](./assets/wikitext.png)
![LAMBADA](./assets/lambada_openai.png)
![WinoGrande](./assets/winogrande.png)
![HellaSwag](./assets/hellaswag.png)
![ARC-Challenge](./assets/arc_challenge.png)
![Speed](./assets/speed.png)
</details>
---
## Findings
**1. Dense won every metric that wasn't already at chance.**
Most clearly on WikiText perplexity β€” 551 vs 1,378, a 2.5Γ— gap.
**2. The routing math is correct.**
Active parameters match Dense almost exactly. Matched active compute simply didn't buy matched quality
at this budget.
**3. Routing stayed balanced.**
The load-balancing term sat on its theoretical floor, so the MoE's gap is not explained by experts
collapsing onto each other.
**4. Extra capacity needs extra tokens.**
The MoE has 38% more parameters but saw the same ~40.7M tokens β€” likely far too few to train 4 experts
per layer, each seeing only a routed fraction of the stream.
---
## Citation
```bibtex
@misc{sundaram2026moestudy,
author = {Sundaram, Oliver},
title = {MoE-Study: Dense vs. Mixture-of-Experts at Matched Active Parameters},
year = {2026},
url = {https://github.com/OliverSundaram/MoE-Study}
}
```
## Acknowledgments
- [lm-evaluation-harness](https://github.com/EleutherAI/lm-evaluation-harness) (EleutherAI) β€” evaluation
- [nampdn-ai/tiny-textbooks](https://huggingface.co/datasets/nampdn-ai/tiny-textbooks) β€” training corpus
- [Hugging Face `transformers`](https://github.com/huggingface/transformers) β€” base classes and tokenizer
## License
MIT