MobileMoE-S-Base / README.md
yanbeic's picture
Add four-stage training recipe figure to model card
23b6f61 verified
|
Raw
History Blame Contribute Delete
9.3 kB
---
license: fair-noncommercial-research-license
extra_gated_fields:
First Name: text
Last Name: text
Date of birth: date_picker
Country: country
Affiliation: text
Job title:
type: select
options:
- Student
- Research Graduate
- AI researcher
- AI developer/engineer
- Reporter
- Other
geo: ip_location
By clicking Submit below I accept the terms of the license and acknowledge that the information I provide will be collected stored processed and shared in accordance with the Meta Privacy Policy: checkbox
extra_gated_description: >-
The information you provide will be collected, stored, processed and shared in
accordance with the [Meta Privacy
Policy](https://www.facebook.com/privacy/policy/).
extra_gated_button_content: Submit
language:
- en
library_name: transformers
tags:
- facebook
- meta
- pytorch
- mixture-of-experts
- MoE
- on-device
---
# MobileMoE-S (Base) Model Card
**MobileMoE** is a family of on-device Mixture-of-Experts (MoE) language models with sub-billion *active* parameters, designed to push the quality–efficiency Pareto frontier for on-device LLMs, including three model scales (S/M/L): 0.3B/0.5B/0.9B active parameters (1.3B/2.8B/5.3B total), with <3 GB INT4 weight footprints to fit in mobile DRAM. Each scale is released in three variants: a **Base** model (pre-training + mid-training), an **SFT** model (supervised fine-tuning), and a **QAT** model (quantization-aware training). You are currently in the **MobileMoE-S-Base** repository — the pre-trained 0.3B-active base model.
| | **S** | **M** | **L** |
|:---|:---:|:---:|:---:|
| Active / total params | 272M / 1.3B | 528M / 2.8B | 922M / 5.3B |
| Layers | 20 | 26 | 32 |
| Model dimension | 768 | 1024 | 1280 |
| Heads (Q / KV) | 12 / 4 | 16 / 4 | 20 / 4 |
| Routed experts | 60 | 60 | 60 |
| Top-k | 4 | 4 | 4 |
| INT4 weight memory | 0.68 GB | 1.48 GB | 2.75 GB |
| Base | **MobileMoE-S-Base** | [MobileMoE-M-Base](https://huggingface.co/facebook/MobileMoE-M-Base) | [MobileMoE-L-Base](https://huggingface.co/facebook/MobileMoE-L-Base) |
| SFT | [MobileMoE-S-SFT](https://huggingface.co/facebook/MobileMoE-S-SFT) | [MobileMoE-M-SFT](https://huggingface.co/facebook/MobileMoE-M-SFT) | [MobileMoE-L-SFT](https://huggingface.co/facebook/MobileMoE-L-SFT) |
| QAT (INT4) | [MobileMoE-S-QAT](https://huggingface.co/facebook/MobileMoE-S-QAT) | [MobileMoE-M-QAT](https://huggingface.co/facebook/MobileMoE-M-QAT) | [MobileMoE-L-QAT](https://huggingface.co/facebook/MobileMoE-L-QAT) |
For the detailed technical report: 📝 [MobileMoE: Scaling On-Device Mixture of Experts](https://arxiv.org/abs/2605.27358)
For more versions, check out the 🤗 [MobileMoE Collection](https://huggingface.co/collections/facebook/mobilemoe)
![MobileMoE establishes a new Pareto frontier for on-device LLMs](mobilemoe_pareto.png)
**MobileMoE establishes a new Pareto frontier for on-device LLMs.** Average benchmark accuracy, computed over 14 benchmarks spanning commonsense, knowledge, science, comprehension, and reasoning, is plotted against (a) per-token inference compute *F*<sub>inf</sub> = 2*N*<sub>act</sub> (GFLOPs) and (b) total parameters *N*<sub>total</sub> (B); in (b), x-axis tick labels show total params (B) | projected INT4 memory (GB).
## Key Features
- **A new Pareto frontier for on-device LLMs.** Across 14 foundational benchmarks, MobileMoE matches or exceeds leading on-device dense LLMs at 2–4× fewer inference FLOPs, and matches or surpasses the state-of-the-art MoE OLMoE-1B-7B with up to 60% fewer parameters.
- **Scaling-law-derived architecture.** The architecture is derived from an on-device MoE scaling law that jointly optimizes under mobile *memory* and *compute* constraints, identifying an on-device sweet spot: moderate sparsity, with fine-grained experts and shared expert.
- **Four-stage recipe.** Pre-training → mid-training → instruction fine-tuning → INT4 quantization-aware training, all on open-source datasets.
## Model Information
**Model:** MobileMoE-S-Base (pre-trained + mid-trained)<br>
**Active Parameters:** 272M<br>
**Total Parameters:** 1.3B<br>
**Layers:** 20<br>
**Model Dimension:** 768<br>
**Attention Heads:** 12<br>
**KV Heads:** 4 (GQA)<br>
**Head Dimension:** 64<br>
**Routed Experts:** 60 (fine-grained, FFN hidden dim 384 each)<br>
**Active Experts per Token:** 4 (top-k sigmoid routing, with normalization)<br>
**Shared Expert:** 1, always on (FFN hidden dim 1536)<br>
**Vocabulary Size:** 128,256<br>
**Other Features:** QK-Norm, tied input/output embeddings, RoPE (θ = 500,000)<br>
**Input Modality:** Text<br>
**Output Modality:** Text<br>
**Languages:** English<br>
**Training Stages:** Pre-training → mid-training<br>
**Context Length:** 8,192 tokens<br>
**Precision:** BF16<br>
**Model Developer:** Meta<br>
**Model Release Date:** Aug 2026<br>
**License:** MobileMoE is FAIR NC licensed
## Results
All numbers below are for the **base (pre-trained)** models, re-evaluated under identical settings with greedy decoding using [`lm-eval`](https://github.com/EleutherAI/lm-evaluation-harness); few-shot counts are given in parentheses after the benchmark name, and benchmarks shown without one are evaluated 0-shot.
### Foundational benchmarks
| Capability | Benchmark | Gemma 3 270M | SmolLM2 360M | **MobileMoE-S** |
|:---|:---|:---:|:---:|:---:|
| *Active / total params* | | 270M | 362M | **272M / 1.3B** |
| Commonsense Reasoning | HellaSwag | 41.4 | 56.5 | **58.9** |
| | PIQA | 68.3 | 71.7 | **75.4** |
| | SIQA | 40.2 | 40.7 | **46.8** |
| | WinoGrande | 53.7 | **59.0** | 58.6 |
| Knowledge | MMLU (5-shot) | 26.7 | 25.2 | **43.7** |
| | NaturalQuestions (5-shot) | 4.1 | 7.4 | **12.6** |
| | TriviaQA (5-shot) | 14.3 | 26.8 | **33.2** |
| Science | ARC-Challenge (25-shot) | 29.4 | 40.5 | **46.5** |
| | ARC-Easy | 56.8 | 68.1 | **73.9** |
| | OpenBookQA | 30.4 | **37.6** | 34.6 |
| Reading | BoolQ | 58.3 | **61.8** | 60.2 |
| | DROP (3-shot) | 14.2 | 17.9 | **39.0** |
| Reasoning | BIG-Bench Hard (3-shot) | 29.5 | 31.7 | **31.8** |
| | GSM8K (8-shot) | 1.8 | 5.3 | **36.2** |
| **Average** | | 33.5 | 39.3 | **46.5** |
## Training
MobileMoE uses a four-stage recipe. This checkpoint is the output of stage 2 (mid-training).
![MobileMoE four-stage training recipe](mobilemoe_recipe.png)
**MobileMoE four-stage training recipe:** pre-training (PT) → mid-training (MT) → instruct supervised fine-tuning (SFT) → quantization-aware training (QAT) with INT4 precision.
| | **Pre-training** | **Mid-training** | SFT | QAT |
|---|---|---|---|---|
| Context length | **2,048** | **8,192** | 8,192 | 8,192 |
| Total tokens | **~6T** | **~500B** | ~126B | ~21B |
| Peak learning rate | **4×10<sup>-4</sup>** | **4×10<sup>-5</sup>** | 4×10<sup>-6</sup> | 4×10<sup>-6</sup> |
| LR schedule | **Cosine** | **Linear** | Cosine | Cosine |
| Token dispatch | **drop-and-pad** | **drop-and-pad** | dropless | dropless |
## How to use
MobileMoE uses a custom architecture (`model_type: mobilemoe`) that is not yet part of upstream `transformers`, so **`trust_remote_code=True` is required**. The modeling code ships in this repo (`configuration_mobilemoe.py`, `modeling_mobilemoe.py`).
### Requirements
```bash
pip install "torch>=2.1" "transformers>=4.57" "safetensors>=0.4" "accelerate>=1.0"
```
Verified with the following versions:
| Package | Version |
|---|---|
| `torch` | 2.8.0 (cu128) |
| `transformers` | 4.57.6 |
| `tokenizers` | 0.22.2 |
| `safetensors` | 0.7.0 |
| `accelerate` | 1.13.0 |
For batch evaluation we recommend vLLM (≥ 0.10.2) with `enforce_eager=True`.
### Text generation
This is a **base model** — it has no chat template and is not instruction-tuned. Prompt it with plain text continuation:
```python
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
MODEL_ID = "facebook/MobileMoE-S-Base"
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
model = AutoModelForCausalLM.from_pretrained(
MODEL_ID,
trust_remote_code=True,
dtype=torch.bfloat16,
)
model.to("cuda" if torch.cuda.is_available() else "cpu")
model.eval()
prompt = "Why are open-source on-device language models great?"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(
input_ids=inputs["input_ids"],
attention_mask=inputs["attention_mask"],
max_new_tokens=64,
do_sample=False,
temperature=None,
top_p=None,
pad_token_id=tokenizer.eos_token_id,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```
**Known issues.** Loading the tokenizer on `transformers` 4.57.6 prints a `fix_mistral_regex=True` warning. Please ignore it and do not set the flag, as MobileMoE uses the Llama-3 tokenizer whose default tokenization is already correct.
## Citation
```bibtex
@article{chen2026mobilemoe,
title={MobileMoE: Scaling On-Device Mixture of Experts},
author={Chen, Yanbei and Huang, Hanxian and Chang, Ernie and Szwejbka, Jacob and Desai, Digant and Liu, Zechun and Chandra, Vikas and Krishnamoorthi, Raghuraman},
journal={arXiv preprint arXiv:2605.27358},
year={2026}
}
```
## License
MobileMoE is distributed under the [FAIR Noncommercial Research License](https://huggingface.co/facebook/MobileMoE-S-Base/blob/main/LICENSE).