Instructions to use facebook/MobileMoE-S-SFT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use facebook/MobileMoE-S-SFT with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="facebook/MobileMoE-S-SFT", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("facebook/MobileMoE-S-SFT", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use facebook/MobileMoE-S-SFT with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "facebook/MobileMoE-S-SFT" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "facebook/MobileMoE-S-SFT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/facebook/MobileMoE-S-SFT
- SGLang
How to use facebook/MobileMoE-S-SFT with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "facebook/MobileMoE-S-SFT" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "facebook/MobileMoE-S-SFT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "facebook/MobileMoE-S-SFT" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "facebook/MobileMoE-S-SFT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use facebook/MobileMoE-S-SFT with Docker Model Runner:
docker model run hf.co/facebook/MobileMoE-S-SFT
You need to agree to share your contact information to access this model
The information you provide will be collected, stored, processed and shared in accordance with the Meta Privacy Policy.
Log in or Sign Up to review the conditions and access this model content.
MobileMoE-S (SFT) Model Card
MobileMoE is a family of on-device Mixture-of-Experts (MoE) language models with sub-billion active parameters, designed to push the quality–efficiency Pareto frontier for on-device LLMs, including three model scales (S/M/L): 0.3B/0.5B/0.9B active parameters (1.3B/2.8B/5.3B total), with <3 GB INT4 weight footprints to fit in mobile DRAM. Each scale is released in three variants: a Base model (pre-training + mid-training), an SFT model (supervised fine-tuning), and a QAT model (quantization-aware training). You are currently in the MobileMoE-S-SFT repository — the instruction-tuned 0.3B-active model.
| S | M | L | |
|---|---|---|---|
| Active / total params | 272M / 1.3B | 528M / 2.8B | 922M / 5.3B |
| Layers | 20 | 26 | 32 |
| Model dimension | 768 | 1024 | 1280 |
| Heads (Q / KV) | 12 / 4 | 16 / 4 | 20 / 4 |
| Routed experts | 60 | 60 | 60 |
| Top-k | 4 | 4 | 4 |
| INT4 weight memory | 0.68 GB | 1.48 GB | 2.75 GB |
| Base | MobileMoE-S-Base | MobileMoE-M-Base | MobileMoE-L-Base |
| SFT | MobileMoE-S-SFT | MobileMoE-M-SFT | MobileMoE-L-SFT |
| QAT (INT4) | MobileMoE-S-QAT | MobileMoE-M-QAT | MobileMoE-L-QAT |
For the detailed technical report: 📝 MobileMoE: Scaling On-Device Mixture of Experts
For more versions, check out the 🤗 MobileMoE Collection
MobileMoE establishes a new Pareto frontier for on-device LLMs. Average benchmark accuracy, computed over 14 benchmarks spanning commonsense, knowledge, science, comprehension, and reasoning, is plotted against (a) per-token inference compute Finf = 2Nact (GFLOPs) and (b) total parameters Ntotal (B); in (b), x-axis tick labels show total params (B) | projected INT4 memory (GB). Accuracy is shown for the instruction-tuned models.
Key Features
- A new Pareto frontier for on-device LLMs. Across 14 foundational benchmarks, MobileMoE matches or exceeds leading on-device dense LLMs at 2–4× fewer inference FLOPs, and matches or surpasses the state-of-the-art MoE OLMoE-1B-7B with up to 60% fewer parameters.
- Scaling-law-derived architecture. The architecture is derived from an on-device MoE scaling law that jointly optimizes under mobile memory and compute constraints, identifying an on-device sweet spot: moderate sparsity, with fine-grained experts and shared expert.
- Four-stage recipe. Pre-training → mid-training → instruction fine-tuning → INT4 quantization-aware training, all on open-source datasets.
Model Information
Model: MobileMoE-S-SFT (instruction-tuned)
Active Parameters: 272M
Total Parameters: 1.3B
Layers: 20
Model Dimension: 768
Attention Heads: 12
KV Heads: 4 (GQA)
Head Dimension: 64
Routed Experts: 60 (fine-grained, FFN hidden dim 384 each)
Active Experts per Token: 4 (top-k sigmoid routing, with normalization)
Shared Expert: 1, always on (FFN hidden dim 1536)
Vocabulary Size: 128,256
Other Features: QK-Norm, tied input/output embeddings, RoPE (θ = 500,000)
Chat Template: Yes (end-of-turn token <|eot|>)
Input Modality: Text
Output Modality: Text
Languages: English
Training Stages: Pre-training → mid-training → supervised fine-tuning (SFT)
Context Length: 8,192 tokens
Precision: BF16
Model Developer: Meta
Model Release Date: Aug 2026
License: MobileMoE is FAIR NC licensed
Results
All results below are for the instruction-tuned (SFT) models. We re-evaluated every model under identical settings in non-thinking mode with greedy decoding, using lm-eval together with the official allenai/IFBench package. Few-shot counts are shown in parentheses after benchmark names; benchmarks without a count are evaluated 0-shot. The MobileMoE results use the exact weights in this repository, which include brief fine-tuning with self-identity beyond the SFT checkpoint in the technical report, resulting in a small difference: a foundational-benchmark average of 47.2 here versus 46.7 in the report.
Foundational benchmarks
| Capability | Benchmark | Gemma 3 270M | SmolLM2 360M | MobileMoE-S |
|---|---|---|---|---|
| Active / total params | 270M | 362M | 272M / 1.3B | |
| Commonsense Reasoning | HellaSwag | 39.4 | 56.9 | 56.1 |
| PIQA | 67.1 | 71.6 | 74.8 | |
| SIQA | 39.6 | 40.6 | 43.1 | |
| WinoGrande | 53.0 | 57.4 | 59.6 | |
| Knowledge | MMLU (5-shot) | 26.5 | 25.9 | 42.9 |
| NaturalQuestions (5-shot) | 2.8 | 6.4 | 10.9 | |
| TriviaQA (5-shot) | 9.1 | 20.4 | 30.5 | |
| Science | ARC-Challenge (25-shot) | 27.7 | 38.8 | 46.2 |
| ARC-Easy | 50.5 | 49.1 | 73.6 | |
| OpenBookQA | 35.0 | 36.2 | 32.6 | |
| Reading | BoolQ | 56.1 | 42.5 | 72.7 |
| DROP (3-shot) | 11.0 | 15.2 | 33.1 | |
| Reasoning | BIG-Bench Hard (3-shot) | 31.8 | 30.5 | 32.5 |
| GSM8K (8-shot) | 5.8 | 10.0 | 52.4 | |
| Average | 32.5 | 35.8 | 47.2 |
Other capabilities
| Capability | Benchmark | Gemma 3 270M | SmolLM2 360M | MobileMoE-S |
|---|---|---|---|---|
| Math | MATH-500 (4-shot) | 7.2 | 3.8 | 18.8 |
| GSM-Plus (5-shot) | 4.3 | 4.6 | 28.9 | |
| Avg | 5.7 | 4.2 | 23.8 | |
| Code | HumanEval | 12.8 | 0.0 | 46.3 |
| MBPP (3-shot) | 9.8 | 22.8 | 27.4 | |
| Avg | 11.3 | 11.4 | 36.9 | |
| Instruction Following | IFEval | 31.2 | 40.2 | 59.5 |
| IFBench | 11.2 | 19.1 | 14.2 | |
| Avg | 21.2 | 29.7 | 36.8 |
Training
MobileMoE uses a four-stage recipe. This checkpoint is the output of stage 3 (supervised fine-tuning).
MobileMoE four-stage training recipe: pre-training (PT) → mid-training (MT) → instruct supervised fine-tuning (SFT) → quantization-aware training (QAT) with INT4 precision.
| Pre-training | Mid-training | SFT | QAT | |
|---|---|---|---|---|
| Context length | 2,048 | 8,192 | 8,192 | 8,192 |
| Total tokens | ~6T | ~500B | ~126B | ~21B |
| Peak learning rate | 4×10-4 | 4×10-5 | 4×10-6 | 4×10-6 |
| LR schedule | Cosine | Linear | Cosine | Cosine |
| Token dispatch | drop-and-pad | drop-and-pad | dropless | dropless |
How to use
MobileMoE uses a custom architecture (model_type: mobilemoe) that is not yet part of upstream transformers, so trust_remote_code=True is required. The modeling code ships in this repo (configuration_mobilemoe.py, modeling_mobilemoe.py).
Requirements
pip install "torch>=2.1" "transformers>=4.57" "safetensors>=0.4" "accelerate>=1.0"
Verified with the following versions:
| Package | Version |
|---|---|
torch |
2.8.0 (cu128) |
transformers |
4.57.6 |
tokenizers |
0.22.2 |
safetensors |
0.7.0 |
accelerate |
1.13.0 |
For batch evaluation we recommend vLLM (≥ 0.10.2) with enforce_eager=True.
Known issues. Loading the tokenizer on transformers 4.57.6 prints a fix_mistral_regex=True warning. Please ignore it and do not set the flag — MobileMoE uses the Llama-3 128k text vocabulary, whose default tokenization is already correct.
Chat
This instruction-tuned model includes a chat template. Format prompts with apply_chat_template; the template uses <|eot|> to mark the end of each turn.
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
MODEL_ID = "facebook/MobileMoE-S-SFT"
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
model = AutoModelForCausalLM.from_pretrained(
MODEL_ID,
trust_remote_code=True,
dtype=torch.bfloat16,
)
model.to("cuda" if torch.cuda.is_available() else "cpu")
model.eval()
messages = [{"role": "user", "content": "Why are open-source on-device language models great?"}]
input_ids = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_tensors="pt"
).to(model.device)
outputs = model.generate(
input_ids,
attention_mask=torch.ones_like(input_ids),
max_new_tokens=1024,
do_sample=False,
temperature=None,
top_p=None,
pad_token_id=tokenizer.eos_token_id,
)
print(tokenizer.decode(outputs[0][input_ids.shape[-1]:], skip_special_tokens=True))
For multi-turn conversations, append each generated reply to messages with the assistant role. This ensures that each subsequent prompt includes the complete conversation history:
messages = []
for user_message in ["Who are you?", "Why are open-source on-device language models great?"]:
messages.append({"role": "user", "content": user_message})
input_ids = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_tensors="pt"
).to(model.device)
outputs = model.generate(
input_ids,
attention_mask=torch.ones_like(input_ids),
max_new_tokens=1024,
do_sample=False,
temperature=None,
top_p=None,
pad_token_id=tokenizer.eos_token_id,
)
reply = tokenizer.decode(outputs[0][input_ids.shape[-1]:], skip_special_tokens=True).strip()
messages.append({"role": "assistant", "content": reply})
print(reply)
Citation
@article{chen2026mobilemoe,
title={MobileMoE: Scaling On-Device Mixture of Experts},
author={Chen, Yanbei and Huang, Hanxian and Chang, Ernie and Szwejbka, Jacob and Desai, Digant and Liu, Zechun and Chandra, Vikas and Krishnamoorthi, Raghuraman},
journal={arXiv preprint arXiv:2605.27358},
year={2026}
}
License
MobileMoE is distributed under the FAIR Noncommercial Research License.
- Downloads last month
- -


# Gated model: Login with a HF token with gated access permission hf auth login