MobileMoE-M-SFT / README.md
yanbeic's picture
Link IFBench to its upstream repository
31d32e6 verified
|
Raw
History Blame Contribute Delete
11.9 kB
metadata
license: fair-noncommercial-research-license
extra_gated_fields:
  First Name: text
  Last Name: text
  Date of birth: date_picker
  Country: country
  Affiliation: text
  Job title:
    type: select
    options:
      - Student
      - Research Graduate
      - AI researcher
      - AI developer/engineer
      - Reporter
      - Other
  geo: ip_location
  By clicking Submit below I accept the terms of the license and acknowledge that the information I provide will be collected stored processed and shared in accordance with the Meta Privacy Policy: checkbox
extra_gated_description: >-
  The information you provide will be collected, stored, processed and shared in
  accordance with the [Meta Privacy
  Policy](https://www.facebook.com/privacy/policy/).
extra_gated_button_content: Submit
language:
  - en
library_name: transformers
tags:
  - facebook
  - meta
  - pytorch
  - mixture-of-experts
  - MoE
  - on-device

MobileMoE-M (SFT) Model Card

MobileMoE is a family of on-device Mixture-of-Experts (MoE) language models with sub-billion active parameters, designed to push the quality–efficiency Pareto frontier for on-device LLMs, including three model scales (S/M/L): 0.3B/0.5B/0.9B active parameters (1.3B/2.8B/5.3B total), with <3 GB INT4 weight footprints to fit in mobile DRAM. Each scale is released in three variants: a Base model (pre-training + mid-training), an SFT model (supervised fine-tuning), and a QAT model (quantization-aware training). You are currently in the MobileMoE-M-SFT repository — the instruction-tuned 0.5B-active model.

S M L
Active / total params 272M / 1.3B 528M / 2.8B 922M / 5.3B
Layers 20 26 32
Model dimension 768 1024 1280
Heads (Q / KV) 12 / 4 16 / 4 20 / 4
Routed experts 60 60 60
Top-k 4 4 4
INT4 weight memory 0.68 GB 1.48 GB 2.75 GB
Base MobileMoE-S-Base MobileMoE-M-Base MobileMoE-L-Base
SFT MobileMoE-S-SFT MobileMoE-M-SFT MobileMoE-L-SFT
QAT (INT4) MobileMoE-S-QAT MobileMoE-M-QAT MobileMoE-L-QAT

For the detailed technical report: 📝 MobileMoE: Scaling On-Device Mixture of Experts

For more versions, check out the 🤗 MobileMoE Collection

MobileMoE establishes a new Pareto frontier for on-device LLMs

MobileMoE establishes a new Pareto frontier for on-device LLMs. Average benchmark accuracy, computed over 14 benchmarks spanning commonsense, knowledge, science, comprehension, and reasoning, is plotted against (a) per-token inference compute Finf = 2Nact (GFLOPs) and (b) total parameters Ntotal (B); in (b), x-axis tick labels show total params (B) | projected INT4 memory (GB). Accuracy is shown for the instruction-tuned models.

Key Features

  • A new Pareto frontier for on-device LLMs. Across 14 foundational benchmarks, MobileMoE matches or exceeds leading on-device dense LLMs at 2–4× fewer inference FLOPs, and matches or surpasses the state-of-the-art MoE OLMoE-1B-7B with up to 60% fewer parameters.
  • Scaling-law-derived architecture. The architecture is derived from an on-device MoE scaling law that jointly optimizes under mobile memory and compute constraints, identifying an on-device sweet spot: moderate sparsity, with fine-grained experts and shared expert.
  • Four-stage recipe. Pre-training → mid-training → instruction fine-tuning → INT4 quantization-aware training, all on open-source datasets.

Model Information

Model: MobileMoE-M-SFT (instruction-tuned)
Active Parameters: 528M
Total Parameters: 2.8B
Layers: 26
Model Dimension: 1024
Attention Heads: 16
KV Heads: 4 (GQA)
Head Dimension: 64
Routed Experts: 60 (fine-grained, FFN hidden dim 512 each)
Active Experts per Token: 4 (top-k sigmoid routing, with normalization)
Shared Expert: 1, always on (FFN hidden dim 2048)
Vocabulary Size: 128,256
Other Features: QK-Norm, tied input/output embeddings, RoPE (θ = 500,000)
Chat Template: Yes (end-of-turn token <|eot|>)
Input Modality: Text
Output Modality: Text
Languages: English
Training Stages: Pre-training → mid-training → supervised fine-tuning (SFT)
Context Length: 8,192 tokens
Precision: BF16
Model Developer: Meta
Model Release Date: Aug 2026
License: MobileMoE is FAIR NC licensed

Results

All results below are for the instruction-tuned (SFT) models. We re-evaluated every model under identical settings in non-thinking mode with greedy decoding, using lm-eval together with the official allenai/IFBench package. Few-shot counts are shown in parentheses after benchmark names; benchmarks without a count are evaluated 0-shot. The MobileMoE results use the exact weights in this repository, which include brief fine-tuning with self-identity beyond the SFT checkpoint in the technical report, resulting in a small difference: a foundational-benchmark average of 55.9 here versus 55.3 in the report.

Foundational benchmarks

Capability Benchmark Gemma 3 270M SmolLM2 360M MobileMoE-S Qwen3.5 0.8B MobileMoE-M
Active / total params 270M 362M 272M / 1.3B 749M 528M / 2.8B
Commonsense Reasoning HellaSwag 39.4 56.9 56.1 49.7 66.7
PIQA 67.1 71.6 74.8 69.4 77.7
SIQA 39.6 40.6 43.1 38.8 49.4
WinoGrande 53.0 57.4 59.6 57.6 62.7
Knowledge MMLU (5-shot) 26.5 25.9 42.9 50.2 54.5
NaturalQuestions (5-shot) 2.8 6.4 10.9 3.3 17.7
TriviaQA (5-shot) 9.1 20.4 30.5 16.1 46.7
Science ARC-Challenge (25-shot) 27.7 38.8 46.2 41.8 52.5
ARC-Easy 50.5 49.1 73.6 61.4 80.1
OpenBookQA 35.0 36.2 32.6 30.8 39.8
Reading BoolQ 56.1 42.5 72.7 62.5 77.7
DROP (3-shot) 11.0 15.2 33.1 33.3 49.9
Reasoning BIG-Bench Hard (3-shot) 31.8 30.5 32.5 37.8 38.8
GSM8K (8-shot) 5.8 10.0 52.4 45.7 67.6
Average 32.5 35.8 47.2 42.7 55.9

Other capabilities

Capability Benchmark Gemma 3 270M SmolLM2 360M MobileMoE-S Qwen3.5 0.8B MobileMoE-M
Math MATH-500 (4-shot) 7.2 3.8 18.8 19.6 26.8
GSM-Plus (5-shot) 4.3 4.6 28.9 26.5 42.6
Avg 5.7 4.2 23.8 23.1 34.7
Code HumanEval 12.8 0.0 46.3 31.1 59.8
MBPP (3-shot) 9.8 22.8 27.4 25.4 44.4
Avg 11.3 11.4 36.9 28.3 52.1
Instruction Following IFEval 31.2 40.2 59.5 59.9 61.6
IFBench 11.2 19.1 14.2 19.8 20.8
Avg 21.2 29.7 36.8 39.8 41.2

Training

MobileMoE uses a four-stage recipe. This checkpoint is the output of stage 3 (supervised fine-tuning).

MobileMoE four-stage training recipe

MobileMoE four-stage training recipe: pre-training (PT) → mid-training (MT) → instruct supervised fine-tuning (SFT) → quantization-aware training (QAT) with INT4 precision.

Pre-training Mid-training SFT QAT
Context length 2,048 8,192 8,192 8,192
Total tokens ~6T ~500B ~126B ~21B
Peak learning rate 4×10-4 4×10-5 4×10-6 4×10-6
LR schedule Cosine Linear Cosine Cosine
Token dispatch drop-and-pad drop-and-pad dropless dropless

How to use

MobileMoE uses a custom architecture (model_type: mobilemoe) that is not yet part of upstream transformers, so trust_remote_code=True is required. The modeling code ships in this repo (configuration_mobilemoe.py, modeling_mobilemoe.py).

Requirements

pip install "torch>=2.1" "transformers>=4.57" "safetensors>=0.4" "accelerate>=1.0"

Verified with the following versions:

Package Version
torch 2.8.0 (cu128)
transformers 4.57.6
tokenizers 0.22.2
safetensors 0.7.0
accelerate 1.13.0

For batch evaluation we recommend vLLM (≥ 0.10.2) with enforce_eager=True.

Known issues. Loading the tokenizer on transformers 4.57.6 prints a fix_mistral_regex=True warning. Please ignore it and do not set the flag — MobileMoE uses the Llama-3 128k text vocabulary, whose default tokenization is already correct.

Chat

This instruction-tuned model includes a chat template. Format prompts with apply_chat_template; the template uses <|eot|> to mark the end of each turn.

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

MODEL_ID = "facebook/MobileMoE-M-SFT"

tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
model = AutoModelForCausalLM.from_pretrained(
    MODEL_ID,
    trust_remote_code=True,
    dtype=torch.bfloat16,
)
model.to("cuda" if torch.cuda.is_available() else "cpu")
model.eval()

messages = [{"role": "user", "content": "Why are open-source on-device language models great?"}]
input_ids = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, return_tensors="pt"
).to(model.device)

outputs = model.generate(
    input_ids,
    attention_mask=torch.ones_like(input_ids),
    max_new_tokens=1024,
    do_sample=False,
    temperature=None,
    top_p=None,
    pad_token_id=tokenizer.eos_token_id,
)
print(tokenizer.decode(outputs[0][input_ids.shape[-1]:], skip_special_tokens=True))

For multi-turn conversations, append each generated reply to messages with the assistant role. This ensures that each subsequent prompt includes the complete conversation history:

messages = []

for user_message in ["Who are you?", "Why are open-source on-device language models great?"]:
    messages.append({"role": "user", "content": user_message})
    input_ids = tokenizer.apply_chat_template(
        messages, add_generation_prompt=True, return_tensors="pt"
    ).to(model.device)
    outputs = model.generate(
        input_ids,
        attention_mask=torch.ones_like(input_ids),
        max_new_tokens=1024,
        do_sample=False,
        temperature=None,
        top_p=None,
        pad_token_id=tokenizer.eos_token_id,
    )
    reply = tokenizer.decode(outputs[0][input_ids.shape[-1]:], skip_special_tokens=True).strip()
    messages.append({"role": "assistant", "content": reply})
    print(reply)

Citation

@article{chen2026mobilemoe,
  title={MobileMoE: Scaling On-Device Mixture of Experts},
  author={Chen, Yanbei and Huang, Hanxian and Chang, Ernie and Szwejbka, Jacob and Desai, Digant and Liu, Zechun and Chandra, Vikas and Krishnamoorthi, Raghuraman},
  journal={arXiv preprint arXiv:2605.27358},
  year={2026}
}

License

MobileMoE is distributed under the FAIR Noncommercial Research License.