Papers
arxiv:2609.35751

How to Loop MoE: Flatten the Experts, Untie the Attention

Published on Sep 28
ยท Submitted by
Debargha Ganguly
on Oct 6
Authors:
,
,
,
,
,
,
,
,

Abstract

Looped Transformers reuse one block of layers several times: by spending extra computation they push a model of fixed size further, and so use its parameters more fully; while sparse mixture-of-experts (MoE) models activate only a few of many experts for each token. Looped MoE bridges these two design philosophies and gives MoE models new potential for better expert usage, but it raises a question: how to loop a MoE? We answer it with Foil. With the expert parameters and the expert compute per token held fixed, Foil (1) flattens the experts, halving the expert layers, doubling the experts per layer and doubling the passes, so that every routing decision chooses from a larger pool, and (2) unties the attention, giving each pass its own attention parameters while the experts and routers stay shared. Experiments show that Foil clearly outperforms the unflattened looped baseline: at 20B tokens every Foil model has lower pretraining loss than the baseline; at 100B tokens the loss improves monotonically with the degree of flattening, the most flattened Foil ending 0.012 nat below the baseline at equal parameters and compute, with downstream accuracy on par or better; untying the attention also yields more balanced and more confident routing at equal shape. Our ablations analyse why Foil works and turn the findings into design guidance for looped MoE: the returns of looping and of widening the expert layers amplify each other, routing confidence tracks healthy expert use better than load balance, and a sparse looped MoE should therefore use more experts per layer and more passes. Code and configurations are available at https://github.com/SR-A-W/how-to-loop-moe.

Community

Paper submitter

Looped Transformers reuse the same block across multiple passes, while sparse Mixture-of-Experts models store many experts but activate only a few for each token. These two ideas are naturally complementary: every new pass gives a token another routing decision, allowing it to reach different experts and expert combinations without adding expert parameters. But this raises an open design question: under a fixed parameter and compute budget, how should experts be distributed across layers and passes; and which components should be shared across passes?

Our new preprint, โ€œHow to Loop MoE: Flatten the Experts, Untie the Attention,โ€ answers this question with Foil.

๐—ข๐˜‚๐—ฟ ๐—ถ๐—ฑ๐—ฒ๐—ฎ: Flatten the experts. Loop more. Untie the attention.

At each flattening step, Foil uses:
โ†’ half as many expert layers
โ†’ twice as many experts per layer
โ†’ twice as many recurrent passes
โ†’ separate attention parameters for each pass
while keeping total parameters, expert compute per token, and effective depth fixed.

We go from:
8 experts ร— 8 layers ร— 2 passes to:
64 experts ร— 1 layer ร— 16 passes

๐—ฅ๐—ฒ๐˜€๐˜‚๐—น๐˜๐˜€:
โ†’ Every Foil configuration achieves lower pretraining loss than the baseline at 20B tokens.
โ†’ ๐—”๐˜ ๐Ÿญ๐Ÿฌ๐Ÿฌ๐—• ๐˜๐—ผ๐—ธ๐—ฒ๐—ป๐˜€, ๐—น๐—ผ๐˜€๐˜€ ๐—ถ๐—บ๐—ฝ๐—ฟ๐—ผ๐˜ƒ๐—ฒ๐˜€ ๐—บ๐—ผ๐—ป๐—ผ๐˜๐—ผ๐—ป๐—ถ๐—ฐ๐—ฎ๐—น๐—น๐˜† ๐˜„๐—ถ๐˜๐—ต ๐˜๐—ต๐—ฒ ๐—ฑ๐—ฒ๐—ด๐—ฟ๐—ฒ๐—ฒ ๐—ผ๐—ณ ๐—ณ๐—น๐—ฎ๐˜๐˜๐—ฒ๐—ป๐—ถ๐—ป๐—ด ๐—ฎ๐˜ ๐—บ๐—ฎ๐˜๐—ฐ๐—ต๐—ฒ๐—ฑ ๐—ฝ๐—ฎ๐—ฟ๐—ฎ๐—บ๐—ฒ๐˜๐—ฒ๐—ฟ๐˜€ ๐—ฎ๐—ป๐—ฑ ๐—ฐ๐—ผ๐—บ๐—ฝ๐˜‚๐˜๐—ฒ. The fully flattened Foil ends 0.012 nat below the baseline, with downstream accuracy on par or better.
โ†’ Untying attention becomes increasingly valuable as the model gets flatter. At the fully flattened shape, it lowers loss by 0.049 nat and improves mean accuracy on three representative downstream tasks by 3.3 points over tied attention at the same shape, with no additional compute.

โ†’ More experts per layer make additional loops more useful; and more loops make additional experts more useful. The two amplify each other.
โ†’ Load balance alone is not enough; it should be read together with routing confidence. In our experiments, routing confidence moves with looping gain, and its per-pass peak may signal diminishing returns from further looping or flattening.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.35751
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.35751 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.35751 in a Space README.md to link it from this page.

Collections including this paper 1