Instructions to use ShourenWSR/how-to-loop-moe with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ShourenWSR/how-to-loop-moe with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ShourenWSR/how-to-loop-moe")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("ShourenWSR/how-to-loop-moe", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ShourenWSR/how-to-loop-moe with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ShourenWSR/how-to-loop-moe" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ShourenWSR/how-to-loop-moe", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/ShourenWSR/how-to-loop-moe
- SGLang
How to use ShourenWSR/how-to-loop-moe with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ShourenWSR/how-to-loop-moe" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ShourenWSR/how-to-loop-moe", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ShourenWSR/how-to-loop-moe" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ShourenWSR/how-to-loop-moe", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use ShourenWSR/how-to-loop-moe with Docker Model Runner:
docker model run hf.co/ShourenWSR/how-to-loop-moe
How to Loop MoE: Flatten the Experts, Untie the Attention
Weights of the four 100B-token models from the paper How to Loop MoE: Flatten the Experts, Untie the Attention.
Each model is in its own folder:
| Folder | Looped block | Experts per looped layer |
|---|---|---|
Foil-1 |
1 layer, looped 16 times | 64 |
Foil-2 |
2 layers, looped 8 times | 32 |
Foil-3 |
4 layers, looped 4 times | 16 |
Base |
8 layers, looped 2 times | 8 |
All four models have:
- 563.2M parameters;
- one MoE layer before the loop and one after it, each with 8 experts;
- top-2 routing, hidden size 1024, 16 attention heads;
- separate attention weights for every pass through the loop (untied attention);
- 18 layers in total once the loop is unrolled;
- context length 4096;
- the SmolLM2 tokenizer (vocabulary 49,152).
They were trained on 100B tokens of FineWeb-Edu (sample-100BT).
Usage
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo, model_name = "ShourenWSR/how-to-loop-moe", "Foil-1" # or Foil-2, Foil-3, Base
tokenizer = AutoTokenizer.from_pretrained(repo, subfolder=model_name, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
repo, subfolder=model_name, trust_remote_code=True, dtype=torch.bfloat16
).eval()
ids = tokenizer("The capital of France is", return_tensors="pt").input_ids
print(tokenizer.decode(model.generate(ids, max_new_tokens=20, do_sample=False)[0]))
trust_remote_code=True is required: the architecture is defined in
modeling_loop_lm.py. The same file is at the repository root and in each
folder; with subfolder=, transformers loads the code from the root.
Labels are not shifted inside the model. When you pass labels to
forward, they must already be the next-token targets:
labels[..., t] == input_ids[..., t + 1], with the last position set to -100.
Passing labels=input_ids, as with most Hugging Face causal language models,
gives a meaningless loss. The returned loss also includes the router auxiliary
losses. To get the language-modelling loss alone, compute it from logits:
logits = model(input_ids=ids).logits.float()
loss = torch.nn.functional.cross_entropy(logits[0, :-1], ids[0, 1:])
No padding masks. The model does not use attention_mask, and it raises
an error if the mask contains padding. Run one sequence at a time, or pad
batches on the right and omit attention_mask; right padding does not change
the outputs for the real tokens. Batched generation with left padding is not
supported.
Intermediate checkpoints
Intermediate checkpoints of each run are in the per-run repositories: Foil-1, Foil-2, Foil-3, Base. The weights here are the final step of those runs (step 254,313).
License
Apache-2.0. The modeling code is ported from the release accompanying
Sparse Layers are Critical to Scaling Looped Language Models
(arXiv:2605.09165); see NOTICE.