File size: 3,244 Bytes
c8448d2 c6c8db8 c8448d2 c6c8db8 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 | ---
license: apache-2.0
datasets:
- HuggingFaceFW/fineweb
language:
- en
library_name: pytorch
tags:
- looped-transformer
- weight-tying
- recurrent-depth
- small-language-model
---
# SimLoop 1+loop x1+1 (10.0M)
A **looped** (weight-tied recurrent) Qwen3-style transformer trained from
scratch on FineWeb under a hard budget of **<= 10M parameters** and
**<= 100M training tokens**.
One middle block is applied **K = 1** times; the layers around it are
ordinary unlooped layers:
```
embed -> [1 layer] -> ( 1 looped layer ) x K -> [1 layer] -> RMSNorm -> tied head
```
| | |
|---|---|
| parameters | **9,962,496** (incl. embeddings, input/output tied) |
| d_model / d_mlp | 384 / 720 |
| heads (GQA) | 6 query / 2 key-value, head_dim 64 |
| context | 512 tokens |
| vocabulary | 16,384 byte-level BPE trained on FineWeb |
| training tokens | 60,014,592 |
| primitives | RMSNorm, RoPE, SwiGLU, GQA, QK-norm |
## Results
| k | CE (nats) | perplexity | bits-per-byte |
|---|---|---|---|
| 0 | 5.4088 | 223.35 | 1.8913 |
| 1 | 4.3966 | 81.17 | 1.5374 **<- best** |
| 2 | 4.5995 | 99.43 | 1.6083 |
| 3 | 4.9413 | 139.95 | 1.7278 |
| 4 | 5.2659 | 193.62 | 1.8413 |
| 5 | 5.5476 | 256.63 | 1.9399 |
| 6 | 5.7900 | 327.01 | 2.0246 |
| 7 | 6.0000 | 403.41 | 2.0980 |
| 8 | 6.1835 | 484.69 | 2.1622 |
| 9 | 6.3452 | 569.77 | 2.2188 |
| 10 | 6.4887 | 657.65 | 2.2689 |
| 11 | 6.6166 | 747.41 | 2.3137 |
| 12 | 6.7313 | 838.20 | 2.3537 |
Reference points on the same validation split: a context-free **unigram** model scores CE 7.5476 (ppl 1896.10); **uniform** over the vocabulary scores CE 9.7041.
`k` is the number of applications of the looped block at inference. The model
is weight-tied, so **any k can be run**; the table is a single checkpoint
evaluated at every depth.
**Measured behaviour:** TRAINED UNLOOPED (K=1): the k>1 rows show how a model trained at one application behaves when looped anyway, not whether looping pays.
- first application of the looped block buys **+1.0122** nats (k=0 -> k=1)
- all further applications buy **+0.0000** nats (k=1 -> k=1)
## Usage
```python
import torch
from simloop.stack import StackConfig, StackedLoop
ck = torch.load("model.pt", map_location="cpu", weights_only=False)
model = StackedLoop(StackConfig(**ck["model_cfg"]))
model.load_state_dict(ck["model"]); model.eval()
ids = torch.tensor([[1, 2, 3]]) # from tokenizer.json
logits = model(ids, K=1) # try other K: the block is tied
```
See `load_example.py`. The tokenizer is a `tokenizers` BPE:
`Tokenizer.from_file("tokenizer.json")`.
## Honest limitations
- Trained on 60,014,592 tokens at ~10M parameters. It is a research artifact
for studying looped depth, **not** a useful general-purpose language model.
- Perplexity is tokenizer-dependent; **bits-per-byte** is the comparable
number and is reported above.
- The looped block **saturates**: past the depth listed as best above, extra
applications make cross-entropy worse, not better. This is measured, not
assumed, and is the central finding of the project.
- English-only, no instruction tuning, no safety filtering beyond FineWeb's.
Full experimental record, including every failed experiment:
https://github.com/brkdrd/SimLoop
|