File size: 3,244 Bytes
c8448d2
c6c8db8
 
 
 
 
 
 
 
 
 
 
c8448d2
c6c8db8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
---
license: apache-2.0
datasets:
  - HuggingFaceFW/fineweb
language:
  - en
library_name: pytorch
tags:
  - looped-transformer
  - weight-tying
  - recurrent-depth
  - small-language-model
---

# SimLoop 1+loop x1+1 (10.0M)

A **looped** (weight-tied recurrent) Qwen3-style transformer trained from
scratch on FineWeb under a hard budget of **<= 10M parameters** and
**<= 100M training tokens**.

One middle block is applied **K = 1** times; the layers around it are
ordinary unlooped layers:

```
embed -> [1 layer] -> ( 1 looped layer ) x K -> [1 layer] -> RMSNorm -> tied head
```

| | |
|---|---|
| parameters | **9,962,496** (incl. embeddings, input/output tied) |
| d_model / d_mlp | 384 / 720 |
| heads (GQA) | 6 query / 2 key-value, head_dim 64 |
| context | 512 tokens |
| vocabulary | 16,384 byte-level BPE trained on FineWeb |
| training tokens | 60,014,592 |
| primitives | RMSNorm, RoPE, SwiGLU, GQA, QK-norm |

## Results

| k | CE (nats) | perplexity | bits-per-byte |
|---|---|---|---|
| 0 | 5.4088 | 223.35 | 1.8913 |
| 1 | 4.3966 | 81.17 | 1.5374 **<- best** |
| 2 | 4.5995 | 99.43 | 1.6083 |
| 3 | 4.9413 | 139.95 | 1.7278 |
| 4 | 5.2659 | 193.62 | 1.8413 |
| 5 | 5.5476 | 256.63 | 1.9399 |
| 6 | 5.7900 | 327.01 | 2.0246 |
| 7 | 6.0000 | 403.41 | 2.0980 |
| 8 | 6.1835 | 484.69 | 2.1622 |
| 9 | 6.3452 | 569.77 | 2.2188 |
| 10 | 6.4887 | 657.65 | 2.2689 |
| 11 | 6.6166 | 747.41 | 2.3137 |
| 12 | 6.7313 | 838.20 | 2.3537 |

Reference points on the same validation split: a context-free **unigram** model scores CE 7.5476 (ppl 1896.10); **uniform** over the vocabulary scores CE 9.7041.


`k` is the number of applications of the looped block at inference. The model
is weight-tied, so **any k can be run**; the table is a single checkpoint
evaluated at every depth.

**Measured behaviour:** TRAINED UNLOOPED (K=1): the k>1 rows show how a model trained at one application behaves when looped anyway, not whether looping pays.

- first application of the looped block buys **+1.0122** nats (k=0 -> k=1)
- all further applications buy **+0.0000** nats (k=1 -> k=1)


## Usage

```python
import torch
from simloop.stack import StackConfig, StackedLoop

ck = torch.load("model.pt", map_location="cpu", weights_only=False)
model = StackedLoop(StackConfig(**ck["model_cfg"]))
model.load_state_dict(ck["model"]); model.eval()

ids = torch.tensor([[1, 2, 3]])            # from tokenizer.json
logits = model(ids, K=1)           # try other K: the block is tied
```

See `load_example.py`. The tokenizer is a `tokenizers` BPE:
`Tokenizer.from_file("tokenizer.json")`.

## Honest limitations

- Trained on 60,014,592 tokens at ~10M parameters. It is a research artifact
  for studying looped depth, **not** a useful general-purpose language model.
- Perplexity is tokenizer-dependent; **bits-per-byte** is the comparable
  number and is reported above.
- The looped block **saturates**: past the depth listed as best above, extra
  applications make cross-entropy worse, not better. This is measured, not
  assumed, and is the central finding of the project.
- English-only, no instruction tuning, no safety filtering beyond FineWeb's.

Full experimental record, including every failed experiment:
https://github.com/brkdrd/SimLoop