File size: 8,214 Bytes
f9346af
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
493177f
 
 
f9346af
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
493177f
 
 
 
f9346af
 
 
 
 
 
 
 
 
 
 
 
 
493177f
 
f9346af
 
 
 
 
493177f
f9346af
 
 
 
493177f
f9346af
 
493177f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f9346af
493177f
f9346af
493177f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f9346af
493177f
 
 
 
 
 
 
 
f9346af
493177f
f9346af
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
493177f
 
 
f9346af
 
493177f
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
---
license: apache-2.0
datasets:
- HuggingFaceFW/fineweb-edu
language:
- en
pipeline_tag: text-generation
library_name: transformers
tags:
- lowonmind
- tiny-lm
- pretrained-from-scratch
- scaling-limits
---

# LowOnMind-5M

A decoder-only language model with **4,920,384 parameters**, pretrained from
scratch on **200M tokens** of
`HuggingFaceFW/fineweb-edu` (sample-10BT).

The largest model in the LowOnMind family and the third point on its scaling
curve, after
[LowOnMind-300k](https://huggingface.co/DedeProGames/LowOnMind-300k) and
[LowOnMind-1M](https://huggingface.co/DedeProGames/LowOnMind-1M). All three
share an **identical tokenizer, dataset, token budget (200M) and schedule
shape**, so validation loss, bits-per-character and benchmark results are
directly comparable across the series.

It is also the first model in the family whose benchmark performance is
statistically distinguishable from chance.

## Architecture

| | 300k | 1M | 5M |
|---|---:|---:|---:|
| parameters | 296,960 | 985,152 | **4,920,384** |
| hidden_size | 64 | 96 | 192 |
| intermediate_size | 136 (2.12x) | 256 (2.67x) | 512 (2.667x) |
| num_hidden_layers | 6 | 9 | 12 |
| heads (q / kv) | 4 / 2 | 6 / 2 | 12 / 4 |
| head_dim | 16 | 16 | 16 |
| aspect ratio | 10.7 | 10.7 | **16.0** |
| embedding share | 22.1% | 10.0% | **4.0%** |
| vocab_size | 1024 | 1024 | 1024 (same tokenizer) |
| context | 512 | 512 | 512 |
| tokens seen | 200M | 200M | 200M |
| tokens/param | 673 | 203 | 41 |

Two deviations from the smaller siblings, both deliberate:

- **Aspect ratio rises from 10.7 to 16.0.** This is the
  normal direction when scaling (GPT-2 small sits at 64). Holding 10.7 at this
  budget would require roughly 18 layers of hidden_size=160 with an
  implausibly wide MLP.
- **intermediate/hidden is now exactly 8/3 = 2.667**, the standard SwiGLU ratio
  used by Llama. LowOnMind-300k was at 2.12 and LowOnMind-1M at 2.67.

The **vocabulary was deliberately left at 1024** rather than raised to something
more appropriate for this scale. A larger vocabulary would compress better
(1024-token byte-level BPE runs about 2.35 characters per token, so 200M tokens
is only ~470MB of text) and would almost certainly improve absolute results.
Keeping it fixed is what makes the three-model comparison valid — the cost is
that this model spends capacity assembling words from fragments that a
4096-token vocabulary would hand it for free.

Modelling code is otherwise byte-identical to the two smaller siblings: GQA,
SwiGLU, RMSNorm, tied embeddings, **QK-Norm** per head, **precomputed RoPE**
with automatic re-expansion, residual projections initialized at
`std / sqrt(2 * num_layers)`.

## Training

| | |
|---|---|
| data | `HuggingFaceFW/fineweb-edu`, sample-10BT |
| tokens | 200M (6,103 steps x 32,768) |
| sequence length | 512 |
| batch size | 64 |
| optimizer | AdamW, betas (0.9, 0.95), wd 0.1 |
| lr | 1.2e-03 peak, cosine to 1.2e-04, 250 warmup |
| grad clip | 1.0 |
| precision | float16 + GradScaler |
| hardware | Tesla T4 |
| wall clock | 27 min |

At 41 tokens per parameter this run is the closest of the three to the
Chinchilla-optimal ratio of roughly 20 — about 2x above it, against 10x for
LowOnMind-1M and 34x for LowOnMind-300k. Train and validation loss tracked
each other throughout; no overfitting.

## Results

| metric | 300k | 1M | 5M |
|---|---:|---:|---:|
| validation loss | 3.2982 | 2.9908 | **2.5828** |
| validation perplexity | 27.06 | 19.90 | **13.23** |
| bits per character | 2.030 | 1.836 | **1.586** |

Perplexity is not comparable across tokenizers, but it is comparable across
these three models because they share one. Bits per character
(loss / ln 2 / 2.35 chars-per-token) is the portable figure.

Deltas: -0.4080 nats from LowOnMind-1M (5.0x the parameters),
-0.7154 nats from LowOnMind-300k (16.6x).

### Real-word rate

With a 1024-token byte-level vocabulary, no long word exists as a single token —
the model has to assemble every one of them from fragments. The fraction of
emitted words that are real English words was introduced to measure this.

| | rate |
|---|---:|
| LowOnMind-1M | 98.0% |
| LowOnMind-5M | 96.3% |
| FineWeb-Edu itself (same lexicon) | 98.4% |

Measured over 64 unconditional samples (5,398 words), using the same reference
lexicon as LowOnMind-1M: words appearing at least 5 times in a 20k-document
sample of the training corpus.

**This number went down, and it should not be read as degraded spelling.** The
drop is statistically real (z = 5.38, not sampling noise), but inspecting the
non-words shows what happened: `illuminator` is an ordinary English word,
`phillipsburg` is a US town, `shima` is a common element of Japanese place
names. They are counted as errors only because they fall below the reference
lexicon's frequency-5 threshold. The remainder (`hymenola`, `almanine`,
`perleti`, `amiravicis`) skew toward proper-noun and Latinate-technical
morphology rather than the malformed common words the metric was built to catch
— LowOnMind-300k produced things like `landship` and `parsetic`, failures of a
different kind.

**The metric has a floor problem as well as a ceiling problem.** As a model
improves it emits rarer real vocabulary — names, places, technical terms — which
a frequency-thresholded lexicon scores as wrong. So the measured rate can fall
while actual quality rises. Comparing against a full dictionary with proper-noun
handling, rather than a corpus-frequency cutoff, would be the fix. The 96.3%
figure is reported as-measured for continuity, but it should not be used to rank
these models.

## BananaMind Base Bench 1.1

Evaluated on [BananaMind/BananaMind-Base-Bench-1.1](https://huggingface.co/datasets/BananaMind/BananaMind-Base-Bench-1.1),
the same 350-item English continuation-likelihood benchmark used across the
family, with identical scoring: context and each of the four continuations
tokenized separately with `add_special_tokens=False`, no BOS, selection by
highest mean conditional token log-probability.

Run validity: dataset SHA-256 matched, full schema validation passed, no context
required truncation against the 512-token window.

| Category | 300k | 1M | 5M | z vs chance (5M) | Elo (5M) |
|---|---:|---:|---:|---:|---:|
| language_completion | 46.0% | 52.0% | **62.0%** | **+6.04** | **1008** |
| world_knowledge | 22.0% | 22.0% | **38.0%** | +2.12 | 881 |
| context_tracking | 14.0% | 24.0% | 32.0% | +1.14 | 851 |
| quantitative | 32.0% | 28.0% | 28.0% | +0.49 | 872 |
| logical_reasoning | 24.0% | 28.0% | 26.0% | +0.16 | 900 |
| commonsense | 34.0% | 28.0% | 24.0% | -0.16 | 758 |
| code_completion | 14.0% | 20.0% | 16.0% | -1.47 | 805 |

| | 300k | 1M | 5M |
|---|---:|---:|---:|
| Overall Elo | 833 | 843 | **863** |
| Chance-level Elo (this grid) | 805 | 805 | 805 |
| Raw accuracy | 26.6% | 28.9% | **32.3%** |
| 95% CI | [22.0, 31.2] | [24.2, 33.6] | **[27.4, 37.2]** |
| z vs. chance | +0.69 | +1.68 | **+3.15** |
| significant vs. chance | no | no | **yes** |

Difficulty split: easy 30.8%, medium 33.3%, hard 32.8%.

## Usage

```python
from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("DedeProGames/LowOnMind-5M")
model = AutoModelForCausalLM.from_pretrained("DedeProGames/LowOnMind-5M", trust_remote_code=True)

ids = tok("The ", return_tensors="pt").input_ids
print(tok.decode(model.generate(ids, max_new_tokens=64, use_cache=False)[0]))
```

`trust_remote_code=True` is required — the architecture ships as custom modeling
code in the repository. `use_cache=False` is required: this implementation has no
KV cache and recomputes the full window at each generation step.

## Limitations

At ~5M parameters this is still a research artifact, not a usable model. Expect
fluent local syntax and register-appropriate structure, but **no reliable
coherence across a paragraph**, no dependable factual knowledge, and no ability
to track state across a passage. Benchmark accuracy of 32.3% is above chance and
far below usefulness. The 1024-token vocabulary caps absolute quality below what
this parameter count could otherwise reach.

The 512-token context and absent KV cache also make it unsuitable for any real
workload.