AbstractPhil commited on
Commit
efe3aa0
·
verified ·
1 Parent(s): acf94d3

mini-beatrix-3: the stage arms (arms off and arms on in one package)

Browse files
README.md CHANGED
@@ -23,10 +23,12 @@ completed 2026-10-05.
23
 
24
  The package ships **both states in one repository**: the bare model (arms
25
  off) and the model with its **stage arms** mounted (arms on). An arm is a
26
- small detachable adapter trained on this exact frozen core. The arms are a
27
- **group**: they were trained switched on together, and they are mounted
28
- together. Each stage also has a **solo arm**, trained alone, for use one at
29
- a time. Arms are guests on the core, never a change to it: the core's
 
 
30
  weights are the same files either way, and detaching restores the bare
31
  model bit for bit.
32
 
@@ -39,7 +41,7 @@ m = AutoModelForCausalLM.from_pretrained(
39
 
40
  m.say("Who are you?") # arms off: the bare core, its own chat frame
41
 
42
- m.mount_arm("stages-1-4") # arms on: the stage arms, all on together
43
  m.detach_arm() # the bare core again, verified bit-exact
44
  ```
45
 
@@ -48,7 +50,7 @@ Arms on from the first call:
48
  ```python
49
  m = AutoModelForCausalLM.from_pretrained(
50
  "AbstractPhil/mini-beatrix-3", trust_remote_code=True,
51
- default_arm="stages-1-4").eval()
52
  ```
53
 
54
  The model reads **raw UTF-8 bytes**: `input_ids` are byte values 0–255.
@@ -80,9 +82,9 @@ There is no tokenizer to download.
80
  The curriculum taught the model nine kinds of text, one stage after another,
81
  and the finished model moved on from each stage's text as the later stages
82
  came. The stage arms bring that text back without touching the core. They
83
- were fitted on the finished, frozen core **as a group**: the arms of stages 1 and 2 were trained first and then held fixed, and the arms of stages 3 and 4 were attached over them, one after the other, and trained in their presence, so the group works with every arm on at once.
84
- `m.arms()` returns the tables below as data, with each row's recipe and full
85
- measurements.
86
 
87
  **The group `stages-1-4`** (4 arms, 54.8M parameters; two seeds):
88
 
@@ -108,6 +110,38 @@ Each member mounted alone, and the whole group, on every stage's held-out text (
108
  | only `stages-1-4/s4_arith` | 0.553 | 0.721 | 0.928 | 1.127 | 0.951 |
109
  | the whole group `stages-1-4` | 0.048 | 0.112 | 0.224 | 0.338 | 0.952 |
110
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
111
  **How to read the tables.** *Stage text* is the loss, in bits per byte, on
112
  held-out text of the arm's own stage, with arms off and with the whole group
113
  on. *This arm alone* is the same loss with only that one arm mounted. *Stage
@@ -116,26 +150,26 @@ correctly, first in the form the stage text uses and then in a form it does
116
  not. The *web text* figure is the change on held-out ordinary web text
117
  (fineweb-edu) with the group on: the arms were trained to leave it alone.
118
 
119
- **The group is the unit.** The group does not divide its work one stage to an arm. The first arm is a generalist: mounted alone it reads every stage's text well below the bare model, and the second holds a part of stages 2 to 4 (second table). The arms of stages 3 and 4 were trained on top of that fixed pair and add what the pair leaves: with only the first two arms on, stage 3 text reads 0.32 bpb and stage 4 text 0.48, and with the whole group they read 0.22 and 0.34. Mounted alone, the last two arms do little. So mount the group as a whole: that is the state it was trained and measured in. For one stage by itself, use its solo arm below. The first two arms also had more training than the last two (3,200 and 2,400 steps against 800), and no arm here is trained to its limit: at 800 steps a solo arm's stage text is still falling.
120
 
121
- **What the arms do, plainly.** The group recovers each stage's own kind of
122
- text. The gain is in each stage's own forms. Whether it carries over to new
123
- forms is shown in the first table's last column, where it was measured.
124
 
125
  ### Arms off and arms on
126
 
127
  ```python
128
  m.arm # None: arms off
129
- h = m.mount_arm("stages-1-4") # arms on: the whole group
130
  with h.all_off(): # every member masked: the bare core's logits
131
  m.say("Hello there.")
132
  m.detach_arm(verify=True) # raises if the restored core is not bit-exact
133
  m.mount_arm("stages-1-4/s1_perspective") # one member alone (the "alone" readings)
134
  ```
135
 
136
- The group mounts its members in the order they were trained, always on, one
137
  after another at every block, with no mixer between them. `default_arm` in
138
- `from_pretrained` (or in `config.json`) mounts the group as the model loads.
139
  `save_pretrained` refuses while an arm is mounted, so an arm can never be
140
  written into the core's weight file.
141
 
@@ -151,6 +185,10 @@ whole stage by itself.
151
  | `solo/s2_concept` | kinds, properties and differences: what a thing is a kind of, what its kind can do, how two things differ | 0.721 → 0.120 (a second seed: 0.120) | +0.0003 | 30% → 84% | 68% → 73% | two |
152
  | `solo/s3_rules` | if-then rules over made-up words, followed step by step to what follows | 0.930 → 0.372 (a second seed: 0.374) | +0.0009 | 16% → 60% | 7% → 12% | two |
153
  | `solo/s4_arith` | small arithmetic worked out in text | 1.234 → 0.389 (a second seed: 0.388) | +0.0007 | 34% → 49% | 24% → 24% | two |
 
 
 
 
154
 
155
  ```python
156
  m.mount_arm("solo/s1_perspective") # one solo arm; mounting another swaps it
@@ -258,7 +296,7 @@ trunk for each stage; those arms stayed nearly empty, because the trunk took
258
  each stage's text in before its arm could, and they were detached at step
259
  148,000 (their files are in the training repo under `mini-beatrix-3/arms/`,
260
  bound to those earlier trunk states). The arms in this repository were fitted
261
- afterwards on the finished, frozen core: as a group in two steps (first the arms of stages 1 and 2, in a run that added the four stages' text one stage at a time with every attached arm training under one pure Adam, 800 steps per stage; then, with those two held fixed and on, a new arm for stage 3 and after it a new arm for stage 4, 800 steps each), every arm kept quiet on ordinary web text and on its partners' stage text; and each alone, 800 steps, kept quiet on web text.
262
 
263
  ## Lineage
264
 
 
23
 
24
  The package ships **both states in one repository**: the bare model (arms
25
  off) and the model with its **stage arms** mounted (arms on). An arm is a
26
+ small detachable adapter trained on this exact frozen core. The arms come as
27
+ **groups**: arms that were trained switched on together and are mounted
28
+ together. There are two, the four arms of stages 1 to 4 and the eight arms
29
+ of stages 1 to 8 (those four, held fixed, with four more fitted over them).
30
+ Each stage also has a **solo arm**, trained alone, for use one at a time.
31
+ Arms are guests on the core, never a change to it: the core's
32
  weights are the same files either way, and detaching restores the bare
33
  model bit for bit.
34
 
 
41
 
42
  m.say("Who are you?") # arms off: the bare core, its own chat frame
43
 
44
+ m.mount_arm("stages-1-8") # arms on: the stage arms, all on together
45
  m.detach_arm() # the bare core again, verified bit-exact
46
  ```
47
 
 
50
  ```python
51
  m = AutoModelForCausalLM.from_pretrained(
52
  "AbstractPhil/mini-beatrix-3", trust_remote_code=True,
53
+ default_arm="stages-1-8").eval()
54
  ```
55
 
56
  The model reads **raw UTF-8 bytes**: `input_ids` are byte values 0–255.
 
82
  The curriculum taught the model nine kinds of text, one stage after another,
83
  and the finished model moved on from each stage's text as the later stages
84
  came. The stage arms bring that text back without touching the core. They
85
+ were fitted on the finished, frozen core **as always-on groups**:
86
+ the arms of stages 1 and 2 were trained first and then held fixed, the arms of stages 3 and 4 were attached over them one after the other and trained in their presence, and then, with all four held fixed and on, the arms of stages 5, 6, 7 and 8 were attached over them the same way, one after the other, each for 4,000 steps and each kept quiet on every other arm's stage text. The eight-arm group's first four arms are the four-arm group unchanged; the four-arm group is also packaged on its own. `m.arms()` returns the tables below as data, with each row's
87
+ recipe and full measurements.
88
 
89
  **The group `stages-1-4`** (4 arms, 54.8M parameters; two seeds):
90
 
 
110
  | only `stages-1-4/s4_arith` | 0.553 | 0.721 | 0.928 | 1.127 | 0.951 |
111
  | the whole group `stages-1-4` | 0.048 | 0.112 | 0.224 | 0.338 | 0.952 |
112
 
113
+ **The group `stages-1-8`** (8 arms, 109.5M parameters; two seeds):
114
+
115
+ | arm | its stage's text | stage text, arms off → the group on (bpb) | this arm alone (bpb) | stage items in the stage's own form, off → group on | stage items in a new form, off → group on |
116
+ |---|---|---|---|---|---|
117
+ | `stages-1-8/s1_perspective` | one small event retold from each side: seen from a person, said to them, told about them | 0.554 → 0.049 | 0.052 | 55% → 89% | 46% → 44% |
118
+ | `stages-1-8/s2_concept` | kinds, properties and differences: what a thing is a kind of, what its kind can do, how two things differ | 0.721 → 0.111 | 0.350 | 30% → 98% | 68% → 72% |
119
+ | `stages-1-8/s3_rules` | if-then rules over made-up words, followed step by step to what follows | 0.930 → 0.224 | 0.824 | 16% → 86% | 7% → 31% |
120
+ | `stages-1-8/s4_arith` | small arithmetic worked out in text | 1.234 → 0.328 | 1.127 | 34% → 64% | 24% → 24% |
121
+ | `stages-1-8/s5_causal` | short cause-and-effect text: what happened and why | 0.660 → 0.043 | 0.091 | not measured | not measured |
122
+ | `stages-1-8/s6_tryfail` | an attempt, its failure and the revised attempt | 0.673 → 0.055 | 0.164 | not measured | not measured |
123
+ | `stages-1-8/s7_mixed` | the mixed stage's diet: the earlier stages' text beside ordinary prose | 0.783 → 0.720 | 0.724 | not measured | not measured |
124
+ | `stages-1-8/s8_register` | the same content in different registers of speech | 0.491 → 0.021 | 0.068 | not measured | not measured |
125
+
126
+ With the group on, held-out web text moves by +0.0019 bpb (the limit set for it was +0.012), and the nine-suite probe mean reads 0.411 with arms off and 0.422 with the group on.
127
+
128
+ The numbers are those of the packaged weights. A second run of the same recipe from other seeds read the stage texts with its group on at 0.048, 0.121, 0.222, 0.327, 0.043, 0.054, 0.720, 0.021 bpb and moved web text by +0.0010.
129
+
130
+ Each member mounted alone, and the whole group, on every stage's held-out text (bpb):
131
+
132
+ | mounted | stage 1 text | stage 2 text | stage 3 text | stage 4 text | stage 5 text | stage 6 text | stage 7 text | stage 8 text | web text |
133
+ |---|---|---|---|---|---|---|---|---|---|
134
+ | nothing (arms off) | 0.554 | 0.721 | 0.930 | 1.234 | 0.660 | 0.673 | 0.783 | 0.491 | 0.951 |
135
+ | only `stages-1-8/s1_perspective` | 0.053 | 0.239 | 0.509 | 0.723 | 0.463 | 0.495 | 0.773 | 0.379 | 0.952 |
136
+ | only `stages-1-8/s2_concept` | 0.527 | 0.350 | 0.619 | 0.880 | 0.659 | 0.673 | 0.783 | 0.469 | 0.952 |
137
+ | only `stages-1-8/s3_rules` | 0.554 | 0.724 | 0.824 | 1.234 | 0.668 | 0.674 | 0.782 | 0.491 | 0.951 |
138
+ | only `stages-1-8/s4_arith` | 0.553 | 0.721 | 0.928 | 1.127 | 0.657 | 0.659 | 0.783 | 0.489 | 0.951 |
139
+ | only `stages-1-8/s5_causal` | 0.536 | 0.710 | 0.917 | 1.225 | 0.091 | 0.666 | 0.782 | 0.484 | 0.951 |
140
+ | only `stages-1-8/s6_tryfail` | 0.552 | 0.722 | 0.930 | 1.237 | 0.651 | 0.164 | 0.783 | 0.490 | 0.951 |
141
+ | only `stages-1-8/s7_mixed` | 0.497 | 0.715 | 0.868 | 1.195 | 0.558 | 0.593 | 0.724 | 0.483 | 0.950 |
142
+ | only `stages-1-8/s8_register` | 0.537 | 0.713 | 0.925 | 1.227 | 0.642 | 0.651 | 0.783 | 0.068 | 0.951 |
143
+ | the whole group `stages-1-8` | 0.049 | 0.110 | 0.224 | 0.328 | 0.043 | 0.055 | 0.720 | 0.021 | 0.952 |
144
+
145
  **How to read the tables.** *Stage text* is the loss, in bits per byte, on
146
  held-out text of the arm's own stage, with arms off and with the whole group
147
  on. *This arm alone* is the same loss with only that one arm mounted. *Stage
 
150
  not. The *web text* figure is the change on held-out ordinary web text
151
  (fineweb-edu) with the group on: the arms were trained to leave it alone.
152
 
153
+ **A group is the unit.** The eight-arm group is the four-arm group with four more arms fitted over it, and it keeps the four as they were: with all eight on, the text of stages 1 to 4 reads within 0.011 bits per byte of the four-arm group's own reading (stage 4 slightly better, the others the same). The new arms take their stages: with the whole group on, stage 5 text reads 0.043 bits per byte against 0.660 bare, stage 6 0.055 against 0.673, stage 8 0.021 against 0.491, on both seeds. Each of those three carries about two thirds to three quarters of its stage's gain itself and the fixed earlier arms supply the rest (the first arm, a generalist, reads stage 5 and 6 text 0.2 and 0.18 below bare on its own), so again the group works as a whole and is measured that way. The mixed stage (7) is the exception: its text is a blend of the other stages' kinds, the bare model reads it at 0.783, and no arm moves it much (the group reads it at 0.720, its own arm alone at 0.724). The group costs 0.001 to 0.002 bits per byte on ordinary web text, no arm costs more than 0.0006 of that, and no arm writes on another arm's stage text above 0.0012. Nothing an earlier arm learned was lost when a later arm trained over it (the largest loss of an earlier stage's gain at any later close was 2 percent). The four new arms trained for 4,000 steps each, about where a solo arm's curve flattens; the first four kept their shorter training (3,200, 2,400, 800 and 800 steps). For one stage by itself, use its solo arm below.
154
 
155
+ **What the arms do, plainly.** A group recovers each of its stages' own kind
156
+ of text. The gain is in each stage's own forms. Whether it carries over to
157
+ new forms is shown in each table's last column, where it was measured.
158
 
159
  ### Arms off and arms on
160
 
161
  ```python
162
  m.arm # None: arms off
163
+ h = m.mount_arm("stages-1-8") # arms on: the whole eight-arm group
164
  with h.all_off(): # every member masked: the bare core's logits
165
  m.say("Hello there.")
166
  m.detach_arm(verify=True) # raises if the restored core is not bit-exact
167
  m.mount_arm("stages-1-4/s1_perspective") # one member alone (the "alone" readings)
168
  ```
169
 
170
+ A group mounts its members in the order they were trained, always on, one
171
  after another at every block, with no mixer between them. `default_arm` in
172
+ `from_pretrained` (or in `config.json`) mounts a group as the model loads.
173
  `save_pretrained` refuses while an arm is mounted, so an arm can never be
174
  written into the core's weight file.
175
 
 
185
  | `solo/s2_concept` | kinds, properties and differences: what a thing is a kind of, what its kind can do, how two things differ | 0.721 → 0.120 (a second seed: 0.120) | +0.0003 | 30% → 84% | 68% → 73% | two |
186
  | `solo/s3_rules` | if-then rules over made-up words, followed step by step to what follows | 0.930 → 0.372 (a second seed: 0.374) | +0.0009 | 16% → 60% | 7% → 12% | two |
187
  | `solo/s4_arith` | small arithmetic worked out in text | 1.234 → 0.389 (a second seed: 0.388) | +0.0007 | 34% → 49% | 24% → 24% | two |
188
+ | `solo/s5_causal` | short cause-and-effect text: what happened and why | 0.660 → 0.068 | +0.0003 | not measured | not measured | one |
189
+ | `solo/s6_tryfail` | an attempt, its failure and the revised attempt | 0.673 → 0.114 | +0.0008 | not measured | not measured | one |
190
+ | `solo/s7_mixed` | the mixed stage's diet: the earlier stages' text beside ordinary prose | 0.783 → 0.755 | +0.0004 | not measured | not measured | one |
191
+ | `solo/s8_register` | the same content in different registers of speech | 0.491 → 0.051 | +0.0008 | not measured | not measured | one |
192
 
193
  ```python
194
  m.mount_arm("solo/s1_perspective") # one solo arm; mounting another swaps it
 
296
  each stage's text in before its arm could, and they were detached at step
297
  148,000 (their files are in the training repo under `mini-beatrix-3/arms/`,
298
  bound to those earlier trunk states). The arms in this repository were fitted
299
+ afterwards on the finished, frozen core: as always-on groups and as solo arms. The four-arm group in two steps (first the arms of stages 1 and 2, in a run that added the four stages' text one stage at a time with every attached arm training under one pure Adam, 800 steps per stage; then, with those two held fixed and on, a new arm for stage 3 and after it a new arm for stage 4, 800 steps each). The eight-arm group by continuing that run: with the four held fixed and on, a new arm for stage 5, then 6, then 7, then 8, each attached over every arm before it and trained for 4,000 steps in their presence. Every group arm was kept quiet on ordinary web text and on the stage text of every other arm in its group. Each solo arm was trained alone, 800 steps, kept quiet on web text.
300
 
301
  ## Lineage
302
 
arms/index.json CHANGED
@@ -6,8 +6,8 @@
6
  "base_model_id": "alephllm/mini-beatrix-3@step245674",
7
  "note": "every packaged arm was trained on this exact frozen core (the default bf16 weight file); arms are trunk-bound"
8
  },
9
- "generated": "2026-10-06 00:09 UTC",
10
- "note": "the stage arms of curriculum stages 1 to 4, fitted on the finished, frozen core in two forms: as a group that is on together, and each alone; the arms trained beside the trunk during the run are on the training repo",
11
  "arms": [
12
  {
13
  "id": "stages-1-4/s1_perspective",
@@ -526,6 +526,1031 @@
526
  },
527
  "examples": []
528
  },
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
529
  {
530
  "id": "solo/s1_perspective",
531
  "name": "solo/s1_perspective",
@@ -849,6 +1874,214 @@
849
  "examples": [
850
  "Count by 3s: 1 3s make 3; 2 3s make 6; 3 3s make 9; 4 3s make 12; 5 3s make 15. So 3 times 5 is 15."
851
  ]
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
852
  }
853
  ]
854
  }
 
6
  "base_model_id": "alephllm/mini-beatrix-3@step245674",
7
  "note": "every packaged arm was trained on this exact frozen core (the default bf16 weight file); arms are trunk-bound"
8
  },
9
+ "generated": "2026-10-06 04:54 UTC",
10
+ "note": "the stage arms of curriculum stages 1 to 8, fitted on the finished, frozen core: two always-on groups (the four arms of stages 1 to 4; the eight arms of stages 1 to 8, which are those four, held fixed, with four more fitted over them) and each stage's arm alone; the arms trained beside the trunk during the run are on the training repo",
11
  "arms": [
12
  {
13
  "id": "stages-1-4/s1_perspective",
 
526
  },
527
  "examples": []
528
  },
529
+ {
530
+ "id": "stages-1-8/s1_perspective",
531
+ "name": "stages-1-8/s1_perspective",
532
+ "group": "stages-1-8",
533
+ "base_model_id": "alephllm/mini-beatrix-3@step245674",
534
+ "step": 245674,
535
+ "params": 13692960,
536
+ "sites": 32,
537
+ "weights": [
538
+ "arms/stages-1-8/s1_perspective.safetensors"
539
+ ],
540
+ "mount": {
541
+ "type": "anchor"
542
+ },
543
+ "precision": "fp32",
544
+ "template": {
545
+ "frame": "raw",
546
+ "template": "raw stage text, continued"
547
+ },
548
+ "decode": {
549
+ "temperature": 0.0,
550
+ "top_p": 1.0,
551
+ "max_new": 120,
552
+ "precision": "fp32"
553
+ },
554
+ "source": {
555
+ "file": "gXA_close8_s1_perspective.safetensors",
556
+ "sha256": "65cb5c57cfd64fdbaaaf47d32a37969adc91bb2b1071cc97e089943041ebe044"
557
+ },
558
+ "recipe": "trained as an always-on group on the frozen final core by continuing the four-arm group's run. The four arms of stages 1 to 4 were loaded as trained (see the group stages-1-4) and held fixed and on. Then a new arm for stage 5 was attached over them and trained for 4,000 steps in their presence; then, with it also held fixed, a new arm for stage 6 the same way, then stage 7, then stage 8. Each step gives the training arm 2 x 4096 bytes of its own stage's text with every attached arm on, and a quiet term at weight 2 on 2 rows that are not its text (one of web text, one of the stage text of another arm of the eight, drawn from every other stage whether or not its arm was attached yet; the mixed stage's text, which contains the other stages' kinds, is drawn as no arm's partner row): the KL from the model with that arm switched off to the model with every arm on. Pure Adam, lr 0.001",
559
+ "title": "Stage 1: perspective",
560
+ "behavior": "one small event retold from each side: seen from a person, said to them, told about them",
561
+ "family": "stage",
562
+ "status": "a member of the group stages-1-8; trained always on with its partners, so alone is not its trained state",
563
+ "score": "alone: own stage text 0.554 -> 0.052 bpb; web text +0.0014 bpb",
564
+ "measured": {
565
+ "own_stage_bpb_arms_off": 0.5538,
566
+ "own_stage_bpb_group_on": 0.0492,
567
+ "own_stage_bpb_this_arm_alone": 0.0525,
568
+ "share_of_the_group_on_own_stage": -0.3832,
569
+ "share_of_the_group_on_own_stage_se": 0.0052,
570
+ "share_of_the_group_on_web_text": 0.00025,
571
+ "alone_on_web_text": 0.00143,
572
+ "share_of_the_group_on_partner_stages": {
573
+ "s2_concept": -0.2174,
574
+ "s3_rules": -0.1832,
575
+ "s4_arith": -0.128,
576
+ "s5_causal": -0.0128,
577
+ "s6_tryfail": -0.028,
578
+ "s7_mixed": -0.0029,
579
+ "s8_register": -0.0272
580
+ },
581
+ "alone_on_partner_stages": {
582
+ "s2_concept": -0.4817,
583
+ "s3_rules": -0.4209,
584
+ "s4_arith": -0.5113,
585
+ "s5_causal": -0.1978,
586
+ "s6_tryfail": -0.1782,
587
+ "s7_mixed": -0.0101,
588
+ "s8_register": -0.1119
589
+ },
590
+ "steps_trained": 3200,
591
+ "stage_items_own_form_accuracy": {
592
+ "arms_off": 0.55,
593
+ "group_on": 0.887,
594
+ "this_arm_alone": 0.838
595
+ },
596
+ "stage_items_new_form_accuracy": {
597
+ "arms_off": 0.463,
598
+ "group_on": 0.438,
599
+ "this_arm_alone": 0.463
600
+ },
601
+ "second_seed": {
602
+ "own_stage_bpb_arms_off": 0.5538,
603
+ "own_stage_bpb_group_on": 0.0484,
604
+ "own_stage_bpb_this_arm_alone": 0.0522,
605
+ "share_of_the_group_on_own_stage": -0.4405,
606
+ "share_of_the_group_on_own_stage_se": 0.0064,
607
+ "share_of_the_group_on_web_text": 0.00043,
608
+ "alone_on_web_text": 0.00134,
609
+ "share_of_the_group_on_partner_stages": {
610
+ "s2_concept": -0.2289,
611
+ "s3_rules": -0.1943,
612
+ "s4_arith": -0.13,
613
+ "s5_causal": -0.0282,
614
+ "s6_tryfail": -0.0206,
615
+ "s7_mixed": -0.0043,
616
+ "s8_register": -0.0214
617
+ },
618
+ "alone_on_partner_stages": {
619
+ "s2_concept": -0.4772,
620
+ "s3_rules": -0.4262,
621
+ "s4_arith": -0.4807,
622
+ "s5_causal": -0.1919,
623
+ "s6_tryfail": -0.169,
624
+ "s7_mixed": -0.0088,
625
+ "s8_register": -0.1105
626
+ },
627
+ "steps_trained": 3200,
628
+ "stage_items_own_form_accuracy": {
629
+ "arms_off": 0.55,
630
+ "group_on": 1.0,
631
+ "this_arm_alone": 0.925
632
+ },
633
+ "stage_items_new_form_accuracy": {
634
+ "arms_off": 0.463,
635
+ "group_on": 0.425,
636
+ "this_arm_alone": 0.463
637
+ }
638
+ }
639
+ },
640
+ "examples": [
641
+ "Rui dropped the acorn near the box. Seen from Rui: I dropped the acorn near the box. Said to Rui: you dropped the acorn near the box. Told about Rui: Rui dropped the acorn near the box. Three ways of saying, one thing that happened."
642
+ ]
643
+ },
644
+ {
645
+ "id": "stages-1-8/s2_concept",
646
+ "name": "stages-1-8/s2_concept",
647
+ "group": "stages-1-8",
648
+ "base_model_id": "alephllm/mini-beatrix-3@step245674",
649
+ "step": 245674,
650
+ "params": 13692960,
651
+ "sites": 32,
652
+ "weights": [
653
+ "arms/stages-1-8/s2_concept.safetensors"
654
+ ],
655
+ "mount": {
656
+ "type": "anchor"
657
+ },
658
+ "precision": "fp32",
659
+ "template": {
660
+ "frame": "raw",
661
+ "template": "raw stage text, continued"
662
+ },
663
+ "decode": {
664
+ "temperature": 0.0,
665
+ "top_p": 1.0,
666
+ "max_new": 120,
667
+ "precision": "fp32"
668
+ },
669
+ "source": {
670
+ "file": "gXA_close8_s2_concept.safetensors",
671
+ "sha256": "f411b5c9ae9994a26e0cbd25475a754fee5af6b7e58fb2af5dedbe790dd602f5"
672
+ },
673
+ "recipe": "trained as an always-on group on the frozen final core by continuing the four-arm group's run. The four arms of stages 1 to 4 were loaded as trained (see the group stages-1-4) and held fixed and on. Then a new arm for stage 5 was attached over them and trained for 4,000 steps in their presence; then, with it also held fixed, a new arm for stage 6 the same way, then stage 7, then stage 8. Each step gives the training arm 2 x 4096 bytes of its own stage's text with every attached arm on, and a quiet term at weight 2 on 2 rows that are not its text (one of web text, one of the stage text of another arm of the eight, drawn from every other stage whether or not its arm was attached yet; the mixed stage's text, which contains the other stages' kinds, is drawn as no arm's partner row): the KL from the model with that arm switched off to the model with every arm on. Pure Adam, lr 0.001",
674
+ "title": "Stage 2: concepts",
675
+ "behavior": "kinds, properties and differences: what a thing is a kind of, what its kind can do, how two things differ",
676
+ "family": "stage",
677
+ "status": "a member of the group stages-1-8; trained always on with its partners, so alone is not its trained state",
678
+ "score": "alone: own stage text 0.721 -> 0.350 bpb; web text +0.0010 bpb",
679
+ "measured": {
680
+ "own_stage_bpb_arms_off": 0.7212,
681
+ "own_stage_bpb_group_on": 0.1105,
682
+ "own_stage_bpb_this_arm_alone": 0.3503,
683
+ "share_of_the_group_on_own_stage": -0.1202,
684
+ "share_of_the_group_on_own_stage_se": 0.0011,
685
+ "share_of_the_group_on_web_text": -0.00043,
686
+ "alone_on_web_text": 0.00105,
687
+ "share_of_the_group_on_partner_stages": {
688
+ "s1_perspective": -0.0057,
689
+ "s3_rules": -0.1086,
690
+ "s4_arith": -0.106,
691
+ "s5_causal": -0.0013,
692
+ "s6_tryfail": -0.0035,
693
+ "s7_mixed": -0.0008,
694
+ "s8_register": -0.0007
695
+ },
696
+ "alone_on_partner_stages": {
697
+ "s1_perspective": -0.0273,
698
+ "s3_rules": -0.3108,
699
+ "s4_arith": -0.3537,
700
+ "s5_causal": -0.0015,
701
+ "s6_tryfail": 0.0004,
702
+ "s7_mixed": -0.0,
703
+ "s8_register": -0.0215
704
+ },
705
+ "steps_trained": 2400,
706
+ "stage_items_own_form_accuracy": {
707
+ "arms_off": 0.298,
708
+ "group_on": 0.976,
709
+ "this_arm_alone": 0.964
710
+ },
711
+ "stage_items_new_form_accuracy": {
712
+ "arms_off": 0.678,
713
+ "group_on": 0.72,
714
+ "this_arm_alone": 0.72
715
+ },
716
+ "second_seed": {
717
+ "own_stage_bpb_arms_off": 0.7212,
718
+ "own_stage_bpb_group_on": 0.1212,
719
+ "own_stage_bpb_this_arm_alone": 0.3603,
720
+ "share_of_the_group_on_own_stage": -0.119,
721
+ "share_of_the_group_on_own_stage_se": 0.0012,
722
+ "share_of_the_group_on_web_text": -0.00073,
723
+ "alone_on_web_text": 0.00086,
724
+ "share_of_the_group_on_partner_stages": {
725
+ "s1_perspective": -0.0042,
726
+ "s3_rules": -0.1006,
727
+ "s4_arith": -0.0957,
728
+ "s5_causal": -0.0009,
729
+ "s6_tryfail": -0.0023,
730
+ "s7_mixed": -0.0004,
731
+ "s8_register": -0.0005
732
+ },
733
+ "alone_on_partner_stages": {
734
+ "s1_perspective": -0.03,
735
+ "s3_rules": -0.3138,
736
+ "s4_arith": -0.3577,
737
+ "s5_causal": -0.0127,
738
+ "s6_tryfail": -0.0062,
739
+ "s7_mixed": 0.0002,
740
+ "s8_register": -0.0193
741
+ },
742
+ "steps_trained": 2400,
743
+ "stage_items_own_form_accuracy": {
744
+ "arms_off": 0.298,
745
+ "group_on": 0.952,
746
+ "this_arm_alone": 0.905
747
+ },
748
+ "stage_items_new_form_accuracy": {
749
+ "arms_off": 0.678,
750
+ "group_on": 0.72,
751
+ "this_arm_alone": 0.703
752
+ }
753
+ }
754
+ },
755
+ "examples": [
756
+ "A whale is a kind of mammal. Most mammals can swim. A whale is not a beetle; they differ in kind."
757
+ ]
758
+ },
759
+ {
760
+ "id": "stages-1-8/s3_rules",
761
+ "name": "stages-1-8/s3_rules",
762
+ "group": "stages-1-8",
763
+ "base_model_id": "alephllm/mini-beatrix-3@step245674",
764
+ "step": 245674,
765
+ "params": 13692960,
766
+ "sites": 32,
767
+ "weights": [
768
+ "arms/stages-1-8/s3_rules.safetensors"
769
+ ],
770
+ "mount": {
771
+ "type": "anchor"
772
+ },
773
+ "precision": "fp32",
774
+ "template": {
775
+ "frame": "raw",
776
+ "template": "raw stage text, continued"
777
+ },
778
+ "decode": {
779
+ "temperature": 0.0,
780
+ "top_p": 1.0,
781
+ "max_new": 120,
782
+ "precision": "fp32"
783
+ },
784
+ "source": {
785
+ "file": "gXA_close8_s3_rules.safetensors",
786
+ "sha256": "bd9f4dd7e15f28ca140e1c86962b0f4ea4b144a8af369219acd270b1368dc5b8"
787
+ },
788
+ "recipe": "trained as an always-on group on the frozen final core by continuing the four-arm group's run. The four arms of stages 1 to 4 were loaded as trained (see the group stages-1-4) and held fixed and on. Then a new arm for stage 5 was attached over them and trained for 4,000 steps in their presence; then, with it also held fixed, a new arm for stage 6 the same way, then stage 7, then stage 8. Each step gives the training arm 2 x 4096 bytes of its own stage's text with every attached arm on, and a quiet term at weight 2 on 2 rows that are not its text (one of web text, one of the stage text of another arm of the eight, drawn from every other stage whether or not its arm was attached yet; the mixed stage's text, which contains the other stages' kinds, is drawn as no arm's partner row): the KL from the model with that arm switched off to the model with every arm on. Pure Adam, lr 0.001",
789
+ "title": "Stage 3: rule chains",
790
+ "behavior": "if-then rules over made-up words, followed step by step to what follows",
791
+ "family": "stage",
792
+ "status": "a member of the group stages-1-8; trained always on with its partners, so alone is not its trained state",
793
+ "score": "alone: own stage text 0.930 -> 0.824 bpb; web text +0.0005 bpb",
794
+ "measured": {
795
+ "own_stage_bpb_arms_off": 0.9296,
796
+ "own_stage_bpb_group_on": 0.2237,
797
+ "own_stage_bpb_this_arm_alone": 0.8236,
798
+ "share_of_the_group_on_own_stage": -0.09,
799
+ "share_of_the_group_on_own_stage_se": 0.0013,
800
+ "share_of_the_group_on_web_text": 9e-05,
801
+ "alone_on_web_text": 0.00053,
802
+ "share_of_the_group_on_partner_stages": {
803
+ "s1_perspective": 0.0001,
804
+ "s2_concept": 0.0005,
805
+ "s4_arith": -0.0002,
806
+ "s5_causal": -0.0001,
807
+ "s6_tryfail": -0.0003,
808
+ "s7_mixed": 0.0001,
809
+ "s8_register": -0.0001
810
+ },
811
+ "alone_on_partner_stages": {
812
+ "s1_perspective": -0.0003,
813
+ "s2_concept": 0.0031,
814
+ "s4_arith": 0.0001,
815
+ "s5_causal": 0.0071,
816
+ "s6_tryfail": 0.001,
817
+ "s7_mixed": -0.0005,
818
+ "s8_register": 0.0002
819
+ },
820
+ "steps_trained": 800,
821
+ "stage_items_own_form_accuracy": {
822
+ "arms_off": 0.155,
823
+ "group_on": 0.855,
824
+ "this_arm_alone": 0.32
825
+ },
826
+ "stage_items_new_form_accuracy": {
827
+ "arms_off": 0.07,
828
+ "group_on": 0.31,
829
+ "this_arm_alone": 0.08
830
+ },
831
+ "second_seed": {
832
+ "own_stage_bpb_arms_off": 0.9296,
833
+ "own_stage_bpb_group_on": 0.2222,
834
+ "own_stage_bpb_this_arm_alone": 0.8702,
835
+ "share_of_the_group_on_own_stage": -0.0915,
836
+ "share_of_the_group_on_own_stage_se": 0.0015,
837
+ "share_of_the_group_on_web_text": 0.00059,
838
+ "alone_on_web_text": 0.00034,
839
+ "share_of_the_group_on_partner_stages": {
840
+ "s1_perspective": 0.0,
841
+ "s2_concept": -0.0005,
842
+ "s4_arith": -0.0004,
843
+ "s5_causal": -0.0001,
844
+ "s6_tryfail": -0.0003,
845
+ "s7_mixed": -0.0001,
846
+ "s8_register": 0.0
847
+ },
848
+ "alone_on_partner_stages": {
849
+ "s1_perspective": -0.0003,
850
+ "s2_concept": 0.0016,
851
+ "s4_arith": -0.0047,
852
+ "s5_causal": 0.006,
853
+ "s6_tryfail": -0.0011,
854
+ "s7_mixed": -0.0007,
855
+ "s8_register": -0.0044
856
+ },
857
+ "steps_trained": 800,
858
+ "stage_items_own_form_accuracy": {
859
+ "arms_off": 0.155,
860
+ "group_on": 0.92,
861
+ "this_arm_alone": 0.335
862
+ },
863
+ "stage_items_new_form_accuracy": {
864
+ "arms_off": 0.07,
865
+ "group_on": 0.37,
866
+ "this_arm_alone": 0.1
867
+ }
868
+ }
869
+ },
870
+ "examples": [
871
+ "If someone is hoream, then they are hortil. If someone is hortil, then they are babig. If someone is babig, then they are banng. Vin is hoream. So Vin is hortil. So Vin is babig. So Vin is banng."
872
+ ]
873
+ },
874
+ {
875
+ "id": "stages-1-8/s4_arith",
876
+ "name": "stages-1-8/s4_arith",
877
+ "group": "stages-1-8",
878
+ "base_model_id": "alephllm/mini-beatrix-3@step245674",
879
+ "step": 245674,
880
+ "params": 13692960,
881
+ "sites": 32,
882
+ "weights": [
883
+ "arms/stages-1-8/s4_arith.safetensors"
884
+ ],
885
+ "mount": {
886
+ "type": "anchor"
887
+ },
888
+ "precision": "fp32",
889
+ "template": {
890
+ "frame": "raw",
891
+ "template": "raw stage text, continued"
892
+ },
893
+ "decode": {
894
+ "temperature": 0.0,
895
+ "top_p": 1.0,
896
+ "max_new": 120,
897
+ "precision": "fp32"
898
+ },
899
+ "source": {
900
+ "file": "gXA_close8_s4_arith.safetensors",
901
+ "sha256": "58552d2cdf11bce31dcad8e1e09fe27b6c760857aef47add6b4c92b0487669b7"
902
+ },
903
+ "recipe": "trained as an always-on group on the frozen final core by continuing the four-arm group's run. The four arms of stages 1 to 4 were loaded as trained (see the group stages-1-4) and held fixed and on. Then a new arm for stage 5 was attached over them and trained for 4,000 steps in their presence; then, with it also held fixed, a new arm for stage 6 the same way, then stage 7, then stage 8. Each step gives the training arm 2 x 4096 bytes of its own stage's text with every attached arm on, and a quiet term at weight 2 on 2 rows that are not its text (one of web text, one of the stage text of another arm of the eight, drawn from every other stage whether or not its arm was attached yet; the mixed stage's text, which contains the other stages' kinds, is drawn as no arm's partner row): the KL from the model with that arm switched off to the model with every arm on. Pure Adam, lr 0.001",
904
+ "title": "Stage 4: arithmetic",
905
+ "behavior": "small arithmetic worked out in text",
906
+ "family": "stage",
907
+ "status": "a member of the group stages-1-8; trained always on with its partners, so alone is not its trained state",
908
+ "score": "alone: own stage text 1.234 -> 1.127 bpb; web text +0.0001 bpb",
909
+ "measured": {
910
+ "own_stage_bpb_arms_off": 1.2339,
911
+ "own_stage_bpb_group_on": 0.3279,
912
+ "own_stage_bpb_this_arm_alone": 1.1267,
913
+ "share_of_the_group_on_own_stage": -0.1431,
914
+ "share_of_the_group_on_own_stage_se": 0.0023,
915
+ "share_of_the_group_on_web_text": 0.0002,
916
+ "alone_on_web_text": 0.00011,
917
+ "share_of_the_group_on_partner_stages": {
918
+ "s1_perspective": 0.0,
919
+ "s2_concept": -0.0006,
920
+ "s3_rules": -0.0001,
921
+ "s5_causal": -0.0,
922
+ "s6_tryfail": -0.0036,
923
+ "s7_mixed": 0.0001,
924
+ "s8_register": -0.0
925
+ },
926
+ "alone_on_partner_stages": {
927
+ "s1_perspective": -0.0009,
928
+ "s2_concept": -0.0004,
929
+ "s3_rules": -0.0016,
930
+ "s5_causal": -0.0033,
931
+ "s6_tryfail": -0.0135,
932
+ "s7_mixed": -0.0002,
933
+ "s8_register": -0.0015
934
+ },
935
+ "steps_trained": 800,
936
+ "stage_items_own_form_accuracy": {
937
+ "arms_off": 0.344,
938
+ "group_on": 0.644,
939
+ "this_arm_alone": 0.381
940
+ },
941
+ "stage_items_new_form_accuracy": {
942
+ "arms_off": 0.242,
943
+ "group_on": 0.242,
944
+ "this_arm_alone": 0.246
945
+ },
946
+ "second_seed": {
947
+ "own_stage_bpb_arms_off": 1.2339,
948
+ "own_stage_bpb_group_on": 0.3266,
949
+ "own_stage_bpb_this_arm_alone": 1.155,
950
+ "share_of_the_group_on_own_stage": -0.1643,
951
+ "share_of_the_group_on_own_stage_se": 0.0019,
952
+ "share_of_the_group_on_web_text": 0.00017,
953
+ "alone_on_web_text": 6e-05,
954
+ "share_of_the_group_on_partner_stages": {
955
+ "s1_perspective": 0.0,
956
+ "s2_concept": -0.0002,
957
+ "s3_rules": 0.0001,
958
+ "s5_causal": -0.0002,
959
+ "s6_tryfail": -0.0039,
960
+ "s7_mixed": -0.0,
961
+ "s8_register": -0.0001
962
+ },
963
+ "alone_on_partner_stages": {
964
+ "s1_perspective": -0.0013,
965
+ "s2_concept": -0.0015,
966
+ "s3_rules": 0.0005,
967
+ "s5_causal": -0.0011,
968
+ "s6_tryfail": -0.0134,
969
+ "s7_mixed": -0.0003,
970
+ "s8_register": -0.0014
971
+ },
972
+ "steps_trained": 800,
973
+ "stage_items_own_form_accuracy": {
974
+ "arms_off": 0.344,
975
+ "group_on": 0.656,
976
+ "this_arm_alone": 0.412
977
+ },
978
+ "stage_items_new_form_accuracy": {
979
+ "arms_off": 0.242,
980
+ "group_on": 0.271,
981
+ "this_arm_alone": 0.237
982
+ }
983
+ }
984
+ },
985
+ "examples": [
986
+ "Count by 3s: 1 3s make 3; 2 3s make 6; 3 3s make 9; 4 3s make 12; 5 3s make 15. So 3 times 5 is 15."
987
+ ]
988
+ },
989
+ {
990
+ "id": "stages-1-8/s5_causal",
991
+ "name": "stages-1-8/s5_causal",
992
+ "group": "stages-1-8",
993
+ "base_model_id": "alephllm/mini-beatrix-3@step245674",
994
+ "step": 245674,
995
+ "params": 13692960,
996
+ "sites": 32,
997
+ "weights": [
998
+ "arms/stages-1-8/s5_causal.safetensors"
999
+ ],
1000
+ "mount": {
1001
+ "type": "anchor"
1002
+ },
1003
+ "precision": "fp32",
1004
+ "template": {
1005
+ "frame": "raw",
1006
+ "template": "raw stage text, continued"
1007
+ },
1008
+ "decode": {
1009
+ "temperature": 0.0,
1010
+ "top_p": 1.0,
1011
+ "max_new": 120,
1012
+ "precision": "fp32"
1013
+ },
1014
+ "source": {
1015
+ "file": "gXA_close8_s5_causal.safetensors",
1016
+ "sha256": "8d1da9acaf7d528243304d7bebf71780242c7b1622ab3a1f72ef39e875b0f732"
1017
+ },
1018
+ "recipe": "trained as an always-on group on the frozen final core by continuing the four-arm group's run. The four arms of stages 1 to 4 were loaded as trained (see the group stages-1-4) and held fixed and on. Then a new arm for stage 5 was attached over them and trained for 4,000 steps in their presence; then, with it also held fixed, a new arm for stage 6 the same way, then stage 7, then stage 8. Each step gives the training arm 2 x 4096 bytes of its own stage's text with every attached arm on, and a quiet term at weight 2 on 2 rows that are not its text (one of web text, one of the stage text of another arm of the eight, drawn from every other stage whether or not its arm was attached yet; the mixed stage's text, which contains the other stages' kinds, is drawn as no arm's partner row): the KL from the model with that arm switched off to the model with every arm on. Pure Adam, lr 0.001",
1019
+ "title": "Stage 5: cause and effect",
1020
+ "behavior": "short cause-and-effect text: what happened and why",
1021
+ "family": "stage",
1022
+ "status": "a member of the group stages-1-8; trained always on with its partners, so alone is not its trained state",
1023
+ "score": "alone: own stage text 0.660 -> 0.091 bpb; web text +0.0003 bpb",
1024
+ "measured": {
1025
+ "own_stage_bpb_arms_off": 0.6604,
1026
+ "own_stage_bpb_group_on": 0.0431,
1027
+ "own_stage_bpb_this_arm_alone": 0.0911,
1028
+ "share_of_the_group_on_own_stage": -0.415,
1029
+ "share_of_the_group_on_own_stage_se": 0.0056,
1030
+ "share_of_the_group_on_web_text": 0.00035,
1031
+ "alone_on_web_text": 0.00028,
1032
+ "share_of_the_group_on_partner_stages": {
1033
+ "s1_perspective": 0.0007,
1034
+ "s2_concept": -0.0005,
1035
+ "s3_rules": 0.0005,
1036
+ "s4_arith": 0.0002,
1037
+ "s6_tryfail": -0.0005,
1038
+ "s7_mixed": 0.0004,
1039
+ "s8_register": -0.0003
1040
+ },
1041
+ "alone_on_partner_stages": {
1042
+ "s1_perspective": -0.0183,
1043
+ "s2_concept": -0.0109,
1044
+ "s3_rules": -0.0122,
1045
+ "s4_arith": -0.0091,
1046
+ "s6_tryfail": -0.0065,
1047
+ "s7_mixed": -0.0011,
1048
+ "s8_register": -0.0064
1049
+ },
1050
+ "steps_trained": 4000,
1051
+ "second_seed": {
1052
+ "own_stage_bpb_arms_off": 0.6604,
1053
+ "own_stage_bpb_group_on": 0.0428,
1054
+ "own_stage_bpb_this_arm_alone": 0.1418,
1055
+ "share_of_the_group_on_own_stage": -0.4245,
1056
+ "share_of_the_group_on_own_stage_se": 0.0045,
1057
+ "share_of_the_group_on_web_text": 0.00029,
1058
+ "alone_on_web_text": 0.0,
1059
+ "share_of_the_group_on_partner_stages": {
1060
+ "s1_perspective": 0.0001,
1061
+ "s2_concept": 0.0002,
1062
+ "s3_rules": 0.0004,
1063
+ "s4_arith": 0.0001,
1064
+ "s6_tryfail": -0.0003,
1065
+ "s7_mixed": 0.0001,
1066
+ "s8_register": -0.0001
1067
+ },
1068
+ "alone_on_partner_stages": {
1069
+ "s1_perspective": -0.0022,
1070
+ "s2_concept": -0.001,
1071
+ "s3_rules": -0.0018,
1072
+ "s4_arith": -0.0022,
1073
+ "s6_tryfail": -0.0023,
1074
+ "s7_mixed": -0.0001,
1075
+ "s8_register": -0.0009
1076
+ },
1077
+ "steps_trained": 4000
1078
+ }
1079
+ },
1080
+ "examples": [
1081
+ "Noor stacked the cups too high. Because of that, the tower leaned. Because of that, the cups crashed down."
1082
+ ]
1083
+ },
1084
+ {
1085
+ "id": "stages-1-8/s6_tryfail",
1086
+ "name": "stages-1-8/s6_tryfail",
1087
+ "group": "stages-1-8",
1088
+ "base_model_id": "alephllm/mini-beatrix-3@step245674",
1089
+ "step": 245674,
1090
+ "params": 13692960,
1091
+ "sites": 32,
1092
+ "weights": [
1093
+ "arms/stages-1-8/s6_tryfail.safetensors"
1094
+ ],
1095
+ "mount": {
1096
+ "type": "anchor"
1097
+ },
1098
+ "precision": "fp32",
1099
+ "template": {
1100
+ "frame": "raw",
1101
+ "template": "raw stage text, continued"
1102
+ },
1103
+ "decode": {
1104
+ "temperature": 0.0,
1105
+ "top_p": 1.0,
1106
+ "max_new": 120,
1107
+ "precision": "fp32"
1108
+ },
1109
+ "source": {
1110
+ "file": "gXA_close8_s6_tryfail.safetensors",
1111
+ "sha256": "2b7942956f7e2ded57b29168ef5a0916424a39bd3826ad7804fabec92050999b"
1112
+ },
1113
+ "recipe": "trained as an always-on group on the frozen final core by continuing the four-arm group's run. The four arms of stages 1 to 4 were loaded as trained (see the group stages-1-4) and held fixed and on. Then a new arm for stage 5 was attached over them and trained for 4,000 steps in their presence; then, with it also held fixed, a new arm for stage 6 the same way, then stage 7, then stage 8. Each step gives the training arm 2 x 4096 bytes of its own stage's text with every attached arm on, and a quiet term at weight 2 on 2 rows that are not its text (one of web text, one of the stage text of another arm of the eight, drawn from every other stage whether or not its arm was attached yet; the mixed stage's text, which contains the other stages' kinds, is drawn as no arm's partner row): the KL from the model with that arm switched off to the model with every arm on. Pure Adam, lr 0.001",
1114
+ "title": "Stage 6: try, fail, revise",
1115
+ "behavior": "an attempt, its failure and the revised attempt",
1116
+ "family": "stage",
1117
+ "status": "a member of the group stages-1-8; trained always on with its partners, so alone is not its trained state",
1118
+ "score": "alone: own stage text 0.673 -> 0.164 bpb; web text +0.0001 bpb",
1119
+ "measured": {
1120
+ "own_stage_bpb_arms_off": 0.6728,
1121
+ "own_stage_bpb_group_on": 0.0548,
1122
+ "own_stage_bpb_this_arm_alone": 0.1638,
1123
+ "share_of_the_group_on_own_stage": -0.4138,
1124
+ "share_of_the_group_on_own_stage_se": 0.0033,
1125
+ "share_of_the_group_on_web_text": -0.00025,
1126
+ "alone_on_web_text": 0.00013,
1127
+ "share_of_the_group_on_partner_stages": {
1128
+ "s1_perspective": -0.0001,
1129
+ "s2_concept": 0.0002,
1130
+ "s3_rules": -0.0002,
1131
+ "s4_arith": -0.0009,
1132
+ "s5_causal": -0.0001,
1133
+ "s7_mixed": 0.0001,
1134
+ "s8_register": -0.0006
1135
+ },
1136
+ "alone_on_partner_stages": {
1137
+ "s1_perspective": -0.002,
1138
+ "s2_concept": 0.0012,
1139
+ "s3_rules": 0.0006,
1140
+ "s4_arith": 0.0036,
1141
+ "s5_causal": -0.0098,
1142
+ "s7_mixed": -0.0,
1143
+ "s8_register": -0.0003
1144
+ },
1145
+ "steps_trained": 4000,
1146
+ "second_seed": {
1147
+ "own_stage_bpb_arms_off": 0.6728,
1148
+ "own_stage_bpb_group_on": 0.0543,
1149
+ "own_stage_bpb_this_arm_alone": 0.1437,
1150
+ "share_of_the_group_on_own_stage": -0.412,
1151
+ "share_of_the_group_on_own_stage_se": 0.0027,
1152
+ "share_of_the_group_on_web_text": 0.00025,
1153
+ "alone_on_web_text": 0.00011,
1154
+ "share_of_the_group_on_partner_stages": {
1155
+ "s1_perspective": 0.0,
1156
+ "s2_concept": 0.0003,
1157
+ "s3_rules": 0.0,
1158
+ "s4_arith": -0.0016,
1159
+ "s5_causal": 0.0,
1160
+ "s7_mixed": -0.0001,
1161
+ "s8_register": -0.0005
1162
+ },
1163
+ "alone_on_partner_stages": {
1164
+ "s1_perspective": 0.0029,
1165
+ "s2_concept": 0.0013,
1166
+ "s3_rules": -0.0003,
1167
+ "s4_arith": -0.001,
1168
+ "s5_causal": -0.0026,
1169
+ "s7_mixed": 0.0001,
1170
+ "s8_register": -0.0001
1171
+ },
1172
+ "steps_trained": 4000
1173
+ }
1174
+ },
1175
+ "examples": [
1176
+ "Talia adds 34 and 23 and writes 65. Talia checks by counting back: 65 is too big. Talia tries again carefully: 34 + 23 = 57. The check works now, and Talia keeps the good method."
1177
+ ]
1178
+ },
1179
+ {
1180
+ "id": "stages-1-8/s7_mixed",
1181
+ "name": "stages-1-8/s7_mixed",
1182
+ "group": "stages-1-8",
1183
+ "base_model_id": "alephllm/mini-beatrix-3@step245674",
1184
+ "step": 245674,
1185
+ "params": 13692960,
1186
+ "sites": 32,
1187
+ "weights": [
1188
+ "arms/stages-1-8/s7_mixed.safetensors"
1189
+ ],
1190
+ "mount": {
1191
+ "type": "anchor"
1192
+ },
1193
+ "precision": "fp32",
1194
+ "template": {
1195
+ "frame": "raw",
1196
+ "template": "raw stage text, continued"
1197
+ },
1198
+ "decode": {
1199
+ "temperature": 0.0,
1200
+ "top_p": 1.0,
1201
+ "max_new": 120,
1202
+ "precision": "fp32"
1203
+ },
1204
+ "source": {
1205
+ "file": "gXA_close8_s7_mixed.safetensors",
1206
+ "sha256": "b7a4210d7ab9690a5fbddf847abbe60a47755d9d8446dbbb04fe6ba469e8f23c"
1207
+ },
1208
+ "recipe": "trained as an always-on group on the frozen final core by continuing the four-arm group's run. The four arms of stages 1 to 4 were loaded as trained (see the group stages-1-4) and held fixed and on. Then a new arm for stage 5 was attached over them and trained for 4,000 steps in their presence; then, with it also held fixed, a new arm for stage 6 the same way, then stage 7, then stage 8. Each step gives the training arm 2 x 4096 bytes of its own stage's text with every attached arm on, and a quiet term at weight 2 on 2 rows that are not its text (one of web text, one of the stage text of another arm of the eight, drawn from every other stage whether or not its arm was attached yet; the mixed stage's text, which contains the other stages' kinds, is drawn as no arm's partner row): the KL from the model with that arm switched off to the model with every arm on. Pure Adam, lr 0.001",
1209
+ "title": "Stage 7: mixed",
1210
+ "behavior": "the mixed stage's diet: the earlier stages' text beside ordinary prose",
1211
+ "family": "stage",
1212
+ "status": "a member of the group stages-1-8; trained always on with its partners, so alone is not its trained state",
1213
+ "score": "alone: own stage text 0.783 -> 0.724 bpb; web text -0.0002 bpb",
1214
+ "measured": {
1215
+ "own_stage_bpb_arms_off": 0.783,
1216
+ "own_stage_bpb_group_on": 0.72,
1217
+ "own_stage_bpb_this_arm_alone": 0.7242,
1218
+ "share_of_the_group_on_own_stage": -0.0568,
1219
+ "share_of_the_group_on_own_stage_se": 0.0328,
1220
+ "share_of_the_group_on_web_text": -0.00096,
1221
+ "alone_on_web_text": -0.00019,
1222
+ "share_of_the_group_on_partner_stages": {
1223
+ "s1_perspective": -0.0004,
1224
+ "s2_concept": -0.0005,
1225
+ "s3_rules": -0.001,
1226
+ "s4_arith": -0.0103,
1227
+ "s5_causal": -0.0002,
1228
+ "s6_tryfail": -0.0004,
1229
+ "s8_register": -0.0001
1230
+ },
1231
+ "alone_on_partner_stages": {
1232
+ "s1_perspective": -0.0568,
1233
+ "s2_concept": -0.0058,
1234
+ "s3_rules": -0.0617,
1235
+ "s4_arith": -0.0385,
1236
+ "s5_causal": -0.1028,
1237
+ "s6_tryfail": -0.0802,
1238
+ "s8_register": -0.008
1239
+ },
1240
+ "steps_trained": 4000,
1241
+ "second_seed": {
1242
+ "own_stage_bpb_arms_off": 0.783,
1243
+ "own_stage_bpb_group_on": 0.7198,
1244
+ "own_stage_bpb_this_arm_alone": 0.7254,
1245
+ "share_of_the_group_on_own_stage": -0.0569,
1246
+ "share_of_the_group_on_own_stage_se": 0.0329,
1247
+ "share_of_the_group_on_web_text": -0.0007,
1248
+ "alone_on_web_text": -0.00032,
1249
+ "share_of_the_group_on_partner_stages": {
1250
+ "s1_perspective": 0.0001,
1251
+ "s2_concept": -0.0003,
1252
+ "s3_rules": -0.0007,
1253
+ "s4_arith": -0.0099,
1254
+ "s5_causal": 0.0,
1255
+ "s6_tryfail": -0.0003,
1256
+ "s8_register": -0.0002
1257
+ },
1258
+ "alone_on_partner_stages": {
1259
+ "s1_perspective": -0.0237,
1260
+ "s2_concept": -0.0022,
1261
+ "s3_rules": -0.0529,
1262
+ "s4_arith": -0.0414,
1263
+ "s5_causal": -0.0976,
1264
+ "s6_tryfail": -0.0774,
1265
+ "s8_register": -0.0061
1266
+ },
1267
+ "steps_trained": 4000
1268
+ }
1269
+ },
1270
+ "examples": [
1271
+ "It has been compared to a chant, a rhythmic divine beauty, a melody, an aria, a toccata, an edification, an exaltation. As poetry is for the tongue, calligraphy is to the page. The authors of The Splendor of Islamic Calligraphy put it best when they said: “Calligraphy is the plainsong of the divine”.\nCalligraphy is the art of the linear graphic, but it is more than that. In Isl"
1272
+ ]
1273
+ },
1274
+ {
1275
+ "id": "stages-1-8/s8_register",
1276
+ "name": "stages-1-8/s8_register",
1277
+ "group": "stages-1-8",
1278
+ "base_model_id": "alephllm/mini-beatrix-3@step245674",
1279
+ "step": 245674,
1280
+ "params": 13692960,
1281
+ "sites": 32,
1282
+ "weights": [
1283
+ "arms/stages-1-8/s8_register.safetensors"
1284
+ ],
1285
+ "mount": {
1286
+ "type": "anchor"
1287
+ },
1288
+ "precision": "fp32",
1289
+ "template": {
1290
+ "frame": "raw",
1291
+ "template": "raw stage text, continued"
1292
+ },
1293
+ "decode": {
1294
+ "temperature": 0.0,
1295
+ "top_p": 1.0,
1296
+ "max_new": 120,
1297
+ "precision": "fp32"
1298
+ },
1299
+ "source": {
1300
+ "file": "gXA_close8_s8_register.safetensors",
1301
+ "sha256": "19b02efab5fb482cd50d10508cda56868d8bb1f7fce835a0b7e9babec182aa3b"
1302
+ },
1303
+ "recipe": "trained as an always-on group on the frozen final core by continuing the four-arm group's run. The four arms of stages 1 to 4 were loaded as trained (see the group stages-1-4) and held fixed and on. Then a new arm for stage 5 was attached over them and trained for 4,000 steps in their presence; then, with it also held fixed, a new arm for stage 6 the same way, then stage 7, then stage 8. Each step gives the training arm 2 x 4096 bytes of its own stage's text with every attached arm on, and a quiet term at weight 2 on 2 rows that are not its text (one of web text, one of the stage text of another arm of the eight, drawn from every other stage whether or not its arm was attached yet; the mixed stage's text, which contains the other stages' kinds, is drawn as no arm's partner row): the KL from the model with that arm switched off to the model with every arm on. Pure Adam, lr 0.001",
1304
+ "title": "Stage 8: register",
1305
+ "behavior": "the same content in different registers of speech",
1306
+ "family": "stage",
1307
+ "status": "a member of the group stages-1-8; trained always on with its partners, so alone is not its trained state",
1308
+ "score": "alone: own stage text 0.491 -> 0.068 bpb; web text +0.0005 bpb",
1309
+ "measured": {
1310
+ "own_stage_bpb_arms_off": 0.4907,
1311
+ "own_stage_bpb_group_on": 0.0212,
1312
+ "own_stage_bpb_this_arm_alone": 0.0678,
1313
+ "share_of_the_group_on_own_stage": -0.3546,
1314
+ "share_of_the_group_on_own_stage_se": 0.0066,
1315
+ "share_of_the_group_on_web_text": 0.00045,
1316
+ "alone_on_web_text": 0.00048,
1317
+ "share_of_the_group_on_partner_stages": {
1318
+ "s1_perspective": 0.0007,
1319
+ "s2_concept": -0.0008,
1320
+ "s3_rules": 0.0005,
1321
+ "s4_arith": 0.0001,
1322
+ "s5_causal": 0.0002,
1323
+ "s6_tryfail": 0.0002,
1324
+ "s7_mixed": 0.0012
1325
+ },
1326
+ "alone_on_partner_stages": {
1327
+ "s1_perspective": -0.0172,
1328
+ "s2_concept": -0.0079,
1329
+ "s3_rules": -0.0048,
1330
+ "s4_arith": -0.0065,
1331
+ "s5_causal": -0.0189,
1332
+ "s6_tryfail": -0.022,
1333
+ "s7_mixed": -0.0002
1334
+ },
1335
+ "steps_trained": 4000,
1336
+ "second_seed": {
1337
+ "own_stage_bpb_arms_off": 0.4907,
1338
+ "own_stage_bpb_group_on": 0.0208,
1339
+ "own_stage_bpb_this_arm_alone": 0.0516,
1340
+ "share_of_the_group_on_own_stage": -0.3573,
1341
+ "share_of_the_group_on_own_stage_se": 0.0067,
1342
+ "share_of_the_group_on_web_text": 0.00046,
1343
+ "alone_on_web_text": 5e-05,
1344
+ "share_of_the_group_on_partner_stages": {
1345
+ "s1_perspective": 0.0001,
1346
+ "s2_concept": 0.0001,
1347
+ "s3_rules": 0.0005,
1348
+ "s4_arith": 0.0001,
1349
+ "s5_causal": 0.0002,
1350
+ "s6_tryfail": 0.0001,
1351
+ "s7_mixed": 0.0002
1352
+ },
1353
+ "alone_on_partner_stages": {
1354
+ "s1_perspective": -0.0045,
1355
+ "s2_concept": -0.005,
1356
+ "s3_rules": -0.0063,
1357
+ "s4_arith": -0.0064,
1358
+ "s5_causal": -0.0139,
1359
+ "s6_tryfail": -0.0187,
1360
+ "s7_mixed": 0.0001
1361
+ },
1362
+ "steps_trained": 4000
1363
+ }
1364
+ },
1365
+ "examples": [
1366
+ "The word 'ancient' means from a very long time ago. Used in a sentence: a ancient is easy to point at once you know the word."
1367
+ ]
1368
+ },
1369
+ {
1370
+ "id": "stages-1-8",
1371
+ "name": "stages-1-8",
1372
+ "base_model_id": "alephllm/mini-beatrix-3@step245674",
1373
+ "step": 245674,
1374
+ "params": 109543680,
1375
+ "sites": 32,
1376
+ "weights": [
1377
+ "arms/stages-1-8/s1_perspective.safetensors",
1378
+ "arms/stages-1-8/s2_concept.safetensors",
1379
+ "arms/stages-1-8/s3_rules.safetensors",
1380
+ "arms/stages-1-8/s4_arith.safetensors",
1381
+ "arms/stages-1-8/s5_causal.safetensors",
1382
+ "arms/stages-1-8/s6_tryfail.safetensors",
1383
+ "arms/stages-1-8/s7_mixed.safetensors",
1384
+ "arms/stages-1-8/s8_register.safetensors"
1385
+ ],
1386
+ "mount": {
1387
+ "type": "stack",
1388
+ "members": [
1389
+ "stages-1-8/s1_perspective",
1390
+ "stages-1-8/s2_concept",
1391
+ "stages-1-8/s3_rules",
1392
+ "stages-1-8/s4_arith",
1393
+ "stages-1-8/s5_causal",
1394
+ "stages-1-8/s6_tryfail",
1395
+ "stages-1-8/s7_mixed",
1396
+ "stages-1-8/s8_register"
1397
+ ]
1398
+ },
1399
+ "precision": "fp32",
1400
+ "template": {
1401
+ "frame": "raw",
1402
+ "template": "raw text, continued"
1403
+ },
1404
+ "decode": {
1405
+ "temperature": 0.0,
1406
+ "top_p": 1.0,
1407
+ "max_new": 120,
1408
+ "precision": "fp32"
1409
+ },
1410
+ "recipe": "trained as an always-on group on the frozen final core by continuing the four-arm group's run. The four arms of stages 1 to 4 were loaded as trained (see the group stages-1-4) and held fixed and on. Then a new arm for stage 5 was attached over them and trained for 4,000 steps in their presence; then, with it also held fixed, a new arm for stage 6 the same way, then stage 7, then stage 8. Each step gives the training arm 2 x 4096 bytes of its own stage's text with every attached arm on, and a quiet term at weight 2 on 2 rows that are not its text (one of web text, one of the stage text of another arm of the eight, drawn from every other stage whether or not its arm was attached yet; the mixed stage's text, which contains the other stages' kinds, is drawn as no arm's partner row): the KL from the model with that arm switched off to the model with every arm on. Pure Adam, lr 0.001",
1411
+ "title": "Stages 1 to 8, on together",
1412
+ "behavior": "the arms of stage 1: perspective, stage 2: concepts, stage 3: rule chains, stage 4: arithmetic, stage 5: cause and effect, stage 6: try, fail, revise, stage 7: mixed, stage 8: register, always on together, as they were trained",
1413
+ "family": "group",
1414
+ "status": "two seeds",
1415
+ "score": "stage text, arms off -> on: s1_perspective 0.554 -> 0.049; s2_concept 0.721 -> 0.111; s3_rules 0.930 -> 0.224; s4_arith 1.234 -> 0.328; s5_causal 0.660 -> 0.043; s6_tryfail 0.673 -> 0.055; s7_mixed 0.783 -> 0.720; s8_register 0.491 -> 0.021; web text +0.0019 bpb",
1416
+ "measured": {
1417
+ "stage_text_bpb": {
1418
+ "s1_perspective": {
1419
+ "arms_off": 0.5538,
1420
+ "group_on": 0.0492,
1421
+ "difference": -0.5046,
1422
+ "difference_se": 0.0048
1423
+ },
1424
+ "s2_concept": {
1425
+ "arms_off": 0.7212,
1426
+ "group_on": 0.1105,
1427
+ "difference": -0.6107,
1428
+ "difference_se": 0.0038
1429
+ },
1430
+ "s3_rules": {
1431
+ "arms_off": 0.9296,
1432
+ "group_on": 0.2237,
1433
+ "difference": -0.7059,
1434
+ "difference_se": 0.0053
1435
+ },
1436
+ "s4_arith": {
1437
+ "arms_off": 1.2339,
1438
+ "group_on": 0.3279,
1439
+ "difference": -0.9059,
1440
+ "difference_se": 0.0062
1441
+ },
1442
+ "s5_causal": {
1443
+ "arms_off": 0.6604,
1444
+ "group_on": 0.0431,
1445
+ "difference": -0.6174,
1446
+ "difference_se": 0.0069
1447
+ },
1448
+ "s6_tryfail": {
1449
+ "arms_off": 0.6728,
1450
+ "group_on": 0.0548,
1451
+ "difference": -0.618,
1452
+ "difference_se": 0.0037
1453
+ },
1454
+ "s7_mixed": {
1455
+ "arms_off": 0.783,
1456
+ "group_on": 0.72,
1457
+ "difference": -0.063,
1458
+ "difference_se": 0.039
1459
+ },
1460
+ "s8_register": {
1461
+ "arms_off": 0.4907,
1462
+ "group_on": 0.0212,
1463
+ "difference": -0.4695,
1464
+ "difference_se": 0.0085
1465
+ }
1466
+ },
1467
+ "web_text_difference": 0.00189,
1468
+ "web_text_difference_se": 0.00032,
1469
+ "steps_trained": {
1470
+ "s1_perspective": 3200,
1471
+ "s2_concept": 2400,
1472
+ "s3_rules": 800,
1473
+ "s4_arith": 800,
1474
+ "s5_causal": 4000,
1475
+ "s6_tryfail": 4000,
1476
+ "s7_mixed": 4000,
1477
+ "s8_register": 4000
1478
+ },
1479
+ "probe_suite_mean": {
1480
+ "arms_off": 0.411,
1481
+ "group_on": 0.422
1482
+ },
1483
+ "second_seed": {
1484
+ "stage_text_bpb": {
1485
+ "s1_perspective": {
1486
+ "arms_off": 0.5538,
1487
+ "group_on": 0.0484,
1488
+ "difference": -0.5055,
1489
+ "difference_se": 0.0048
1490
+ },
1491
+ "s2_concept": {
1492
+ "arms_off": 0.7212,
1493
+ "group_on": 0.1212,
1494
+ "difference": -0.6,
1495
+ "difference_se": 0.0038
1496
+ },
1497
+ "s3_rules": {
1498
+ "arms_off": 0.9296,
1499
+ "group_on": 0.2222,
1500
+ "difference": -0.7074,
1501
+ "difference_se": 0.0055
1502
+ },
1503
+ "s4_arith": {
1504
+ "arms_off": 1.2339,
1505
+ "group_on": 0.3266,
1506
+ "difference": -0.9073,
1507
+ "difference_se": 0.0062
1508
+ },
1509
+ "s5_causal": {
1510
+ "arms_off": 0.6604,
1511
+ "group_on": 0.0428,
1512
+ "difference": -0.6176,
1513
+ "difference_se": 0.0069
1514
+ },
1515
+ "s6_tryfail": {
1516
+ "arms_off": 0.6728,
1517
+ "group_on": 0.0543,
1518
+ "difference": -0.6185,
1519
+ "difference_se": 0.0036
1520
+ },
1521
+ "s7_mixed": {
1522
+ "arms_off": 0.783,
1523
+ "group_on": 0.7198,
1524
+ "difference": -0.0632,
1525
+ "difference_se": 0.039
1526
+ },
1527
+ "s8_register": {
1528
+ "arms_off": 0.4907,
1529
+ "group_on": 0.0208,
1530
+ "difference": -0.4699,
1531
+ "difference_se": 0.0086
1532
+ }
1533
+ },
1534
+ "web_text_difference": 0.00103,
1535
+ "web_text_difference_se": 0.0002,
1536
+ "steps_trained": {
1537
+ "s1_perspective": 3200,
1538
+ "s2_concept": 2400,
1539
+ "s3_rules": 800,
1540
+ "s4_arith": 800,
1541
+ "s5_causal": 4000,
1542
+ "s6_tryfail": 4000,
1543
+ "s7_mixed": 4000,
1544
+ "s8_register": 4000
1545
+ },
1546
+ "probe_suite_mean": {
1547
+ "arms_off": 0.411,
1548
+ "group_on": 0.43
1549
+ }
1550
+ }
1551
+ },
1552
+ "examples": []
1553
+ },
1554
  {
1555
  "id": "solo/s1_perspective",
1556
  "name": "solo/s1_perspective",
 
1874
  "examples": [
1875
  "Count by 3s: 1 3s make 3; 2 3s make 6; 3 3s make 9; 4 3s make 12; 5 3s make 15. So 3 times 5 is 15."
1876
  ]
1877
+ },
1878
+ {
1879
+ "id": "solo/s5_causal",
1880
+ "name": "solo/s5_causal",
1881
+ "group": null,
1882
+ "base_model_id": "alephllm/mini-beatrix-3@step245674",
1883
+ "step": 245674,
1884
+ "params": 13692960,
1885
+ "sites": 32,
1886
+ "weights": [
1887
+ "arms/solo/s5_causal.safetensors"
1888
+ ],
1889
+ "mount": {
1890
+ "type": "anchor"
1891
+ },
1892
+ "precision": "fp32",
1893
+ "template": {
1894
+ "frame": "raw",
1895
+ "template": "raw stage text, continued"
1896
+ },
1897
+ "decode": {
1898
+ "temperature": 0.0,
1899
+ "top_p": 1.0,
1900
+ "max_new": 120,
1901
+ "precision": "fp32"
1902
+ },
1903
+ "source": {
1904
+ "file": "s5_causal_fresh_seed5_d20261005_step800.safetensors",
1905
+ "sha256": "b4369a46c13a3c2b425255e030a9c2b851fa3cf50046394dcadfd4f88fdc2a6a"
1906
+ },
1907
+ "recipe": "trained alone on the frozen final core: 800 steps of 2 x 4096 bytes of the stage's own text, pure Adam (lr 0.001), with a quiet term at weight 2 on held-out web text (the KL from the model with the arm off to the model with it on, one chunk beside every task chunk)",
1908
+ "title": "Stage 5: cause and effect (solo)",
1909
+ "behavior": "short cause-and-effect text: what happened and why",
1910
+ "family": "solo",
1911
+ "status": "one seed; trained alone: mount one solo arm at a time",
1912
+ "score": "own stage text 0.660 -> 0.068 bpb; web text +0.0003 bpb",
1913
+ "measured": {
1914
+ "steps": 800,
1915
+ "own_stage_bpb_arms_off": 0.6604,
1916
+ "own_stage_bpb_arm_on": 0.0681,
1917
+ "own_stage_difference": -0.5924,
1918
+ "own_stage_difference_se": 0.0068,
1919
+ "web_text_difference": 0.00032,
1920
+ "web_text_difference_se": 0.00017,
1921
+ "probe_suite_mean": {
1922
+ "arms_off": 0.411,
1923
+ "arm_on": 0.407
1924
+ }
1925
+ },
1926
+ "examples": [
1927
+ "Noor stacked the cups too high. Because of that, the tower leaned. Because of that, the cups crashed down."
1928
+ ]
1929
+ },
1930
+ {
1931
+ "id": "solo/s6_tryfail",
1932
+ "name": "solo/s6_tryfail",
1933
+ "group": null,
1934
+ "base_model_id": "alephllm/mini-beatrix-3@step245674",
1935
+ "step": 245674,
1936
+ "params": 13692960,
1937
+ "sites": 32,
1938
+ "weights": [
1939
+ "arms/solo/s6_tryfail.safetensors"
1940
+ ],
1941
+ "mount": {
1942
+ "type": "anchor"
1943
+ },
1944
+ "precision": "fp32",
1945
+ "template": {
1946
+ "frame": "raw",
1947
+ "template": "raw stage text, continued"
1948
+ },
1949
+ "decode": {
1950
+ "temperature": 0.0,
1951
+ "top_p": 1.0,
1952
+ "max_new": 120,
1953
+ "precision": "fp32"
1954
+ },
1955
+ "source": {
1956
+ "file": "s6_tryfail_fresh_seed6_d20261005_step800.safetensors",
1957
+ "sha256": "35db26313e181e40be63f5fca7cba883dd25cc88fd9c71e150d1a67363959bbf"
1958
+ },
1959
+ "recipe": "trained alone on the frozen final core: 800 steps of 2 x 4096 bytes of the stage's own text, pure Adam (lr 0.001), with a quiet term at weight 2 on held-out web text (the KL from the model with the arm off to the model with it on, one chunk beside every task chunk)",
1960
+ "title": "Stage 6: try, fail, revise (solo)",
1961
+ "behavior": "an attempt, its failure and the revised attempt",
1962
+ "family": "solo",
1963
+ "status": "one seed; trained alone: mount one solo arm at a time",
1964
+ "score": "own stage text 0.673 -> 0.114 bpb; web text +0.0008 bpb",
1965
+ "measured": {
1966
+ "steps": 800,
1967
+ "own_stage_bpb_arms_off": 0.6728,
1968
+ "own_stage_bpb_arm_on": 0.1142,
1969
+ "own_stage_difference": -0.5587,
1970
+ "own_stage_difference_se": 0.0045,
1971
+ "web_text_difference": 0.00079,
1972
+ "web_text_difference_se": 0.00016,
1973
+ "probe_suite_mean": {
1974
+ "arms_off": 0.411,
1975
+ "arm_on": 0.4
1976
+ }
1977
+ },
1978
+ "examples": [
1979
+ "Talia adds 34 and 23 and writes 65. Talia checks by counting back: 65 is too big. Talia tries again carefully: 34 + 23 = 57. The check works now, and Talia keeps the good method."
1980
+ ]
1981
+ },
1982
+ {
1983
+ "id": "solo/s7_mixed",
1984
+ "name": "solo/s7_mixed",
1985
+ "group": null,
1986
+ "base_model_id": "alephllm/mini-beatrix-3@step245674",
1987
+ "step": 245674,
1988
+ "params": 13692960,
1989
+ "sites": 32,
1990
+ "weights": [
1991
+ "arms/solo/s7_mixed.safetensors"
1992
+ ],
1993
+ "mount": {
1994
+ "type": "anchor"
1995
+ },
1996
+ "precision": "fp32",
1997
+ "template": {
1998
+ "frame": "raw",
1999
+ "template": "raw stage text, continued"
2000
+ },
2001
+ "decode": {
2002
+ "temperature": 0.0,
2003
+ "top_p": 1.0,
2004
+ "max_new": 120,
2005
+ "precision": "fp32"
2006
+ },
2007
+ "source": {
2008
+ "file": "s7_mixed_fresh_seed7_d20261005_step800.safetensors",
2009
+ "sha256": "6dc2c0c2a436a6052d8e82180a17248c04f8ef55822b5374f6f12f555268c412"
2010
+ },
2011
+ "recipe": "trained alone on the frozen final core: 800 steps of 2 x 4096 bytes of the stage's own text, pure Adam (lr 0.001), with a quiet term at weight 2 on held-out web text (the KL from the model with the arm off to the model with it on, one chunk beside every task chunk)",
2012
+ "title": "Stage 7: mixed (solo)",
2013
+ "behavior": "the mixed stage's diet: the earlier stages' text beside ordinary prose",
2014
+ "family": "solo",
2015
+ "status": "one seed; trained alone: mount one solo arm at a time",
2016
+ "score": "own stage text 0.783 -> 0.755 bpb; web text +0.0004 bpb",
2017
+ "measured": {
2018
+ "steps": 800,
2019
+ "own_stage_bpb_arms_off": 0.783,
2020
+ "own_stage_bpb_arm_on": 0.755,
2021
+ "own_stage_difference": -0.028,
2022
+ "own_stage_difference_se": 0.018,
2023
+ "web_text_difference": 0.00039,
2024
+ "web_text_difference_se": 0.00013,
2025
+ "probe_suite_mean": {
2026
+ "arms_off": 0.411,
2027
+ "arm_on": 0.378
2028
+ }
2029
+ },
2030
+ "examples": [
2031
+ "It has been compared to a chant, a rhythmic divine beauty, a melody, an aria, a toccata, an edification, an exaltation. As poetry is for the tongue, calligraphy is to the page. The authors of The Splendor of Islamic Calligraphy put it best when they said: “Calligraphy is the plainsong of the divine”.\nCalligraphy is the art of the linear graphic, but it is more than that. In Isl"
2032
+ ]
2033
+ },
2034
+ {
2035
+ "id": "solo/s8_register",
2036
+ "name": "solo/s8_register",
2037
+ "group": null,
2038
+ "base_model_id": "alephllm/mini-beatrix-3@step245674",
2039
+ "step": 245674,
2040
+ "params": 13692960,
2041
+ "sites": 32,
2042
+ "weights": [
2043
+ "arms/solo/s8_register.safetensors"
2044
+ ],
2045
+ "mount": {
2046
+ "type": "anchor"
2047
+ },
2048
+ "precision": "fp32",
2049
+ "template": {
2050
+ "frame": "raw",
2051
+ "template": "raw stage text, continued"
2052
+ },
2053
+ "decode": {
2054
+ "temperature": 0.0,
2055
+ "top_p": 1.0,
2056
+ "max_new": 120,
2057
+ "precision": "fp32"
2058
+ },
2059
+ "source": {
2060
+ "file": "s8_register_fresh_seed8_d20261005_step800.safetensors",
2061
+ "sha256": "a5ae8506d4752cf9d619e789e396f2b379317e20316ad32a02f5ee0af541b4c1"
2062
+ },
2063
+ "recipe": "trained alone on the frozen final core: 800 steps of 2 x 4096 bytes of the stage's own text, pure Adam (lr 0.001), with a quiet term at weight 2 on held-out web text (the KL from the model with the arm off to the model with it on, one chunk beside every task chunk)",
2064
+ "title": "Stage 8: register (solo)",
2065
+ "behavior": "the same content in different registers of speech",
2066
+ "family": "solo",
2067
+ "status": "one seed; trained alone: mount one solo arm at a time",
2068
+ "score": "own stage text 0.491 -> 0.051 bpb; web text +0.0008 bpb",
2069
+ "measured": {
2070
+ "steps": 800,
2071
+ "own_stage_bpb_arms_off": 0.4907,
2072
+ "own_stage_bpb_arm_on": 0.051,
2073
+ "own_stage_difference": -0.4396,
2074
+ "own_stage_difference_se": 0.0078,
2075
+ "web_text_difference": 0.00075,
2076
+ "web_text_difference_se": 9e-05,
2077
+ "probe_suite_mean": {
2078
+ "arms_off": 0.411,
2079
+ "arm_on": 0.378
2080
+ }
2081
+ },
2082
+ "examples": [
2083
+ "The word 'ancient' means from a very long time ago. Used in a sentence: a ancient is easy to point at once you know the word."
2084
+ ]
2085
  }
2086
  ]
2087
  }
arms/solo/s3_rules.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:d26218539c5f1328f30a2ce40c4d5ccc5cf01a64d03c50469183ad64c565553c
3
  size 54817680
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e068430f1da2f0de1d6f703ccf552ee9633ad5ad02754fef90d4fa924cba9144
3
  size 54817680
arms/solo/s4_arith.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:358c996148f9bb8c785f6fc2ec565ca5f08ba2cb252cb85b3ae388655e5626ed
3
  size 54817552
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:17f398b0d0100e48cd8ff5c3aba10c9ee2d57f769c94914d40c6bc1f396396b0
3
  size 54817552
arms/solo/s5_causal.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b871df2812ce5bdbd61df53e0c4000543ded6685c198e436ddffb5e76dea58f7
3
+ size 54816968
arms/solo/s6_tryfail.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e53255c4a1748cc950da6c577ace85edd7650170c88acca019c00d2d31395e5a
3
+ size 54817040
arms/solo/s7_mixed.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:310baca6f2ab7fbfb3e1070e7c233871d9f5452c83d0062432b0d27ca4b3504f
3
+ size 54817256
arms/solo/s8_register.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4ce9d6a1bb0aa19b8ad8b20deaa4f036d74792f0d62af8c4a96e3dbd84091b02
3
+ size 54816984
arms/stages-1-4/s1_perspective.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:656d03203726ea77aa758a636424db6f21a39e379c34697fa10debbe017c48f9
3
  size 54819000
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:bcf835c13de8bd0b446e65f21aad9515c2cab769b701936165c143815091b321
3
  size 54819000
arms/stages-1-4/s2_concept.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:436ce53452aa7264ef3fb5fea28b6a7077b843c4e65f1538932114b99fe380d5
3
  size 54818880
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:3aa85cc0047416499b5f00288875803899dec2ec897d4f834b307ff64cdc6fb6
3
  size 54818880
arms/stages-1-4/s3_rules.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:c76ee31256d3c36597b7791ab0deab31507d625ff68ccdac10f28a5c4a007a07
3
  size 54818928
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e409aa5684064ff03c54bb6136ffde682df5b067cf191715ab95e44748a120c3
3
  size 54818928
arms/stages-1-4/s4_arith.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:9c6080918b0ea2b107fa8392901cc26628934c1d8ec7c095db6cd1f86b212cc8
3
  size 54818792
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a85c82a5d02951a81882146d2df38523e4bddd7a07910eaeba23a99ae86c1ce5
3
  size 54818792
arms/stages-1-8/s1_perspective.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:7faa377430e9be0b4600683ea4c63ecc8a3c84e7d536f3df830324065b76f97d
3
+ size 54819424
arms/stages-1-8/s2_concept.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e305183f35aa6adb76b8bee4c57d802f377acd2b0ad92b1f7578cd2de8618ad5
3
+ size 54819304
arms/stages-1-8/s3_rules.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6069cac6f21653ab9da229b2daf25f2271633603fb60bf99f04bfc3acfc73ee9
3
+ size 54819344
arms/stages-1-8/s4_arith.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6e29b2fcc48ae351a519a783bed637f282597c16bddb903f412948dd89ac4102
3
+ size 54819216
arms/stages-1-8/s5_causal.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:00e6374b72a19253a3130001daf4c67fddc761497ae519e8d3d4c0a38eac56e1
3
+ size 54818832
arms/stages-1-8/s6_tryfail.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:da74a8218506923537757416610e94f075d72615a30aaaba35f3bc5e9dedffd8
3
+ size 54818896
arms/stages-1-8/s7_mixed.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b54beaf81c1f6f31dbc61ed312cf1e36a068f334f222da8c6152208ae741c9b2
3
+ size 54819136
arms/stages-1-8/s8_register.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c55b64e9c6a8084cadafca177fff3b1079c0747a4cf8542a2037e4fe18eb6ff0
3
+ size 54818840