mini-beatrix-3: the stage arms (arms off and arms on in one package)
Browse files- README.md +55 -17
- arms/index.json +1235 -2
- arms/solo/s3_rules.safetensors +1 -1
- arms/solo/s4_arith.safetensors +1 -1
- arms/solo/s5_causal.safetensors +3 -0
- arms/solo/s6_tryfail.safetensors +3 -0
- arms/solo/s7_mixed.safetensors +3 -0
- arms/solo/s8_register.safetensors +3 -0
- arms/stages-1-4/s1_perspective.safetensors +1 -1
- arms/stages-1-4/s2_concept.safetensors +1 -1
- arms/stages-1-4/s3_rules.safetensors +1 -1
- arms/stages-1-4/s4_arith.safetensors +1 -1
- arms/stages-1-8/s1_perspective.safetensors +3 -0
- arms/stages-1-8/s2_concept.safetensors +3 -0
- arms/stages-1-8/s3_rules.safetensors +3 -0
- arms/stages-1-8/s4_arith.safetensors +3 -0
- arms/stages-1-8/s5_causal.safetensors +3 -0
- arms/stages-1-8/s6_tryfail.safetensors +3 -0
- arms/stages-1-8/s7_mixed.safetensors +3 -0
- arms/stages-1-8/s8_register.safetensors +3 -0
README.md
CHANGED
|
@@ -23,10 +23,12 @@ completed 2026-10-05.
|
|
| 23 |
|
| 24 |
The package ships **both states in one repository**: the bare model (arms
|
| 25 |
off) and the model with its **stage arms** mounted (arms on). An arm is a
|
| 26 |
-
small detachable adapter trained on this exact frozen core. The arms
|
| 27 |
-
**
|
| 28 |
-
together.
|
| 29 |
-
|
|
|
|
|
|
|
| 30 |
weights are the same files either way, and detaching restores the bare
|
| 31 |
model bit for bit.
|
| 32 |
|
|
@@ -39,7 +41,7 @@ m = AutoModelForCausalLM.from_pretrained(
|
|
| 39 |
|
| 40 |
m.say("Who are you?") # arms off: the bare core, its own chat frame
|
| 41 |
|
| 42 |
-
m.mount_arm("stages-1-
|
| 43 |
m.detach_arm() # the bare core again, verified bit-exact
|
| 44 |
```
|
| 45 |
|
|
@@ -48,7 +50,7 @@ Arms on from the first call:
|
|
| 48 |
```python
|
| 49 |
m = AutoModelForCausalLM.from_pretrained(
|
| 50 |
"AbstractPhil/mini-beatrix-3", trust_remote_code=True,
|
| 51 |
-
default_arm="stages-1-
|
| 52 |
```
|
| 53 |
|
| 54 |
The model reads **raw UTF-8 bytes**: `input_ids` are byte values 0–255.
|
|
@@ -80,9 +82,9 @@ There is no tokenizer to download.
|
|
| 80 |
The curriculum taught the model nine kinds of text, one stage after another,
|
| 81 |
and the finished model moved on from each stage's text as the later stages
|
| 82 |
came. The stage arms bring that text back without touching the core. They
|
| 83 |
-
were fitted on the finished, frozen core **as
|
| 84 |
-
`m.arms()` returns the tables below as data, with each row's
|
| 85 |
-
measurements.
|
| 86 |
|
| 87 |
**The group `stages-1-4`** (4 arms, 54.8M parameters; two seeds):
|
| 88 |
|
|
@@ -108,6 +110,38 @@ Each member mounted alone, and the whole group, on every stage's held-out text (
|
|
| 108 |
| only `stages-1-4/s4_arith` | 0.553 | 0.721 | 0.928 | 1.127 | 0.951 |
|
| 109 |
| the whole group `stages-1-4` | 0.048 | 0.112 | 0.224 | 0.338 | 0.952 |
|
| 110 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 111 |
**How to read the tables.** *Stage text* is the loss, in bits per byte, on
|
| 112 |
held-out text of the arm's own stage, with arms off and with the whole group
|
| 113 |
on. *This arm alone* is the same loss with only that one arm mounted. *Stage
|
|
@@ -116,26 +150,26 @@ correctly, first in the form the stage text uses and then in a form it does
|
|
| 116 |
not. The *web text* figure is the change on held-out ordinary web text
|
| 117 |
(fineweb-edu) with the group on: the arms were trained to leave it alone.
|
| 118 |
|
| 119 |
-
**
|
| 120 |
|
| 121 |
-
**What the arms do, plainly.**
|
| 122 |
-
text. The gain is in each stage's own forms. Whether it carries over to
|
| 123 |
-
forms is shown in
|
| 124 |
|
| 125 |
### Arms off and arms on
|
| 126 |
|
| 127 |
```python
|
| 128 |
m.arm # None: arms off
|
| 129 |
-
h = m.mount_arm("stages-1-
|
| 130 |
with h.all_off(): # every member masked: the bare core's logits
|
| 131 |
m.say("Hello there.")
|
| 132 |
m.detach_arm(verify=True) # raises if the restored core is not bit-exact
|
| 133 |
m.mount_arm("stages-1-4/s1_perspective") # one member alone (the "alone" readings)
|
| 134 |
```
|
| 135 |
|
| 136 |
-
|
| 137 |
after another at every block, with no mixer between them. `default_arm` in
|
| 138 |
-
`from_pretrained` (or in `config.json`) mounts
|
| 139 |
`save_pretrained` refuses while an arm is mounted, so an arm can never be
|
| 140 |
written into the core's weight file.
|
| 141 |
|
|
@@ -151,6 +185,10 @@ whole stage by itself.
|
|
| 151 |
| `solo/s2_concept` | kinds, properties and differences: what a thing is a kind of, what its kind can do, how two things differ | 0.721 → 0.120 (a second seed: 0.120) | +0.0003 | 30% → 84% | 68% → 73% | two |
|
| 152 |
| `solo/s3_rules` | if-then rules over made-up words, followed step by step to what follows | 0.930 → 0.372 (a second seed: 0.374) | +0.0009 | 16% → 60% | 7% → 12% | two |
|
| 153 |
| `solo/s4_arith` | small arithmetic worked out in text | 1.234 → 0.389 (a second seed: 0.388) | +0.0007 | 34% → 49% | 24% → 24% | two |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 154 |
|
| 155 |
```python
|
| 156 |
m.mount_arm("solo/s1_perspective") # one solo arm; mounting another swaps it
|
|
@@ -258,7 +296,7 @@ trunk for each stage; those arms stayed nearly empty, because the trunk took
|
|
| 258 |
each stage's text in before its arm could, and they were detached at step
|
| 259 |
148,000 (their files are in the training repo under `mini-beatrix-3/arms/`,
|
| 260 |
bound to those earlier trunk states). The arms in this repository were fitted
|
| 261 |
-
afterwards on the finished, frozen core: as
|
| 262 |
|
| 263 |
## Lineage
|
| 264 |
|
|
|
|
| 23 |
|
| 24 |
The package ships **both states in one repository**: the bare model (arms
|
| 25 |
off) and the model with its **stage arms** mounted (arms on). An arm is a
|
| 26 |
+
small detachable adapter trained on this exact frozen core. The arms come as
|
| 27 |
+
**groups**: arms that were trained switched on together and are mounted
|
| 28 |
+
together. There are two, the four arms of stages 1 to 4 and the eight arms
|
| 29 |
+
of stages 1 to 8 (those four, held fixed, with four more fitted over them).
|
| 30 |
+
Each stage also has a **solo arm**, trained alone, for use one at a time.
|
| 31 |
+
Arms are guests on the core, never a change to it: the core's
|
| 32 |
weights are the same files either way, and detaching restores the bare
|
| 33 |
model bit for bit.
|
| 34 |
|
|
|
|
| 41 |
|
| 42 |
m.say("Who are you?") # arms off: the bare core, its own chat frame
|
| 43 |
|
| 44 |
+
m.mount_arm("stages-1-8") # arms on: the stage arms, all on together
|
| 45 |
m.detach_arm() # the bare core again, verified bit-exact
|
| 46 |
```
|
| 47 |
|
|
|
|
| 50 |
```python
|
| 51 |
m = AutoModelForCausalLM.from_pretrained(
|
| 52 |
"AbstractPhil/mini-beatrix-3", trust_remote_code=True,
|
| 53 |
+
default_arm="stages-1-8").eval()
|
| 54 |
```
|
| 55 |
|
| 56 |
The model reads **raw UTF-8 bytes**: `input_ids` are byte values 0–255.
|
|
|
|
| 82 |
The curriculum taught the model nine kinds of text, one stage after another,
|
| 83 |
and the finished model moved on from each stage's text as the later stages
|
| 84 |
came. The stage arms bring that text back without touching the core. They
|
| 85 |
+
were fitted on the finished, frozen core **as always-on groups**:
|
| 86 |
+
the arms of stages 1 and 2 were trained first and then held fixed, the arms of stages 3 and 4 were attached over them one after the other and trained in their presence, and then, with all four held fixed and on, the arms of stages 5, 6, 7 and 8 were attached over them the same way, one after the other, each for 4,000 steps and each kept quiet on every other arm's stage text. The eight-arm group's first four arms are the four-arm group unchanged; the four-arm group is also packaged on its own. `m.arms()` returns the tables below as data, with each row's
|
| 87 |
+
recipe and full measurements.
|
| 88 |
|
| 89 |
**The group `stages-1-4`** (4 arms, 54.8M parameters; two seeds):
|
| 90 |
|
|
|
|
| 110 |
| only `stages-1-4/s4_arith` | 0.553 | 0.721 | 0.928 | 1.127 | 0.951 |
|
| 111 |
| the whole group `stages-1-4` | 0.048 | 0.112 | 0.224 | 0.338 | 0.952 |
|
| 112 |
|
| 113 |
+
**The group `stages-1-8`** (8 arms, 109.5M parameters; two seeds):
|
| 114 |
+
|
| 115 |
+
| arm | its stage's text | stage text, arms off → the group on (bpb) | this arm alone (bpb) | stage items in the stage's own form, off → group on | stage items in a new form, off → group on |
|
| 116 |
+
|---|---|---|---|---|---|
|
| 117 |
+
| `stages-1-8/s1_perspective` | one small event retold from each side: seen from a person, said to them, told about them | 0.554 → 0.049 | 0.052 | 55% → 89% | 46% → 44% |
|
| 118 |
+
| `stages-1-8/s2_concept` | kinds, properties and differences: what a thing is a kind of, what its kind can do, how two things differ | 0.721 → 0.111 | 0.350 | 30% → 98% | 68% → 72% |
|
| 119 |
+
| `stages-1-8/s3_rules` | if-then rules over made-up words, followed step by step to what follows | 0.930 → 0.224 | 0.824 | 16% → 86% | 7% → 31% |
|
| 120 |
+
| `stages-1-8/s4_arith` | small arithmetic worked out in text | 1.234 → 0.328 | 1.127 | 34% → 64% | 24% → 24% |
|
| 121 |
+
| `stages-1-8/s5_causal` | short cause-and-effect text: what happened and why | 0.660 → 0.043 | 0.091 | not measured | not measured |
|
| 122 |
+
| `stages-1-8/s6_tryfail` | an attempt, its failure and the revised attempt | 0.673 → 0.055 | 0.164 | not measured | not measured |
|
| 123 |
+
| `stages-1-8/s7_mixed` | the mixed stage's diet: the earlier stages' text beside ordinary prose | 0.783 → 0.720 | 0.724 | not measured | not measured |
|
| 124 |
+
| `stages-1-8/s8_register` | the same content in different registers of speech | 0.491 → 0.021 | 0.068 | not measured | not measured |
|
| 125 |
+
|
| 126 |
+
With the group on, held-out web text moves by +0.0019 bpb (the limit set for it was +0.012), and the nine-suite probe mean reads 0.411 with arms off and 0.422 with the group on.
|
| 127 |
+
|
| 128 |
+
The numbers are those of the packaged weights. A second run of the same recipe from other seeds read the stage texts with its group on at 0.048, 0.121, 0.222, 0.327, 0.043, 0.054, 0.720, 0.021 bpb and moved web text by +0.0010.
|
| 129 |
+
|
| 130 |
+
Each member mounted alone, and the whole group, on every stage's held-out text (bpb):
|
| 131 |
+
|
| 132 |
+
| mounted | stage 1 text | stage 2 text | stage 3 text | stage 4 text | stage 5 text | stage 6 text | stage 7 text | stage 8 text | web text |
|
| 133 |
+
|---|---|---|---|---|---|---|---|---|---|
|
| 134 |
+
| nothing (arms off) | 0.554 | 0.721 | 0.930 | 1.234 | 0.660 | 0.673 | 0.783 | 0.491 | 0.951 |
|
| 135 |
+
| only `stages-1-8/s1_perspective` | 0.053 | 0.239 | 0.509 | 0.723 | 0.463 | 0.495 | 0.773 | 0.379 | 0.952 |
|
| 136 |
+
| only `stages-1-8/s2_concept` | 0.527 | 0.350 | 0.619 | 0.880 | 0.659 | 0.673 | 0.783 | 0.469 | 0.952 |
|
| 137 |
+
| only `stages-1-8/s3_rules` | 0.554 | 0.724 | 0.824 | 1.234 | 0.668 | 0.674 | 0.782 | 0.491 | 0.951 |
|
| 138 |
+
| only `stages-1-8/s4_arith` | 0.553 | 0.721 | 0.928 | 1.127 | 0.657 | 0.659 | 0.783 | 0.489 | 0.951 |
|
| 139 |
+
| only `stages-1-8/s5_causal` | 0.536 | 0.710 | 0.917 | 1.225 | 0.091 | 0.666 | 0.782 | 0.484 | 0.951 |
|
| 140 |
+
| only `stages-1-8/s6_tryfail` | 0.552 | 0.722 | 0.930 | 1.237 | 0.651 | 0.164 | 0.783 | 0.490 | 0.951 |
|
| 141 |
+
| only `stages-1-8/s7_mixed` | 0.497 | 0.715 | 0.868 | 1.195 | 0.558 | 0.593 | 0.724 | 0.483 | 0.950 |
|
| 142 |
+
| only `stages-1-8/s8_register` | 0.537 | 0.713 | 0.925 | 1.227 | 0.642 | 0.651 | 0.783 | 0.068 | 0.951 |
|
| 143 |
+
| the whole group `stages-1-8` | 0.049 | 0.110 | 0.224 | 0.328 | 0.043 | 0.055 | 0.720 | 0.021 | 0.952 |
|
| 144 |
+
|
| 145 |
**How to read the tables.** *Stage text* is the loss, in bits per byte, on
|
| 146 |
held-out text of the arm's own stage, with arms off and with the whole group
|
| 147 |
on. *This arm alone* is the same loss with only that one arm mounted. *Stage
|
|
|
|
| 150 |
not. The *web text* figure is the change on held-out ordinary web text
|
| 151 |
(fineweb-edu) with the group on: the arms were trained to leave it alone.
|
| 152 |
|
| 153 |
+
**A group is the unit.** The eight-arm group is the four-arm group with four more arms fitted over it, and it keeps the four as they were: with all eight on, the text of stages 1 to 4 reads within 0.011 bits per byte of the four-arm group's own reading (stage 4 slightly better, the others the same). The new arms take their stages: with the whole group on, stage 5 text reads 0.043 bits per byte against 0.660 bare, stage 6 0.055 against 0.673, stage 8 0.021 against 0.491, on both seeds. Each of those three carries about two thirds to three quarters of its stage's gain itself and the fixed earlier arms supply the rest (the first arm, a generalist, reads stage 5 and 6 text 0.2 and 0.18 below bare on its own), so again the group works as a whole and is measured that way. The mixed stage (7) is the exception: its text is a blend of the other stages' kinds, the bare model reads it at 0.783, and no arm moves it much (the group reads it at 0.720, its own arm alone at 0.724). The group costs 0.001 to 0.002 bits per byte on ordinary web text, no arm costs more than 0.0006 of that, and no arm writes on another arm's stage text above 0.0012. Nothing an earlier arm learned was lost when a later arm trained over it (the largest loss of an earlier stage's gain at any later close was 2 percent). The four new arms trained for 4,000 steps each, about where a solo arm's curve flattens; the first four kept their shorter training (3,200, 2,400, 800 and 800 steps). For one stage by itself, use its solo arm below.
|
| 154 |
|
| 155 |
+
**What the arms do, plainly.** A group recovers each of its stages' own kind
|
| 156 |
+
of text. The gain is in each stage's own forms. Whether it carries over to
|
| 157 |
+
new forms is shown in each table's last column, where it was measured.
|
| 158 |
|
| 159 |
### Arms off and arms on
|
| 160 |
|
| 161 |
```python
|
| 162 |
m.arm # None: arms off
|
| 163 |
+
h = m.mount_arm("stages-1-8") # arms on: the whole eight-arm group
|
| 164 |
with h.all_off(): # every member masked: the bare core's logits
|
| 165 |
m.say("Hello there.")
|
| 166 |
m.detach_arm(verify=True) # raises if the restored core is not bit-exact
|
| 167 |
m.mount_arm("stages-1-4/s1_perspective") # one member alone (the "alone" readings)
|
| 168 |
```
|
| 169 |
|
| 170 |
+
A group mounts its members in the order they were trained, always on, one
|
| 171 |
after another at every block, with no mixer between them. `default_arm` in
|
| 172 |
+
`from_pretrained` (or in `config.json`) mounts a group as the model loads.
|
| 173 |
`save_pretrained` refuses while an arm is mounted, so an arm can never be
|
| 174 |
written into the core's weight file.
|
| 175 |
|
|
|
|
| 185 |
| `solo/s2_concept` | kinds, properties and differences: what a thing is a kind of, what its kind can do, how two things differ | 0.721 → 0.120 (a second seed: 0.120) | +0.0003 | 30% → 84% | 68% → 73% | two |
|
| 186 |
| `solo/s3_rules` | if-then rules over made-up words, followed step by step to what follows | 0.930 → 0.372 (a second seed: 0.374) | +0.0009 | 16% → 60% | 7% → 12% | two |
|
| 187 |
| `solo/s4_arith` | small arithmetic worked out in text | 1.234 → 0.389 (a second seed: 0.388) | +0.0007 | 34% → 49% | 24% → 24% | two |
|
| 188 |
+
| `solo/s5_causal` | short cause-and-effect text: what happened and why | 0.660 → 0.068 | +0.0003 | not measured | not measured | one |
|
| 189 |
+
| `solo/s6_tryfail` | an attempt, its failure and the revised attempt | 0.673 → 0.114 | +0.0008 | not measured | not measured | one |
|
| 190 |
+
| `solo/s7_mixed` | the mixed stage's diet: the earlier stages' text beside ordinary prose | 0.783 → 0.755 | +0.0004 | not measured | not measured | one |
|
| 191 |
+
| `solo/s8_register` | the same content in different registers of speech | 0.491 → 0.051 | +0.0008 | not measured | not measured | one |
|
| 192 |
|
| 193 |
```python
|
| 194 |
m.mount_arm("solo/s1_perspective") # one solo arm; mounting another swaps it
|
|
|
|
| 296 |
each stage's text in before its arm could, and they were detached at step
|
| 297 |
148,000 (their files are in the training repo under `mini-beatrix-3/arms/`,
|
| 298 |
bound to those earlier trunk states). The arms in this repository were fitted
|
| 299 |
+
afterwards on the finished, frozen core: as always-on groups and as solo arms. The four-arm group in two steps (first the arms of stages 1 and 2, in a run that added the four stages' text one stage at a time with every attached arm training under one pure Adam, 800 steps per stage; then, with those two held fixed and on, a new arm for stage 3 and after it a new arm for stage 4, 800 steps each). The eight-arm group by continuing that run: with the four held fixed and on, a new arm for stage 5, then 6, then 7, then 8, each attached over every arm before it and trained for 4,000 steps in their presence. Every group arm was kept quiet on ordinary web text and on the stage text of every other arm in its group. Each solo arm was trained alone, 800 steps, kept quiet on web text.
|
| 300 |
|
| 301 |
## Lineage
|
| 302 |
|
arms/index.json
CHANGED
|
@@ -6,8 +6,8 @@
|
|
| 6 |
"base_model_id": "alephllm/mini-beatrix-3@step245674",
|
| 7 |
"note": "every packaged arm was trained on this exact frozen core (the default bf16 weight file); arms are trunk-bound"
|
| 8 |
},
|
| 9 |
-
"generated": "2026-10-06
|
| 10 |
-
"note": "the stage arms of curriculum stages 1 to
|
| 11 |
"arms": [
|
| 12 |
{
|
| 13 |
"id": "stages-1-4/s1_perspective",
|
|
@@ -526,6 +526,1031 @@
|
|
| 526 |
},
|
| 527 |
"examples": []
|
| 528 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 529 |
{
|
| 530 |
"id": "solo/s1_perspective",
|
| 531 |
"name": "solo/s1_perspective",
|
|
@@ -849,6 +1874,214 @@
|
|
| 849 |
"examples": [
|
| 850 |
"Count by 3s: 1 3s make 3; 2 3s make 6; 3 3s make 9; 4 3s make 12; 5 3s make 15. So 3 times 5 is 15."
|
| 851 |
]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 852 |
}
|
| 853 |
]
|
| 854 |
}
|
|
|
|
| 6 |
"base_model_id": "alephllm/mini-beatrix-3@step245674",
|
| 7 |
"note": "every packaged arm was trained on this exact frozen core (the default bf16 weight file); arms are trunk-bound"
|
| 8 |
},
|
| 9 |
+
"generated": "2026-10-06 04:54 UTC",
|
| 10 |
+
"note": "the stage arms of curriculum stages 1 to 8, fitted on the finished, frozen core: two always-on groups (the four arms of stages 1 to 4; the eight arms of stages 1 to 8, which are those four, held fixed, with four more fitted over them) and each stage's arm alone; the arms trained beside the trunk during the run are on the training repo",
|
| 11 |
"arms": [
|
| 12 |
{
|
| 13 |
"id": "stages-1-4/s1_perspective",
|
|
|
|
| 526 |
},
|
| 527 |
"examples": []
|
| 528 |
},
|
| 529 |
+
{
|
| 530 |
+
"id": "stages-1-8/s1_perspective",
|
| 531 |
+
"name": "stages-1-8/s1_perspective",
|
| 532 |
+
"group": "stages-1-8",
|
| 533 |
+
"base_model_id": "alephllm/mini-beatrix-3@step245674",
|
| 534 |
+
"step": 245674,
|
| 535 |
+
"params": 13692960,
|
| 536 |
+
"sites": 32,
|
| 537 |
+
"weights": [
|
| 538 |
+
"arms/stages-1-8/s1_perspective.safetensors"
|
| 539 |
+
],
|
| 540 |
+
"mount": {
|
| 541 |
+
"type": "anchor"
|
| 542 |
+
},
|
| 543 |
+
"precision": "fp32",
|
| 544 |
+
"template": {
|
| 545 |
+
"frame": "raw",
|
| 546 |
+
"template": "raw stage text, continued"
|
| 547 |
+
},
|
| 548 |
+
"decode": {
|
| 549 |
+
"temperature": 0.0,
|
| 550 |
+
"top_p": 1.0,
|
| 551 |
+
"max_new": 120,
|
| 552 |
+
"precision": "fp32"
|
| 553 |
+
},
|
| 554 |
+
"source": {
|
| 555 |
+
"file": "gXA_close8_s1_perspective.safetensors",
|
| 556 |
+
"sha256": "65cb5c57cfd64fdbaaaf47d32a37969adc91bb2b1071cc97e089943041ebe044"
|
| 557 |
+
},
|
| 558 |
+
"recipe": "trained as an always-on group on the frozen final core by continuing the four-arm group's run. The four arms of stages 1 to 4 were loaded as trained (see the group stages-1-4) and held fixed and on. Then a new arm for stage 5 was attached over them and trained for 4,000 steps in their presence; then, with it also held fixed, a new arm for stage 6 the same way, then stage 7, then stage 8. Each step gives the training arm 2 x 4096 bytes of its own stage's text with every attached arm on, and a quiet term at weight 2 on 2 rows that are not its text (one of web text, one of the stage text of another arm of the eight, drawn from every other stage whether or not its arm was attached yet; the mixed stage's text, which contains the other stages' kinds, is drawn as no arm's partner row): the KL from the model with that arm switched off to the model with every arm on. Pure Adam, lr 0.001",
|
| 559 |
+
"title": "Stage 1: perspective",
|
| 560 |
+
"behavior": "one small event retold from each side: seen from a person, said to them, told about them",
|
| 561 |
+
"family": "stage",
|
| 562 |
+
"status": "a member of the group stages-1-8; trained always on with its partners, so alone is not its trained state",
|
| 563 |
+
"score": "alone: own stage text 0.554 -> 0.052 bpb; web text +0.0014 bpb",
|
| 564 |
+
"measured": {
|
| 565 |
+
"own_stage_bpb_arms_off": 0.5538,
|
| 566 |
+
"own_stage_bpb_group_on": 0.0492,
|
| 567 |
+
"own_stage_bpb_this_arm_alone": 0.0525,
|
| 568 |
+
"share_of_the_group_on_own_stage": -0.3832,
|
| 569 |
+
"share_of_the_group_on_own_stage_se": 0.0052,
|
| 570 |
+
"share_of_the_group_on_web_text": 0.00025,
|
| 571 |
+
"alone_on_web_text": 0.00143,
|
| 572 |
+
"share_of_the_group_on_partner_stages": {
|
| 573 |
+
"s2_concept": -0.2174,
|
| 574 |
+
"s3_rules": -0.1832,
|
| 575 |
+
"s4_arith": -0.128,
|
| 576 |
+
"s5_causal": -0.0128,
|
| 577 |
+
"s6_tryfail": -0.028,
|
| 578 |
+
"s7_mixed": -0.0029,
|
| 579 |
+
"s8_register": -0.0272
|
| 580 |
+
},
|
| 581 |
+
"alone_on_partner_stages": {
|
| 582 |
+
"s2_concept": -0.4817,
|
| 583 |
+
"s3_rules": -0.4209,
|
| 584 |
+
"s4_arith": -0.5113,
|
| 585 |
+
"s5_causal": -0.1978,
|
| 586 |
+
"s6_tryfail": -0.1782,
|
| 587 |
+
"s7_mixed": -0.0101,
|
| 588 |
+
"s8_register": -0.1119
|
| 589 |
+
},
|
| 590 |
+
"steps_trained": 3200,
|
| 591 |
+
"stage_items_own_form_accuracy": {
|
| 592 |
+
"arms_off": 0.55,
|
| 593 |
+
"group_on": 0.887,
|
| 594 |
+
"this_arm_alone": 0.838
|
| 595 |
+
},
|
| 596 |
+
"stage_items_new_form_accuracy": {
|
| 597 |
+
"arms_off": 0.463,
|
| 598 |
+
"group_on": 0.438,
|
| 599 |
+
"this_arm_alone": 0.463
|
| 600 |
+
},
|
| 601 |
+
"second_seed": {
|
| 602 |
+
"own_stage_bpb_arms_off": 0.5538,
|
| 603 |
+
"own_stage_bpb_group_on": 0.0484,
|
| 604 |
+
"own_stage_bpb_this_arm_alone": 0.0522,
|
| 605 |
+
"share_of_the_group_on_own_stage": -0.4405,
|
| 606 |
+
"share_of_the_group_on_own_stage_se": 0.0064,
|
| 607 |
+
"share_of_the_group_on_web_text": 0.00043,
|
| 608 |
+
"alone_on_web_text": 0.00134,
|
| 609 |
+
"share_of_the_group_on_partner_stages": {
|
| 610 |
+
"s2_concept": -0.2289,
|
| 611 |
+
"s3_rules": -0.1943,
|
| 612 |
+
"s4_arith": -0.13,
|
| 613 |
+
"s5_causal": -0.0282,
|
| 614 |
+
"s6_tryfail": -0.0206,
|
| 615 |
+
"s7_mixed": -0.0043,
|
| 616 |
+
"s8_register": -0.0214
|
| 617 |
+
},
|
| 618 |
+
"alone_on_partner_stages": {
|
| 619 |
+
"s2_concept": -0.4772,
|
| 620 |
+
"s3_rules": -0.4262,
|
| 621 |
+
"s4_arith": -0.4807,
|
| 622 |
+
"s5_causal": -0.1919,
|
| 623 |
+
"s6_tryfail": -0.169,
|
| 624 |
+
"s7_mixed": -0.0088,
|
| 625 |
+
"s8_register": -0.1105
|
| 626 |
+
},
|
| 627 |
+
"steps_trained": 3200,
|
| 628 |
+
"stage_items_own_form_accuracy": {
|
| 629 |
+
"arms_off": 0.55,
|
| 630 |
+
"group_on": 1.0,
|
| 631 |
+
"this_arm_alone": 0.925
|
| 632 |
+
},
|
| 633 |
+
"stage_items_new_form_accuracy": {
|
| 634 |
+
"arms_off": 0.463,
|
| 635 |
+
"group_on": 0.425,
|
| 636 |
+
"this_arm_alone": 0.463
|
| 637 |
+
}
|
| 638 |
+
}
|
| 639 |
+
},
|
| 640 |
+
"examples": [
|
| 641 |
+
"Rui dropped the acorn near the box. Seen from Rui: I dropped the acorn near the box. Said to Rui: you dropped the acorn near the box. Told about Rui: Rui dropped the acorn near the box. Three ways of saying, one thing that happened."
|
| 642 |
+
]
|
| 643 |
+
},
|
| 644 |
+
{
|
| 645 |
+
"id": "stages-1-8/s2_concept",
|
| 646 |
+
"name": "stages-1-8/s2_concept",
|
| 647 |
+
"group": "stages-1-8",
|
| 648 |
+
"base_model_id": "alephllm/mini-beatrix-3@step245674",
|
| 649 |
+
"step": 245674,
|
| 650 |
+
"params": 13692960,
|
| 651 |
+
"sites": 32,
|
| 652 |
+
"weights": [
|
| 653 |
+
"arms/stages-1-8/s2_concept.safetensors"
|
| 654 |
+
],
|
| 655 |
+
"mount": {
|
| 656 |
+
"type": "anchor"
|
| 657 |
+
},
|
| 658 |
+
"precision": "fp32",
|
| 659 |
+
"template": {
|
| 660 |
+
"frame": "raw",
|
| 661 |
+
"template": "raw stage text, continued"
|
| 662 |
+
},
|
| 663 |
+
"decode": {
|
| 664 |
+
"temperature": 0.0,
|
| 665 |
+
"top_p": 1.0,
|
| 666 |
+
"max_new": 120,
|
| 667 |
+
"precision": "fp32"
|
| 668 |
+
},
|
| 669 |
+
"source": {
|
| 670 |
+
"file": "gXA_close8_s2_concept.safetensors",
|
| 671 |
+
"sha256": "f411b5c9ae9994a26e0cbd25475a754fee5af6b7e58fb2af5dedbe790dd602f5"
|
| 672 |
+
},
|
| 673 |
+
"recipe": "trained as an always-on group on the frozen final core by continuing the four-arm group's run. The four arms of stages 1 to 4 were loaded as trained (see the group stages-1-4) and held fixed and on. Then a new arm for stage 5 was attached over them and trained for 4,000 steps in their presence; then, with it also held fixed, a new arm for stage 6 the same way, then stage 7, then stage 8. Each step gives the training arm 2 x 4096 bytes of its own stage's text with every attached arm on, and a quiet term at weight 2 on 2 rows that are not its text (one of web text, one of the stage text of another arm of the eight, drawn from every other stage whether or not its arm was attached yet; the mixed stage's text, which contains the other stages' kinds, is drawn as no arm's partner row): the KL from the model with that arm switched off to the model with every arm on. Pure Adam, lr 0.001",
|
| 674 |
+
"title": "Stage 2: concepts",
|
| 675 |
+
"behavior": "kinds, properties and differences: what a thing is a kind of, what its kind can do, how two things differ",
|
| 676 |
+
"family": "stage",
|
| 677 |
+
"status": "a member of the group stages-1-8; trained always on with its partners, so alone is not its trained state",
|
| 678 |
+
"score": "alone: own stage text 0.721 -> 0.350 bpb; web text +0.0010 bpb",
|
| 679 |
+
"measured": {
|
| 680 |
+
"own_stage_bpb_arms_off": 0.7212,
|
| 681 |
+
"own_stage_bpb_group_on": 0.1105,
|
| 682 |
+
"own_stage_bpb_this_arm_alone": 0.3503,
|
| 683 |
+
"share_of_the_group_on_own_stage": -0.1202,
|
| 684 |
+
"share_of_the_group_on_own_stage_se": 0.0011,
|
| 685 |
+
"share_of_the_group_on_web_text": -0.00043,
|
| 686 |
+
"alone_on_web_text": 0.00105,
|
| 687 |
+
"share_of_the_group_on_partner_stages": {
|
| 688 |
+
"s1_perspective": -0.0057,
|
| 689 |
+
"s3_rules": -0.1086,
|
| 690 |
+
"s4_arith": -0.106,
|
| 691 |
+
"s5_causal": -0.0013,
|
| 692 |
+
"s6_tryfail": -0.0035,
|
| 693 |
+
"s7_mixed": -0.0008,
|
| 694 |
+
"s8_register": -0.0007
|
| 695 |
+
},
|
| 696 |
+
"alone_on_partner_stages": {
|
| 697 |
+
"s1_perspective": -0.0273,
|
| 698 |
+
"s3_rules": -0.3108,
|
| 699 |
+
"s4_arith": -0.3537,
|
| 700 |
+
"s5_causal": -0.0015,
|
| 701 |
+
"s6_tryfail": 0.0004,
|
| 702 |
+
"s7_mixed": -0.0,
|
| 703 |
+
"s8_register": -0.0215
|
| 704 |
+
},
|
| 705 |
+
"steps_trained": 2400,
|
| 706 |
+
"stage_items_own_form_accuracy": {
|
| 707 |
+
"arms_off": 0.298,
|
| 708 |
+
"group_on": 0.976,
|
| 709 |
+
"this_arm_alone": 0.964
|
| 710 |
+
},
|
| 711 |
+
"stage_items_new_form_accuracy": {
|
| 712 |
+
"arms_off": 0.678,
|
| 713 |
+
"group_on": 0.72,
|
| 714 |
+
"this_arm_alone": 0.72
|
| 715 |
+
},
|
| 716 |
+
"second_seed": {
|
| 717 |
+
"own_stage_bpb_arms_off": 0.7212,
|
| 718 |
+
"own_stage_bpb_group_on": 0.1212,
|
| 719 |
+
"own_stage_bpb_this_arm_alone": 0.3603,
|
| 720 |
+
"share_of_the_group_on_own_stage": -0.119,
|
| 721 |
+
"share_of_the_group_on_own_stage_se": 0.0012,
|
| 722 |
+
"share_of_the_group_on_web_text": -0.00073,
|
| 723 |
+
"alone_on_web_text": 0.00086,
|
| 724 |
+
"share_of_the_group_on_partner_stages": {
|
| 725 |
+
"s1_perspective": -0.0042,
|
| 726 |
+
"s3_rules": -0.1006,
|
| 727 |
+
"s4_arith": -0.0957,
|
| 728 |
+
"s5_causal": -0.0009,
|
| 729 |
+
"s6_tryfail": -0.0023,
|
| 730 |
+
"s7_mixed": -0.0004,
|
| 731 |
+
"s8_register": -0.0005
|
| 732 |
+
},
|
| 733 |
+
"alone_on_partner_stages": {
|
| 734 |
+
"s1_perspective": -0.03,
|
| 735 |
+
"s3_rules": -0.3138,
|
| 736 |
+
"s4_arith": -0.3577,
|
| 737 |
+
"s5_causal": -0.0127,
|
| 738 |
+
"s6_tryfail": -0.0062,
|
| 739 |
+
"s7_mixed": 0.0002,
|
| 740 |
+
"s8_register": -0.0193
|
| 741 |
+
},
|
| 742 |
+
"steps_trained": 2400,
|
| 743 |
+
"stage_items_own_form_accuracy": {
|
| 744 |
+
"arms_off": 0.298,
|
| 745 |
+
"group_on": 0.952,
|
| 746 |
+
"this_arm_alone": 0.905
|
| 747 |
+
},
|
| 748 |
+
"stage_items_new_form_accuracy": {
|
| 749 |
+
"arms_off": 0.678,
|
| 750 |
+
"group_on": 0.72,
|
| 751 |
+
"this_arm_alone": 0.703
|
| 752 |
+
}
|
| 753 |
+
}
|
| 754 |
+
},
|
| 755 |
+
"examples": [
|
| 756 |
+
"A whale is a kind of mammal. Most mammals can swim. A whale is not a beetle; they differ in kind."
|
| 757 |
+
]
|
| 758 |
+
},
|
| 759 |
+
{
|
| 760 |
+
"id": "stages-1-8/s3_rules",
|
| 761 |
+
"name": "stages-1-8/s3_rules",
|
| 762 |
+
"group": "stages-1-8",
|
| 763 |
+
"base_model_id": "alephllm/mini-beatrix-3@step245674",
|
| 764 |
+
"step": 245674,
|
| 765 |
+
"params": 13692960,
|
| 766 |
+
"sites": 32,
|
| 767 |
+
"weights": [
|
| 768 |
+
"arms/stages-1-8/s3_rules.safetensors"
|
| 769 |
+
],
|
| 770 |
+
"mount": {
|
| 771 |
+
"type": "anchor"
|
| 772 |
+
},
|
| 773 |
+
"precision": "fp32",
|
| 774 |
+
"template": {
|
| 775 |
+
"frame": "raw",
|
| 776 |
+
"template": "raw stage text, continued"
|
| 777 |
+
},
|
| 778 |
+
"decode": {
|
| 779 |
+
"temperature": 0.0,
|
| 780 |
+
"top_p": 1.0,
|
| 781 |
+
"max_new": 120,
|
| 782 |
+
"precision": "fp32"
|
| 783 |
+
},
|
| 784 |
+
"source": {
|
| 785 |
+
"file": "gXA_close8_s3_rules.safetensors",
|
| 786 |
+
"sha256": "bd9f4dd7e15f28ca140e1c86962b0f4ea4b144a8af369219acd270b1368dc5b8"
|
| 787 |
+
},
|
| 788 |
+
"recipe": "trained as an always-on group on the frozen final core by continuing the four-arm group's run. The four arms of stages 1 to 4 were loaded as trained (see the group stages-1-4) and held fixed and on. Then a new arm for stage 5 was attached over them and trained for 4,000 steps in their presence; then, with it also held fixed, a new arm for stage 6 the same way, then stage 7, then stage 8. Each step gives the training arm 2 x 4096 bytes of its own stage's text with every attached arm on, and a quiet term at weight 2 on 2 rows that are not its text (one of web text, one of the stage text of another arm of the eight, drawn from every other stage whether or not its arm was attached yet; the mixed stage's text, which contains the other stages' kinds, is drawn as no arm's partner row): the KL from the model with that arm switched off to the model with every arm on. Pure Adam, lr 0.001",
|
| 789 |
+
"title": "Stage 3: rule chains",
|
| 790 |
+
"behavior": "if-then rules over made-up words, followed step by step to what follows",
|
| 791 |
+
"family": "stage",
|
| 792 |
+
"status": "a member of the group stages-1-8; trained always on with its partners, so alone is not its trained state",
|
| 793 |
+
"score": "alone: own stage text 0.930 -> 0.824 bpb; web text +0.0005 bpb",
|
| 794 |
+
"measured": {
|
| 795 |
+
"own_stage_bpb_arms_off": 0.9296,
|
| 796 |
+
"own_stage_bpb_group_on": 0.2237,
|
| 797 |
+
"own_stage_bpb_this_arm_alone": 0.8236,
|
| 798 |
+
"share_of_the_group_on_own_stage": -0.09,
|
| 799 |
+
"share_of_the_group_on_own_stage_se": 0.0013,
|
| 800 |
+
"share_of_the_group_on_web_text": 9e-05,
|
| 801 |
+
"alone_on_web_text": 0.00053,
|
| 802 |
+
"share_of_the_group_on_partner_stages": {
|
| 803 |
+
"s1_perspective": 0.0001,
|
| 804 |
+
"s2_concept": 0.0005,
|
| 805 |
+
"s4_arith": -0.0002,
|
| 806 |
+
"s5_causal": -0.0001,
|
| 807 |
+
"s6_tryfail": -0.0003,
|
| 808 |
+
"s7_mixed": 0.0001,
|
| 809 |
+
"s8_register": -0.0001
|
| 810 |
+
},
|
| 811 |
+
"alone_on_partner_stages": {
|
| 812 |
+
"s1_perspective": -0.0003,
|
| 813 |
+
"s2_concept": 0.0031,
|
| 814 |
+
"s4_arith": 0.0001,
|
| 815 |
+
"s5_causal": 0.0071,
|
| 816 |
+
"s6_tryfail": 0.001,
|
| 817 |
+
"s7_mixed": -0.0005,
|
| 818 |
+
"s8_register": 0.0002
|
| 819 |
+
},
|
| 820 |
+
"steps_trained": 800,
|
| 821 |
+
"stage_items_own_form_accuracy": {
|
| 822 |
+
"arms_off": 0.155,
|
| 823 |
+
"group_on": 0.855,
|
| 824 |
+
"this_arm_alone": 0.32
|
| 825 |
+
},
|
| 826 |
+
"stage_items_new_form_accuracy": {
|
| 827 |
+
"arms_off": 0.07,
|
| 828 |
+
"group_on": 0.31,
|
| 829 |
+
"this_arm_alone": 0.08
|
| 830 |
+
},
|
| 831 |
+
"second_seed": {
|
| 832 |
+
"own_stage_bpb_arms_off": 0.9296,
|
| 833 |
+
"own_stage_bpb_group_on": 0.2222,
|
| 834 |
+
"own_stage_bpb_this_arm_alone": 0.8702,
|
| 835 |
+
"share_of_the_group_on_own_stage": -0.0915,
|
| 836 |
+
"share_of_the_group_on_own_stage_se": 0.0015,
|
| 837 |
+
"share_of_the_group_on_web_text": 0.00059,
|
| 838 |
+
"alone_on_web_text": 0.00034,
|
| 839 |
+
"share_of_the_group_on_partner_stages": {
|
| 840 |
+
"s1_perspective": 0.0,
|
| 841 |
+
"s2_concept": -0.0005,
|
| 842 |
+
"s4_arith": -0.0004,
|
| 843 |
+
"s5_causal": -0.0001,
|
| 844 |
+
"s6_tryfail": -0.0003,
|
| 845 |
+
"s7_mixed": -0.0001,
|
| 846 |
+
"s8_register": 0.0
|
| 847 |
+
},
|
| 848 |
+
"alone_on_partner_stages": {
|
| 849 |
+
"s1_perspective": -0.0003,
|
| 850 |
+
"s2_concept": 0.0016,
|
| 851 |
+
"s4_arith": -0.0047,
|
| 852 |
+
"s5_causal": 0.006,
|
| 853 |
+
"s6_tryfail": -0.0011,
|
| 854 |
+
"s7_mixed": -0.0007,
|
| 855 |
+
"s8_register": -0.0044
|
| 856 |
+
},
|
| 857 |
+
"steps_trained": 800,
|
| 858 |
+
"stage_items_own_form_accuracy": {
|
| 859 |
+
"arms_off": 0.155,
|
| 860 |
+
"group_on": 0.92,
|
| 861 |
+
"this_arm_alone": 0.335
|
| 862 |
+
},
|
| 863 |
+
"stage_items_new_form_accuracy": {
|
| 864 |
+
"arms_off": 0.07,
|
| 865 |
+
"group_on": 0.37,
|
| 866 |
+
"this_arm_alone": 0.1
|
| 867 |
+
}
|
| 868 |
+
}
|
| 869 |
+
},
|
| 870 |
+
"examples": [
|
| 871 |
+
"If someone is hoream, then they are hortil. If someone is hortil, then they are babig. If someone is babig, then they are banng. Vin is hoream. So Vin is hortil. So Vin is babig. So Vin is banng."
|
| 872 |
+
]
|
| 873 |
+
},
|
| 874 |
+
{
|
| 875 |
+
"id": "stages-1-8/s4_arith",
|
| 876 |
+
"name": "stages-1-8/s4_arith",
|
| 877 |
+
"group": "stages-1-8",
|
| 878 |
+
"base_model_id": "alephllm/mini-beatrix-3@step245674",
|
| 879 |
+
"step": 245674,
|
| 880 |
+
"params": 13692960,
|
| 881 |
+
"sites": 32,
|
| 882 |
+
"weights": [
|
| 883 |
+
"arms/stages-1-8/s4_arith.safetensors"
|
| 884 |
+
],
|
| 885 |
+
"mount": {
|
| 886 |
+
"type": "anchor"
|
| 887 |
+
},
|
| 888 |
+
"precision": "fp32",
|
| 889 |
+
"template": {
|
| 890 |
+
"frame": "raw",
|
| 891 |
+
"template": "raw stage text, continued"
|
| 892 |
+
},
|
| 893 |
+
"decode": {
|
| 894 |
+
"temperature": 0.0,
|
| 895 |
+
"top_p": 1.0,
|
| 896 |
+
"max_new": 120,
|
| 897 |
+
"precision": "fp32"
|
| 898 |
+
},
|
| 899 |
+
"source": {
|
| 900 |
+
"file": "gXA_close8_s4_arith.safetensors",
|
| 901 |
+
"sha256": "58552d2cdf11bce31dcad8e1e09fe27b6c760857aef47add6b4c92b0487669b7"
|
| 902 |
+
},
|
| 903 |
+
"recipe": "trained as an always-on group on the frozen final core by continuing the four-arm group's run. The four arms of stages 1 to 4 were loaded as trained (see the group stages-1-4) and held fixed and on. Then a new arm for stage 5 was attached over them and trained for 4,000 steps in their presence; then, with it also held fixed, a new arm for stage 6 the same way, then stage 7, then stage 8. Each step gives the training arm 2 x 4096 bytes of its own stage's text with every attached arm on, and a quiet term at weight 2 on 2 rows that are not its text (one of web text, one of the stage text of another arm of the eight, drawn from every other stage whether or not its arm was attached yet; the mixed stage's text, which contains the other stages' kinds, is drawn as no arm's partner row): the KL from the model with that arm switched off to the model with every arm on. Pure Adam, lr 0.001",
|
| 904 |
+
"title": "Stage 4: arithmetic",
|
| 905 |
+
"behavior": "small arithmetic worked out in text",
|
| 906 |
+
"family": "stage",
|
| 907 |
+
"status": "a member of the group stages-1-8; trained always on with its partners, so alone is not its trained state",
|
| 908 |
+
"score": "alone: own stage text 1.234 -> 1.127 bpb; web text +0.0001 bpb",
|
| 909 |
+
"measured": {
|
| 910 |
+
"own_stage_bpb_arms_off": 1.2339,
|
| 911 |
+
"own_stage_bpb_group_on": 0.3279,
|
| 912 |
+
"own_stage_bpb_this_arm_alone": 1.1267,
|
| 913 |
+
"share_of_the_group_on_own_stage": -0.1431,
|
| 914 |
+
"share_of_the_group_on_own_stage_se": 0.0023,
|
| 915 |
+
"share_of_the_group_on_web_text": 0.0002,
|
| 916 |
+
"alone_on_web_text": 0.00011,
|
| 917 |
+
"share_of_the_group_on_partner_stages": {
|
| 918 |
+
"s1_perspective": 0.0,
|
| 919 |
+
"s2_concept": -0.0006,
|
| 920 |
+
"s3_rules": -0.0001,
|
| 921 |
+
"s5_causal": -0.0,
|
| 922 |
+
"s6_tryfail": -0.0036,
|
| 923 |
+
"s7_mixed": 0.0001,
|
| 924 |
+
"s8_register": -0.0
|
| 925 |
+
},
|
| 926 |
+
"alone_on_partner_stages": {
|
| 927 |
+
"s1_perspective": -0.0009,
|
| 928 |
+
"s2_concept": -0.0004,
|
| 929 |
+
"s3_rules": -0.0016,
|
| 930 |
+
"s5_causal": -0.0033,
|
| 931 |
+
"s6_tryfail": -0.0135,
|
| 932 |
+
"s7_mixed": -0.0002,
|
| 933 |
+
"s8_register": -0.0015
|
| 934 |
+
},
|
| 935 |
+
"steps_trained": 800,
|
| 936 |
+
"stage_items_own_form_accuracy": {
|
| 937 |
+
"arms_off": 0.344,
|
| 938 |
+
"group_on": 0.644,
|
| 939 |
+
"this_arm_alone": 0.381
|
| 940 |
+
},
|
| 941 |
+
"stage_items_new_form_accuracy": {
|
| 942 |
+
"arms_off": 0.242,
|
| 943 |
+
"group_on": 0.242,
|
| 944 |
+
"this_arm_alone": 0.246
|
| 945 |
+
},
|
| 946 |
+
"second_seed": {
|
| 947 |
+
"own_stage_bpb_arms_off": 1.2339,
|
| 948 |
+
"own_stage_bpb_group_on": 0.3266,
|
| 949 |
+
"own_stage_bpb_this_arm_alone": 1.155,
|
| 950 |
+
"share_of_the_group_on_own_stage": -0.1643,
|
| 951 |
+
"share_of_the_group_on_own_stage_se": 0.0019,
|
| 952 |
+
"share_of_the_group_on_web_text": 0.00017,
|
| 953 |
+
"alone_on_web_text": 6e-05,
|
| 954 |
+
"share_of_the_group_on_partner_stages": {
|
| 955 |
+
"s1_perspective": 0.0,
|
| 956 |
+
"s2_concept": -0.0002,
|
| 957 |
+
"s3_rules": 0.0001,
|
| 958 |
+
"s5_causal": -0.0002,
|
| 959 |
+
"s6_tryfail": -0.0039,
|
| 960 |
+
"s7_mixed": -0.0,
|
| 961 |
+
"s8_register": -0.0001
|
| 962 |
+
},
|
| 963 |
+
"alone_on_partner_stages": {
|
| 964 |
+
"s1_perspective": -0.0013,
|
| 965 |
+
"s2_concept": -0.0015,
|
| 966 |
+
"s3_rules": 0.0005,
|
| 967 |
+
"s5_causal": -0.0011,
|
| 968 |
+
"s6_tryfail": -0.0134,
|
| 969 |
+
"s7_mixed": -0.0003,
|
| 970 |
+
"s8_register": -0.0014
|
| 971 |
+
},
|
| 972 |
+
"steps_trained": 800,
|
| 973 |
+
"stage_items_own_form_accuracy": {
|
| 974 |
+
"arms_off": 0.344,
|
| 975 |
+
"group_on": 0.656,
|
| 976 |
+
"this_arm_alone": 0.412
|
| 977 |
+
},
|
| 978 |
+
"stage_items_new_form_accuracy": {
|
| 979 |
+
"arms_off": 0.242,
|
| 980 |
+
"group_on": 0.271,
|
| 981 |
+
"this_arm_alone": 0.237
|
| 982 |
+
}
|
| 983 |
+
}
|
| 984 |
+
},
|
| 985 |
+
"examples": [
|
| 986 |
+
"Count by 3s: 1 3s make 3; 2 3s make 6; 3 3s make 9; 4 3s make 12; 5 3s make 15. So 3 times 5 is 15."
|
| 987 |
+
]
|
| 988 |
+
},
|
| 989 |
+
{
|
| 990 |
+
"id": "stages-1-8/s5_causal",
|
| 991 |
+
"name": "stages-1-8/s5_causal",
|
| 992 |
+
"group": "stages-1-8",
|
| 993 |
+
"base_model_id": "alephllm/mini-beatrix-3@step245674",
|
| 994 |
+
"step": 245674,
|
| 995 |
+
"params": 13692960,
|
| 996 |
+
"sites": 32,
|
| 997 |
+
"weights": [
|
| 998 |
+
"arms/stages-1-8/s5_causal.safetensors"
|
| 999 |
+
],
|
| 1000 |
+
"mount": {
|
| 1001 |
+
"type": "anchor"
|
| 1002 |
+
},
|
| 1003 |
+
"precision": "fp32",
|
| 1004 |
+
"template": {
|
| 1005 |
+
"frame": "raw",
|
| 1006 |
+
"template": "raw stage text, continued"
|
| 1007 |
+
},
|
| 1008 |
+
"decode": {
|
| 1009 |
+
"temperature": 0.0,
|
| 1010 |
+
"top_p": 1.0,
|
| 1011 |
+
"max_new": 120,
|
| 1012 |
+
"precision": "fp32"
|
| 1013 |
+
},
|
| 1014 |
+
"source": {
|
| 1015 |
+
"file": "gXA_close8_s5_causal.safetensors",
|
| 1016 |
+
"sha256": "8d1da9acaf7d528243304d7bebf71780242c7b1622ab3a1f72ef39e875b0f732"
|
| 1017 |
+
},
|
| 1018 |
+
"recipe": "trained as an always-on group on the frozen final core by continuing the four-arm group's run. The four arms of stages 1 to 4 were loaded as trained (see the group stages-1-4) and held fixed and on. Then a new arm for stage 5 was attached over them and trained for 4,000 steps in their presence; then, with it also held fixed, a new arm for stage 6 the same way, then stage 7, then stage 8. Each step gives the training arm 2 x 4096 bytes of its own stage's text with every attached arm on, and a quiet term at weight 2 on 2 rows that are not its text (one of web text, one of the stage text of another arm of the eight, drawn from every other stage whether or not its arm was attached yet; the mixed stage's text, which contains the other stages' kinds, is drawn as no arm's partner row): the KL from the model with that arm switched off to the model with every arm on. Pure Adam, lr 0.001",
|
| 1019 |
+
"title": "Stage 5: cause and effect",
|
| 1020 |
+
"behavior": "short cause-and-effect text: what happened and why",
|
| 1021 |
+
"family": "stage",
|
| 1022 |
+
"status": "a member of the group stages-1-8; trained always on with its partners, so alone is not its trained state",
|
| 1023 |
+
"score": "alone: own stage text 0.660 -> 0.091 bpb; web text +0.0003 bpb",
|
| 1024 |
+
"measured": {
|
| 1025 |
+
"own_stage_bpb_arms_off": 0.6604,
|
| 1026 |
+
"own_stage_bpb_group_on": 0.0431,
|
| 1027 |
+
"own_stage_bpb_this_arm_alone": 0.0911,
|
| 1028 |
+
"share_of_the_group_on_own_stage": -0.415,
|
| 1029 |
+
"share_of_the_group_on_own_stage_se": 0.0056,
|
| 1030 |
+
"share_of_the_group_on_web_text": 0.00035,
|
| 1031 |
+
"alone_on_web_text": 0.00028,
|
| 1032 |
+
"share_of_the_group_on_partner_stages": {
|
| 1033 |
+
"s1_perspective": 0.0007,
|
| 1034 |
+
"s2_concept": -0.0005,
|
| 1035 |
+
"s3_rules": 0.0005,
|
| 1036 |
+
"s4_arith": 0.0002,
|
| 1037 |
+
"s6_tryfail": -0.0005,
|
| 1038 |
+
"s7_mixed": 0.0004,
|
| 1039 |
+
"s8_register": -0.0003
|
| 1040 |
+
},
|
| 1041 |
+
"alone_on_partner_stages": {
|
| 1042 |
+
"s1_perspective": -0.0183,
|
| 1043 |
+
"s2_concept": -0.0109,
|
| 1044 |
+
"s3_rules": -0.0122,
|
| 1045 |
+
"s4_arith": -0.0091,
|
| 1046 |
+
"s6_tryfail": -0.0065,
|
| 1047 |
+
"s7_mixed": -0.0011,
|
| 1048 |
+
"s8_register": -0.0064
|
| 1049 |
+
},
|
| 1050 |
+
"steps_trained": 4000,
|
| 1051 |
+
"second_seed": {
|
| 1052 |
+
"own_stage_bpb_arms_off": 0.6604,
|
| 1053 |
+
"own_stage_bpb_group_on": 0.0428,
|
| 1054 |
+
"own_stage_bpb_this_arm_alone": 0.1418,
|
| 1055 |
+
"share_of_the_group_on_own_stage": -0.4245,
|
| 1056 |
+
"share_of_the_group_on_own_stage_se": 0.0045,
|
| 1057 |
+
"share_of_the_group_on_web_text": 0.00029,
|
| 1058 |
+
"alone_on_web_text": 0.0,
|
| 1059 |
+
"share_of_the_group_on_partner_stages": {
|
| 1060 |
+
"s1_perspective": 0.0001,
|
| 1061 |
+
"s2_concept": 0.0002,
|
| 1062 |
+
"s3_rules": 0.0004,
|
| 1063 |
+
"s4_arith": 0.0001,
|
| 1064 |
+
"s6_tryfail": -0.0003,
|
| 1065 |
+
"s7_mixed": 0.0001,
|
| 1066 |
+
"s8_register": -0.0001
|
| 1067 |
+
},
|
| 1068 |
+
"alone_on_partner_stages": {
|
| 1069 |
+
"s1_perspective": -0.0022,
|
| 1070 |
+
"s2_concept": -0.001,
|
| 1071 |
+
"s3_rules": -0.0018,
|
| 1072 |
+
"s4_arith": -0.0022,
|
| 1073 |
+
"s6_tryfail": -0.0023,
|
| 1074 |
+
"s7_mixed": -0.0001,
|
| 1075 |
+
"s8_register": -0.0009
|
| 1076 |
+
},
|
| 1077 |
+
"steps_trained": 4000
|
| 1078 |
+
}
|
| 1079 |
+
},
|
| 1080 |
+
"examples": [
|
| 1081 |
+
"Noor stacked the cups too high. Because of that, the tower leaned. Because of that, the cups crashed down."
|
| 1082 |
+
]
|
| 1083 |
+
},
|
| 1084 |
+
{
|
| 1085 |
+
"id": "stages-1-8/s6_tryfail",
|
| 1086 |
+
"name": "stages-1-8/s6_tryfail",
|
| 1087 |
+
"group": "stages-1-8",
|
| 1088 |
+
"base_model_id": "alephllm/mini-beatrix-3@step245674",
|
| 1089 |
+
"step": 245674,
|
| 1090 |
+
"params": 13692960,
|
| 1091 |
+
"sites": 32,
|
| 1092 |
+
"weights": [
|
| 1093 |
+
"arms/stages-1-8/s6_tryfail.safetensors"
|
| 1094 |
+
],
|
| 1095 |
+
"mount": {
|
| 1096 |
+
"type": "anchor"
|
| 1097 |
+
},
|
| 1098 |
+
"precision": "fp32",
|
| 1099 |
+
"template": {
|
| 1100 |
+
"frame": "raw",
|
| 1101 |
+
"template": "raw stage text, continued"
|
| 1102 |
+
},
|
| 1103 |
+
"decode": {
|
| 1104 |
+
"temperature": 0.0,
|
| 1105 |
+
"top_p": 1.0,
|
| 1106 |
+
"max_new": 120,
|
| 1107 |
+
"precision": "fp32"
|
| 1108 |
+
},
|
| 1109 |
+
"source": {
|
| 1110 |
+
"file": "gXA_close8_s6_tryfail.safetensors",
|
| 1111 |
+
"sha256": "2b7942956f7e2ded57b29168ef5a0916424a39bd3826ad7804fabec92050999b"
|
| 1112 |
+
},
|
| 1113 |
+
"recipe": "trained as an always-on group on the frozen final core by continuing the four-arm group's run. The four arms of stages 1 to 4 were loaded as trained (see the group stages-1-4) and held fixed and on. Then a new arm for stage 5 was attached over them and trained for 4,000 steps in their presence; then, with it also held fixed, a new arm for stage 6 the same way, then stage 7, then stage 8. Each step gives the training arm 2 x 4096 bytes of its own stage's text with every attached arm on, and a quiet term at weight 2 on 2 rows that are not its text (one of web text, one of the stage text of another arm of the eight, drawn from every other stage whether or not its arm was attached yet; the mixed stage's text, which contains the other stages' kinds, is drawn as no arm's partner row): the KL from the model with that arm switched off to the model with every arm on. Pure Adam, lr 0.001",
|
| 1114 |
+
"title": "Stage 6: try, fail, revise",
|
| 1115 |
+
"behavior": "an attempt, its failure and the revised attempt",
|
| 1116 |
+
"family": "stage",
|
| 1117 |
+
"status": "a member of the group stages-1-8; trained always on with its partners, so alone is not its trained state",
|
| 1118 |
+
"score": "alone: own stage text 0.673 -> 0.164 bpb; web text +0.0001 bpb",
|
| 1119 |
+
"measured": {
|
| 1120 |
+
"own_stage_bpb_arms_off": 0.6728,
|
| 1121 |
+
"own_stage_bpb_group_on": 0.0548,
|
| 1122 |
+
"own_stage_bpb_this_arm_alone": 0.1638,
|
| 1123 |
+
"share_of_the_group_on_own_stage": -0.4138,
|
| 1124 |
+
"share_of_the_group_on_own_stage_se": 0.0033,
|
| 1125 |
+
"share_of_the_group_on_web_text": -0.00025,
|
| 1126 |
+
"alone_on_web_text": 0.00013,
|
| 1127 |
+
"share_of_the_group_on_partner_stages": {
|
| 1128 |
+
"s1_perspective": -0.0001,
|
| 1129 |
+
"s2_concept": 0.0002,
|
| 1130 |
+
"s3_rules": -0.0002,
|
| 1131 |
+
"s4_arith": -0.0009,
|
| 1132 |
+
"s5_causal": -0.0001,
|
| 1133 |
+
"s7_mixed": 0.0001,
|
| 1134 |
+
"s8_register": -0.0006
|
| 1135 |
+
},
|
| 1136 |
+
"alone_on_partner_stages": {
|
| 1137 |
+
"s1_perspective": -0.002,
|
| 1138 |
+
"s2_concept": 0.0012,
|
| 1139 |
+
"s3_rules": 0.0006,
|
| 1140 |
+
"s4_arith": 0.0036,
|
| 1141 |
+
"s5_causal": -0.0098,
|
| 1142 |
+
"s7_mixed": -0.0,
|
| 1143 |
+
"s8_register": -0.0003
|
| 1144 |
+
},
|
| 1145 |
+
"steps_trained": 4000,
|
| 1146 |
+
"second_seed": {
|
| 1147 |
+
"own_stage_bpb_arms_off": 0.6728,
|
| 1148 |
+
"own_stage_bpb_group_on": 0.0543,
|
| 1149 |
+
"own_stage_bpb_this_arm_alone": 0.1437,
|
| 1150 |
+
"share_of_the_group_on_own_stage": -0.412,
|
| 1151 |
+
"share_of_the_group_on_own_stage_se": 0.0027,
|
| 1152 |
+
"share_of_the_group_on_web_text": 0.00025,
|
| 1153 |
+
"alone_on_web_text": 0.00011,
|
| 1154 |
+
"share_of_the_group_on_partner_stages": {
|
| 1155 |
+
"s1_perspective": 0.0,
|
| 1156 |
+
"s2_concept": 0.0003,
|
| 1157 |
+
"s3_rules": 0.0,
|
| 1158 |
+
"s4_arith": -0.0016,
|
| 1159 |
+
"s5_causal": 0.0,
|
| 1160 |
+
"s7_mixed": -0.0001,
|
| 1161 |
+
"s8_register": -0.0005
|
| 1162 |
+
},
|
| 1163 |
+
"alone_on_partner_stages": {
|
| 1164 |
+
"s1_perspective": 0.0029,
|
| 1165 |
+
"s2_concept": 0.0013,
|
| 1166 |
+
"s3_rules": -0.0003,
|
| 1167 |
+
"s4_arith": -0.001,
|
| 1168 |
+
"s5_causal": -0.0026,
|
| 1169 |
+
"s7_mixed": 0.0001,
|
| 1170 |
+
"s8_register": -0.0001
|
| 1171 |
+
},
|
| 1172 |
+
"steps_trained": 4000
|
| 1173 |
+
}
|
| 1174 |
+
},
|
| 1175 |
+
"examples": [
|
| 1176 |
+
"Talia adds 34 and 23 and writes 65. Talia checks by counting back: 65 is too big. Talia tries again carefully: 34 + 23 = 57. The check works now, and Talia keeps the good method."
|
| 1177 |
+
]
|
| 1178 |
+
},
|
| 1179 |
+
{
|
| 1180 |
+
"id": "stages-1-8/s7_mixed",
|
| 1181 |
+
"name": "stages-1-8/s7_mixed",
|
| 1182 |
+
"group": "stages-1-8",
|
| 1183 |
+
"base_model_id": "alephllm/mini-beatrix-3@step245674",
|
| 1184 |
+
"step": 245674,
|
| 1185 |
+
"params": 13692960,
|
| 1186 |
+
"sites": 32,
|
| 1187 |
+
"weights": [
|
| 1188 |
+
"arms/stages-1-8/s7_mixed.safetensors"
|
| 1189 |
+
],
|
| 1190 |
+
"mount": {
|
| 1191 |
+
"type": "anchor"
|
| 1192 |
+
},
|
| 1193 |
+
"precision": "fp32",
|
| 1194 |
+
"template": {
|
| 1195 |
+
"frame": "raw",
|
| 1196 |
+
"template": "raw stage text, continued"
|
| 1197 |
+
},
|
| 1198 |
+
"decode": {
|
| 1199 |
+
"temperature": 0.0,
|
| 1200 |
+
"top_p": 1.0,
|
| 1201 |
+
"max_new": 120,
|
| 1202 |
+
"precision": "fp32"
|
| 1203 |
+
},
|
| 1204 |
+
"source": {
|
| 1205 |
+
"file": "gXA_close8_s7_mixed.safetensors",
|
| 1206 |
+
"sha256": "b7a4210d7ab9690a5fbddf847abbe60a47755d9d8446dbbb04fe6ba469e8f23c"
|
| 1207 |
+
},
|
| 1208 |
+
"recipe": "trained as an always-on group on the frozen final core by continuing the four-arm group's run. The four arms of stages 1 to 4 were loaded as trained (see the group stages-1-4) and held fixed and on. Then a new arm for stage 5 was attached over them and trained for 4,000 steps in their presence; then, with it also held fixed, a new arm for stage 6 the same way, then stage 7, then stage 8. Each step gives the training arm 2 x 4096 bytes of its own stage's text with every attached arm on, and a quiet term at weight 2 on 2 rows that are not its text (one of web text, one of the stage text of another arm of the eight, drawn from every other stage whether or not its arm was attached yet; the mixed stage's text, which contains the other stages' kinds, is drawn as no arm's partner row): the KL from the model with that arm switched off to the model with every arm on. Pure Adam, lr 0.001",
|
| 1209 |
+
"title": "Stage 7: mixed",
|
| 1210 |
+
"behavior": "the mixed stage's diet: the earlier stages' text beside ordinary prose",
|
| 1211 |
+
"family": "stage",
|
| 1212 |
+
"status": "a member of the group stages-1-8; trained always on with its partners, so alone is not its trained state",
|
| 1213 |
+
"score": "alone: own stage text 0.783 -> 0.724 bpb; web text -0.0002 bpb",
|
| 1214 |
+
"measured": {
|
| 1215 |
+
"own_stage_bpb_arms_off": 0.783,
|
| 1216 |
+
"own_stage_bpb_group_on": 0.72,
|
| 1217 |
+
"own_stage_bpb_this_arm_alone": 0.7242,
|
| 1218 |
+
"share_of_the_group_on_own_stage": -0.0568,
|
| 1219 |
+
"share_of_the_group_on_own_stage_se": 0.0328,
|
| 1220 |
+
"share_of_the_group_on_web_text": -0.00096,
|
| 1221 |
+
"alone_on_web_text": -0.00019,
|
| 1222 |
+
"share_of_the_group_on_partner_stages": {
|
| 1223 |
+
"s1_perspective": -0.0004,
|
| 1224 |
+
"s2_concept": -0.0005,
|
| 1225 |
+
"s3_rules": -0.001,
|
| 1226 |
+
"s4_arith": -0.0103,
|
| 1227 |
+
"s5_causal": -0.0002,
|
| 1228 |
+
"s6_tryfail": -0.0004,
|
| 1229 |
+
"s8_register": -0.0001
|
| 1230 |
+
},
|
| 1231 |
+
"alone_on_partner_stages": {
|
| 1232 |
+
"s1_perspective": -0.0568,
|
| 1233 |
+
"s2_concept": -0.0058,
|
| 1234 |
+
"s3_rules": -0.0617,
|
| 1235 |
+
"s4_arith": -0.0385,
|
| 1236 |
+
"s5_causal": -0.1028,
|
| 1237 |
+
"s6_tryfail": -0.0802,
|
| 1238 |
+
"s8_register": -0.008
|
| 1239 |
+
},
|
| 1240 |
+
"steps_trained": 4000,
|
| 1241 |
+
"second_seed": {
|
| 1242 |
+
"own_stage_bpb_arms_off": 0.783,
|
| 1243 |
+
"own_stage_bpb_group_on": 0.7198,
|
| 1244 |
+
"own_stage_bpb_this_arm_alone": 0.7254,
|
| 1245 |
+
"share_of_the_group_on_own_stage": -0.0569,
|
| 1246 |
+
"share_of_the_group_on_own_stage_se": 0.0329,
|
| 1247 |
+
"share_of_the_group_on_web_text": -0.0007,
|
| 1248 |
+
"alone_on_web_text": -0.00032,
|
| 1249 |
+
"share_of_the_group_on_partner_stages": {
|
| 1250 |
+
"s1_perspective": 0.0001,
|
| 1251 |
+
"s2_concept": -0.0003,
|
| 1252 |
+
"s3_rules": -0.0007,
|
| 1253 |
+
"s4_arith": -0.0099,
|
| 1254 |
+
"s5_causal": 0.0,
|
| 1255 |
+
"s6_tryfail": -0.0003,
|
| 1256 |
+
"s8_register": -0.0002
|
| 1257 |
+
},
|
| 1258 |
+
"alone_on_partner_stages": {
|
| 1259 |
+
"s1_perspective": -0.0237,
|
| 1260 |
+
"s2_concept": -0.0022,
|
| 1261 |
+
"s3_rules": -0.0529,
|
| 1262 |
+
"s4_arith": -0.0414,
|
| 1263 |
+
"s5_causal": -0.0976,
|
| 1264 |
+
"s6_tryfail": -0.0774,
|
| 1265 |
+
"s8_register": -0.0061
|
| 1266 |
+
},
|
| 1267 |
+
"steps_trained": 4000
|
| 1268 |
+
}
|
| 1269 |
+
},
|
| 1270 |
+
"examples": [
|
| 1271 |
+
"It has been compared to a chant, a rhythmic divine beauty, a melody, an aria, a toccata, an edification, an exaltation. As poetry is for the tongue, calligraphy is to the page. The authors of The Splendor of Islamic Calligraphy put it best when they said: “Calligraphy is the plainsong of the divine”.\nCalligraphy is the art of the linear graphic, but it is more than that. In Isl"
|
| 1272 |
+
]
|
| 1273 |
+
},
|
| 1274 |
+
{
|
| 1275 |
+
"id": "stages-1-8/s8_register",
|
| 1276 |
+
"name": "stages-1-8/s8_register",
|
| 1277 |
+
"group": "stages-1-8",
|
| 1278 |
+
"base_model_id": "alephllm/mini-beatrix-3@step245674",
|
| 1279 |
+
"step": 245674,
|
| 1280 |
+
"params": 13692960,
|
| 1281 |
+
"sites": 32,
|
| 1282 |
+
"weights": [
|
| 1283 |
+
"arms/stages-1-8/s8_register.safetensors"
|
| 1284 |
+
],
|
| 1285 |
+
"mount": {
|
| 1286 |
+
"type": "anchor"
|
| 1287 |
+
},
|
| 1288 |
+
"precision": "fp32",
|
| 1289 |
+
"template": {
|
| 1290 |
+
"frame": "raw",
|
| 1291 |
+
"template": "raw stage text, continued"
|
| 1292 |
+
},
|
| 1293 |
+
"decode": {
|
| 1294 |
+
"temperature": 0.0,
|
| 1295 |
+
"top_p": 1.0,
|
| 1296 |
+
"max_new": 120,
|
| 1297 |
+
"precision": "fp32"
|
| 1298 |
+
},
|
| 1299 |
+
"source": {
|
| 1300 |
+
"file": "gXA_close8_s8_register.safetensors",
|
| 1301 |
+
"sha256": "19b02efab5fb482cd50d10508cda56868d8bb1f7fce835a0b7e9babec182aa3b"
|
| 1302 |
+
},
|
| 1303 |
+
"recipe": "trained as an always-on group on the frozen final core by continuing the four-arm group's run. The four arms of stages 1 to 4 were loaded as trained (see the group stages-1-4) and held fixed and on. Then a new arm for stage 5 was attached over them and trained for 4,000 steps in their presence; then, with it also held fixed, a new arm for stage 6 the same way, then stage 7, then stage 8. Each step gives the training arm 2 x 4096 bytes of its own stage's text with every attached arm on, and a quiet term at weight 2 on 2 rows that are not its text (one of web text, one of the stage text of another arm of the eight, drawn from every other stage whether or not its arm was attached yet; the mixed stage's text, which contains the other stages' kinds, is drawn as no arm's partner row): the KL from the model with that arm switched off to the model with every arm on. Pure Adam, lr 0.001",
|
| 1304 |
+
"title": "Stage 8: register",
|
| 1305 |
+
"behavior": "the same content in different registers of speech",
|
| 1306 |
+
"family": "stage",
|
| 1307 |
+
"status": "a member of the group stages-1-8; trained always on with its partners, so alone is not its trained state",
|
| 1308 |
+
"score": "alone: own stage text 0.491 -> 0.068 bpb; web text +0.0005 bpb",
|
| 1309 |
+
"measured": {
|
| 1310 |
+
"own_stage_bpb_arms_off": 0.4907,
|
| 1311 |
+
"own_stage_bpb_group_on": 0.0212,
|
| 1312 |
+
"own_stage_bpb_this_arm_alone": 0.0678,
|
| 1313 |
+
"share_of_the_group_on_own_stage": -0.3546,
|
| 1314 |
+
"share_of_the_group_on_own_stage_se": 0.0066,
|
| 1315 |
+
"share_of_the_group_on_web_text": 0.00045,
|
| 1316 |
+
"alone_on_web_text": 0.00048,
|
| 1317 |
+
"share_of_the_group_on_partner_stages": {
|
| 1318 |
+
"s1_perspective": 0.0007,
|
| 1319 |
+
"s2_concept": -0.0008,
|
| 1320 |
+
"s3_rules": 0.0005,
|
| 1321 |
+
"s4_arith": 0.0001,
|
| 1322 |
+
"s5_causal": 0.0002,
|
| 1323 |
+
"s6_tryfail": 0.0002,
|
| 1324 |
+
"s7_mixed": 0.0012
|
| 1325 |
+
},
|
| 1326 |
+
"alone_on_partner_stages": {
|
| 1327 |
+
"s1_perspective": -0.0172,
|
| 1328 |
+
"s2_concept": -0.0079,
|
| 1329 |
+
"s3_rules": -0.0048,
|
| 1330 |
+
"s4_arith": -0.0065,
|
| 1331 |
+
"s5_causal": -0.0189,
|
| 1332 |
+
"s6_tryfail": -0.022,
|
| 1333 |
+
"s7_mixed": -0.0002
|
| 1334 |
+
},
|
| 1335 |
+
"steps_trained": 4000,
|
| 1336 |
+
"second_seed": {
|
| 1337 |
+
"own_stage_bpb_arms_off": 0.4907,
|
| 1338 |
+
"own_stage_bpb_group_on": 0.0208,
|
| 1339 |
+
"own_stage_bpb_this_arm_alone": 0.0516,
|
| 1340 |
+
"share_of_the_group_on_own_stage": -0.3573,
|
| 1341 |
+
"share_of_the_group_on_own_stage_se": 0.0067,
|
| 1342 |
+
"share_of_the_group_on_web_text": 0.00046,
|
| 1343 |
+
"alone_on_web_text": 5e-05,
|
| 1344 |
+
"share_of_the_group_on_partner_stages": {
|
| 1345 |
+
"s1_perspective": 0.0001,
|
| 1346 |
+
"s2_concept": 0.0001,
|
| 1347 |
+
"s3_rules": 0.0005,
|
| 1348 |
+
"s4_arith": 0.0001,
|
| 1349 |
+
"s5_causal": 0.0002,
|
| 1350 |
+
"s6_tryfail": 0.0001,
|
| 1351 |
+
"s7_mixed": 0.0002
|
| 1352 |
+
},
|
| 1353 |
+
"alone_on_partner_stages": {
|
| 1354 |
+
"s1_perspective": -0.0045,
|
| 1355 |
+
"s2_concept": -0.005,
|
| 1356 |
+
"s3_rules": -0.0063,
|
| 1357 |
+
"s4_arith": -0.0064,
|
| 1358 |
+
"s5_causal": -0.0139,
|
| 1359 |
+
"s6_tryfail": -0.0187,
|
| 1360 |
+
"s7_mixed": 0.0001
|
| 1361 |
+
},
|
| 1362 |
+
"steps_trained": 4000
|
| 1363 |
+
}
|
| 1364 |
+
},
|
| 1365 |
+
"examples": [
|
| 1366 |
+
"The word 'ancient' means from a very long time ago. Used in a sentence: a ancient is easy to point at once you know the word."
|
| 1367 |
+
]
|
| 1368 |
+
},
|
| 1369 |
+
{
|
| 1370 |
+
"id": "stages-1-8",
|
| 1371 |
+
"name": "stages-1-8",
|
| 1372 |
+
"base_model_id": "alephllm/mini-beatrix-3@step245674",
|
| 1373 |
+
"step": 245674,
|
| 1374 |
+
"params": 109543680,
|
| 1375 |
+
"sites": 32,
|
| 1376 |
+
"weights": [
|
| 1377 |
+
"arms/stages-1-8/s1_perspective.safetensors",
|
| 1378 |
+
"arms/stages-1-8/s2_concept.safetensors",
|
| 1379 |
+
"arms/stages-1-8/s3_rules.safetensors",
|
| 1380 |
+
"arms/stages-1-8/s4_arith.safetensors",
|
| 1381 |
+
"arms/stages-1-8/s5_causal.safetensors",
|
| 1382 |
+
"arms/stages-1-8/s6_tryfail.safetensors",
|
| 1383 |
+
"arms/stages-1-8/s7_mixed.safetensors",
|
| 1384 |
+
"arms/stages-1-8/s8_register.safetensors"
|
| 1385 |
+
],
|
| 1386 |
+
"mount": {
|
| 1387 |
+
"type": "stack",
|
| 1388 |
+
"members": [
|
| 1389 |
+
"stages-1-8/s1_perspective",
|
| 1390 |
+
"stages-1-8/s2_concept",
|
| 1391 |
+
"stages-1-8/s3_rules",
|
| 1392 |
+
"stages-1-8/s4_arith",
|
| 1393 |
+
"stages-1-8/s5_causal",
|
| 1394 |
+
"stages-1-8/s6_tryfail",
|
| 1395 |
+
"stages-1-8/s7_mixed",
|
| 1396 |
+
"stages-1-8/s8_register"
|
| 1397 |
+
]
|
| 1398 |
+
},
|
| 1399 |
+
"precision": "fp32",
|
| 1400 |
+
"template": {
|
| 1401 |
+
"frame": "raw",
|
| 1402 |
+
"template": "raw text, continued"
|
| 1403 |
+
},
|
| 1404 |
+
"decode": {
|
| 1405 |
+
"temperature": 0.0,
|
| 1406 |
+
"top_p": 1.0,
|
| 1407 |
+
"max_new": 120,
|
| 1408 |
+
"precision": "fp32"
|
| 1409 |
+
},
|
| 1410 |
+
"recipe": "trained as an always-on group on the frozen final core by continuing the four-arm group's run. The four arms of stages 1 to 4 were loaded as trained (see the group stages-1-4) and held fixed and on. Then a new arm for stage 5 was attached over them and trained for 4,000 steps in their presence; then, with it also held fixed, a new arm for stage 6 the same way, then stage 7, then stage 8. Each step gives the training arm 2 x 4096 bytes of its own stage's text with every attached arm on, and a quiet term at weight 2 on 2 rows that are not its text (one of web text, one of the stage text of another arm of the eight, drawn from every other stage whether or not its arm was attached yet; the mixed stage's text, which contains the other stages' kinds, is drawn as no arm's partner row): the KL from the model with that arm switched off to the model with every arm on. Pure Adam, lr 0.001",
|
| 1411 |
+
"title": "Stages 1 to 8, on together",
|
| 1412 |
+
"behavior": "the arms of stage 1: perspective, stage 2: concepts, stage 3: rule chains, stage 4: arithmetic, stage 5: cause and effect, stage 6: try, fail, revise, stage 7: mixed, stage 8: register, always on together, as they were trained",
|
| 1413 |
+
"family": "group",
|
| 1414 |
+
"status": "two seeds",
|
| 1415 |
+
"score": "stage text, arms off -> on: s1_perspective 0.554 -> 0.049; s2_concept 0.721 -> 0.111; s3_rules 0.930 -> 0.224; s4_arith 1.234 -> 0.328; s5_causal 0.660 -> 0.043; s6_tryfail 0.673 -> 0.055; s7_mixed 0.783 -> 0.720; s8_register 0.491 -> 0.021; web text +0.0019 bpb",
|
| 1416 |
+
"measured": {
|
| 1417 |
+
"stage_text_bpb": {
|
| 1418 |
+
"s1_perspective": {
|
| 1419 |
+
"arms_off": 0.5538,
|
| 1420 |
+
"group_on": 0.0492,
|
| 1421 |
+
"difference": -0.5046,
|
| 1422 |
+
"difference_se": 0.0048
|
| 1423 |
+
},
|
| 1424 |
+
"s2_concept": {
|
| 1425 |
+
"arms_off": 0.7212,
|
| 1426 |
+
"group_on": 0.1105,
|
| 1427 |
+
"difference": -0.6107,
|
| 1428 |
+
"difference_se": 0.0038
|
| 1429 |
+
},
|
| 1430 |
+
"s3_rules": {
|
| 1431 |
+
"arms_off": 0.9296,
|
| 1432 |
+
"group_on": 0.2237,
|
| 1433 |
+
"difference": -0.7059,
|
| 1434 |
+
"difference_se": 0.0053
|
| 1435 |
+
},
|
| 1436 |
+
"s4_arith": {
|
| 1437 |
+
"arms_off": 1.2339,
|
| 1438 |
+
"group_on": 0.3279,
|
| 1439 |
+
"difference": -0.9059,
|
| 1440 |
+
"difference_se": 0.0062
|
| 1441 |
+
},
|
| 1442 |
+
"s5_causal": {
|
| 1443 |
+
"arms_off": 0.6604,
|
| 1444 |
+
"group_on": 0.0431,
|
| 1445 |
+
"difference": -0.6174,
|
| 1446 |
+
"difference_se": 0.0069
|
| 1447 |
+
},
|
| 1448 |
+
"s6_tryfail": {
|
| 1449 |
+
"arms_off": 0.6728,
|
| 1450 |
+
"group_on": 0.0548,
|
| 1451 |
+
"difference": -0.618,
|
| 1452 |
+
"difference_se": 0.0037
|
| 1453 |
+
},
|
| 1454 |
+
"s7_mixed": {
|
| 1455 |
+
"arms_off": 0.783,
|
| 1456 |
+
"group_on": 0.72,
|
| 1457 |
+
"difference": -0.063,
|
| 1458 |
+
"difference_se": 0.039
|
| 1459 |
+
},
|
| 1460 |
+
"s8_register": {
|
| 1461 |
+
"arms_off": 0.4907,
|
| 1462 |
+
"group_on": 0.0212,
|
| 1463 |
+
"difference": -0.4695,
|
| 1464 |
+
"difference_se": 0.0085
|
| 1465 |
+
}
|
| 1466 |
+
},
|
| 1467 |
+
"web_text_difference": 0.00189,
|
| 1468 |
+
"web_text_difference_se": 0.00032,
|
| 1469 |
+
"steps_trained": {
|
| 1470 |
+
"s1_perspective": 3200,
|
| 1471 |
+
"s2_concept": 2400,
|
| 1472 |
+
"s3_rules": 800,
|
| 1473 |
+
"s4_arith": 800,
|
| 1474 |
+
"s5_causal": 4000,
|
| 1475 |
+
"s6_tryfail": 4000,
|
| 1476 |
+
"s7_mixed": 4000,
|
| 1477 |
+
"s8_register": 4000
|
| 1478 |
+
},
|
| 1479 |
+
"probe_suite_mean": {
|
| 1480 |
+
"arms_off": 0.411,
|
| 1481 |
+
"group_on": 0.422
|
| 1482 |
+
},
|
| 1483 |
+
"second_seed": {
|
| 1484 |
+
"stage_text_bpb": {
|
| 1485 |
+
"s1_perspective": {
|
| 1486 |
+
"arms_off": 0.5538,
|
| 1487 |
+
"group_on": 0.0484,
|
| 1488 |
+
"difference": -0.5055,
|
| 1489 |
+
"difference_se": 0.0048
|
| 1490 |
+
},
|
| 1491 |
+
"s2_concept": {
|
| 1492 |
+
"arms_off": 0.7212,
|
| 1493 |
+
"group_on": 0.1212,
|
| 1494 |
+
"difference": -0.6,
|
| 1495 |
+
"difference_se": 0.0038
|
| 1496 |
+
},
|
| 1497 |
+
"s3_rules": {
|
| 1498 |
+
"arms_off": 0.9296,
|
| 1499 |
+
"group_on": 0.2222,
|
| 1500 |
+
"difference": -0.7074,
|
| 1501 |
+
"difference_se": 0.0055
|
| 1502 |
+
},
|
| 1503 |
+
"s4_arith": {
|
| 1504 |
+
"arms_off": 1.2339,
|
| 1505 |
+
"group_on": 0.3266,
|
| 1506 |
+
"difference": -0.9073,
|
| 1507 |
+
"difference_se": 0.0062
|
| 1508 |
+
},
|
| 1509 |
+
"s5_causal": {
|
| 1510 |
+
"arms_off": 0.6604,
|
| 1511 |
+
"group_on": 0.0428,
|
| 1512 |
+
"difference": -0.6176,
|
| 1513 |
+
"difference_se": 0.0069
|
| 1514 |
+
},
|
| 1515 |
+
"s6_tryfail": {
|
| 1516 |
+
"arms_off": 0.6728,
|
| 1517 |
+
"group_on": 0.0543,
|
| 1518 |
+
"difference": -0.6185,
|
| 1519 |
+
"difference_se": 0.0036
|
| 1520 |
+
},
|
| 1521 |
+
"s7_mixed": {
|
| 1522 |
+
"arms_off": 0.783,
|
| 1523 |
+
"group_on": 0.7198,
|
| 1524 |
+
"difference": -0.0632,
|
| 1525 |
+
"difference_se": 0.039
|
| 1526 |
+
},
|
| 1527 |
+
"s8_register": {
|
| 1528 |
+
"arms_off": 0.4907,
|
| 1529 |
+
"group_on": 0.0208,
|
| 1530 |
+
"difference": -0.4699,
|
| 1531 |
+
"difference_se": 0.0086
|
| 1532 |
+
}
|
| 1533 |
+
},
|
| 1534 |
+
"web_text_difference": 0.00103,
|
| 1535 |
+
"web_text_difference_se": 0.0002,
|
| 1536 |
+
"steps_trained": {
|
| 1537 |
+
"s1_perspective": 3200,
|
| 1538 |
+
"s2_concept": 2400,
|
| 1539 |
+
"s3_rules": 800,
|
| 1540 |
+
"s4_arith": 800,
|
| 1541 |
+
"s5_causal": 4000,
|
| 1542 |
+
"s6_tryfail": 4000,
|
| 1543 |
+
"s7_mixed": 4000,
|
| 1544 |
+
"s8_register": 4000
|
| 1545 |
+
},
|
| 1546 |
+
"probe_suite_mean": {
|
| 1547 |
+
"arms_off": 0.411,
|
| 1548 |
+
"group_on": 0.43
|
| 1549 |
+
}
|
| 1550 |
+
}
|
| 1551 |
+
},
|
| 1552 |
+
"examples": []
|
| 1553 |
+
},
|
| 1554 |
{
|
| 1555 |
"id": "solo/s1_perspective",
|
| 1556 |
"name": "solo/s1_perspective",
|
|
|
|
| 1874 |
"examples": [
|
| 1875 |
"Count by 3s: 1 3s make 3; 2 3s make 6; 3 3s make 9; 4 3s make 12; 5 3s make 15. So 3 times 5 is 15."
|
| 1876 |
]
|
| 1877 |
+
},
|
| 1878 |
+
{
|
| 1879 |
+
"id": "solo/s5_causal",
|
| 1880 |
+
"name": "solo/s5_causal",
|
| 1881 |
+
"group": null,
|
| 1882 |
+
"base_model_id": "alephllm/mini-beatrix-3@step245674",
|
| 1883 |
+
"step": 245674,
|
| 1884 |
+
"params": 13692960,
|
| 1885 |
+
"sites": 32,
|
| 1886 |
+
"weights": [
|
| 1887 |
+
"arms/solo/s5_causal.safetensors"
|
| 1888 |
+
],
|
| 1889 |
+
"mount": {
|
| 1890 |
+
"type": "anchor"
|
| 1891 |
+
},
|
| 1892 |
+
"precision": "fp32",
|
| 1893 |
+
"template": {
|
| 1894 |
+
"frame": "raw",
|
| 1895 |
+
"template": "raw stage text, continued"
|
| 1896 |
+
},
|
| 1897 |
+
"decode": {
|
| 1898 |
+
"temperature": 0.0,
|
| 1899 |
+
"top_p": 1.0,
|
| 1900 |
+
"max_new": 120,
|
| 1901 |
+
"precision": "fp32"
|
| 1902 |
+
},
|
| 1903 |
+
"source": {
|
| 1904 |
+
"file": "s5_causal_fresh_seed5_d20261005_step800.safetensors",
|
| 1905 |
+
"sha256": "b4369a46c13a3c2b425255e030a9c2b851fa3cf50046394dcadfd4f88fdc2a6a"
|
| 1906 |
+
},
|
| 1907 |
+
"recipe": "trained alone on the frozen final core: 800 steps of 2 x 4096 bytes of the stage's own text, pure Adam (lr 0.001), with a quiet term at weight 2 on held-out web text (the KL from the model with the arm off to the model with it on, one chunk beside every task chunk)",
|
| 1908 |
+
"title": "Stage 5: cause and effect (solo)",
|
| 1909 |
+
"behavior": "short cause-and-effect text: what happened and why",
|
| 1910 |
+
"family": "solo",
|
| 1911 |
+
"status": "one seed; trained alone: mount one solo arm at a time",
|
| 1912 |
+
"score": "own stage text 0.660 -> 0.068 bpb; web text +0.0003 bpb",
|
| 1913 |
+
"measured": {
|
| 1914 |
+
"steps": 800,
|
| 1915 |
+
"own_stage_bpb_arms_off": 0.6604,
|
| 1916 |
+
"own_stage_bpb_arm_on": 0.0681,
|
| 1917 |
+
"own_stage_difference": -0.5924,
|
| 1918 |
+
"own_stage_difference_se": 0.0068,
|
| 1919 |
+
"web_text_difference": 0.00032,
|
| 1920 |
+
"web_text_difference_se": 0.00017,
|
| 1921 |
+
"probe_suite_mean": {
|
| 1922 |
+
"arms_off": 0.411,
|
| 1923 |
+
"arm_on": 0.407
|
| 1924 |
+
}
|
| 1925 |
+
},
|
| 1926 |
+
"examples": [
|
| 1927 |
+
"Noor stacked the cups too high. Because of that, the tower leaned. Because of that, the cups crashed down."
|
| 1928 |
+
]
|
| 1929 |
+
},
|
| 1930 |
+
{
|
| 1931 |
+
"id": "solo/s6_tryfail",
|
| 1932 |
+
"name": "solo/s6_tryfail",
|
| 1933 |
+
"group": null,
|
| 1934 |
+
"base_model_id": "alephllm/mini-beatrix-3@step245674",
|
| 1935 |
+
"step": 245674,
|
| 1936 |
+
"params": 13692960,
|
| 1937 |
+
"sites": 32,
|
| 1938 |
+
"weights": [
|
| 1939 |
+
"arms/solo/s6_tryfail.safetensors"
|
| 1940 |
+
],
|
| 1941 |
+
"mount": {
|
| 1942 |
+
"type": "anchor"
|
| 1943 |
+
},
|
| 1944 |
+
"precision": "fp32",
|
| 1945 |
+
"template": {
|
| 1946 |
+
"frame": "raw",
|
| 1947 |
+
"template": "raw stage text, continued"
|
| 1948 |
+
},
|
| 1949 |
+
"decode": {
|
| 1950 |
+
"temperature": 0.0,
|
| 1951 |
+
"top_p": 1.0,
|
| 1952 |
+
"max_new": 120,
|
| 1953 |
+
"precision": "fp32"
|
| 1954 |
+
},
|
| 1955 |
+
"source": {
|
| 1956 |
+
"file": "s6_tryfail_fresh_seed6_d20261005_step800.safetensors",
|
| 1957 |
+
"sha256": "35db26313e181e40be63f5fca7cba883dd25cc88fd9c71e150d1a67363959bbf"
|
| 1958 |
+
},
|
| 1959 |
+
"recipe": "trained alone on the frozen final core: 800 steps of 2 x 4096 bytes of the stage's own text, pure Adam (lr 0.001), with a quiet term at weight 2 on held-out web text (the KL from the model with the arm off to the model with it on, one chunk beside every task chunk)",
|
| 1960 |
+
"title": "Stage 6: try, fail, revise (solo)",
|
| 1961 |
+
"behavior": "an attempt, its failure and the revised attempt",
|
| 1962 |
+
"family": "solo",
|
| 1963 |
+
"status": "one seed; trained alone: mount one solo arm at a time",
|
| 1964 |
+
"score": "own stage text 0.673 -> 0.114 bpb; web text +0.0008 bpb",
|
| 1965 |
+
"measured": {
|
| 1966 |
+
"steps": 800,
|
| 1967 |
+
"own_stage_bpb_arms_off": 0.6728,
|
| 1968 |
+
"own_stage_bpb_arm_on": 0.1142,
|
| 1969 |
+
"own_stage_difference": -0.5587,
|
| 1970 |
+
"own_stage_difference_se": 0.0045,
|
| 1971 |
+
"web_text_difference": 0.00079,
|
| 1972 |
+
"web_text_difference_se": 0.00016,
|
| 1973 |
+
"probe_suite_mean": {
|
| 1974 |
+
"arms_off": 0.411,
|
| 1975 |
+
"arm_on": 0.4
|
| 1976 |
+
}
|
| 1977 |
+
},
|
| 1978 |
+
"examples": [
|
| 1979 |
+
"Talia adds 34 and 23 and writes 65. Talia checks by counting back: 65 is too big. Talia tries again carefully: 34 + 23 = 57. The check works now, and Talia keeps the good method."
|
| 1980 |
+
]
|
| 1981 |
+
},
|
| 1982 |
+
{
|
| 1983 |
+
"id": "solo/s7_mixed",
|
| 1984 |
+
"name": "solo/s7_mixed",
|
| 1985 |
+
"group": null,
|
| 1986 |
+
"base_model_id": "alephllm/mini-beatrix-3@step245674",
|
| 1987 |
+
"step": 245674,
|
| 1988 |
+
"params": 13692960,
|
| 1989 |
+
"sites": 32,
|
| 1990 |
+
"weights": [
|
| 1991 |
+
"arms/solo/s7_mixed.safetensors"
|
| 1992 |
+
],
|
| 1993 |
+
"mount": {
|
| 1994 |
+
"type": "anchor"
|
| 1995 |
+
},
|
| 1996 |
+
"precision": "fp32",
|
| 1997 |
+
"template": {
|
| 1998 |
+
"frame": "raw",
|
| 1999 |
+
"template": "raw stage text, continued"
|
| 2000 |
+
},
|
| 2001 |
+
"decode": {
|
| 2002 |
+
"temperature": 0.0,
|
| 2003 |
+
"top_p": 1.0,
|
| 2004 |
+
"max_new": 120,
|
| 2005 |
+
"precision": "fp32"
|
| 2006 |
+
},
|
| 2007 |
+
"source": {
|
| 2008 |
+
"file": "s7_mixed_fresh_seed7_d20261005_step800.safetensors",
|
| 2009 |
+
"sha256": "6dc2c0c2a436a6052d8e82180a17248c04f8ef55822b5374f6f12f555268c412"
|
| 2010 |
+
},
|
| 2011 |
+
"recipe": "trained alone on the frozen final core: 800 steps of 2 x 4096 bytes of the stage's own text, pure Adam (lr 0.001), with a quiet term at weight 2 on held-out web text (the KL from the model with the arm off to the model with it on, one chunk beside every task chunk)",
|
| 2012 |
+
"title": "Stage 7: mixed (solo)",
|
| 2013 |
+
"behavior": "the mixed stage's diet: the earlier stages' text beside ordinary prose",
|
| 2014 |
+
"family": "solo",
|
| 2015 |
+
"status": "one seed; trained alone: mount one solo arm at a time",
|
| 2016 |
+
"score": "own stage text 0.783 -> 0.755 bpb; web text +0.0004 bpb",
|
| 2017 |
+
"measured": {
|
| 2018 |
+
"steps": 800,
|
| 2019 |
+
"own_stage_bpb_arms_off": 0.783,
|
| 2020 |
+
"own_stage_bpb_arm_on": 0.755,
|
| 2021 |
+
"own_stage_difference": -0.028,
|
| 2022 |
+
"own_stage_difference_se": 0.018,
|
| 2023 |
+
"web_text_difference": 0.00039,
|
| 2024 |
+
"web_text_difference_se": 0.00013,
|
| 2025 |
+
"probe_suite_mean": {
|
| 2026 |
+
"arms_off": 0.411,
|
| 2027 |
+
"arm_on": 0.378
|
| 2028 |
+
}
|
| 2029 |
+
},
|
| 2030 |
+
"examples": [
|
| 2031 |
+
"It has been compared to a chant, a rhythmic divine beauty, a melody, an aria, a toccata, an edification, an exaltation. As poetry is for the tongue, calligraphy is to the page. The authors of The Splendor of Islamic Calligraphy put it best when they said: “Calligraphy is the plainsong of the divine”.\nCalligraphy is the art of the linear graphic, but it is more than that. In Isl"
|
| 2032 |
+
]
|
| 2033 |
+
},
|
| 2034 |
+
{
|
| 2035 |
+
"id": "solo/s8_register",
|
| 2036 |
+
"name": "solo/s8_register",
|
| 2037 |
+
"group": null,
|
| 2038 |
+
"base_model_id": "alephllm/mini-beatrix-3@step245674",
|
| 2039 |
+
"step": 245674,
|
| 2040 |
+
"params": 13692960,
|
| 2041 |
+
"sites": 32,
|
| 2042 |
+
"weights": [
|
| 2043 |
+
"arms/solo/s8_register.safetensors"
|
| 2044 |
+
],
|
| 2045 |
+
"mount": {
|
| 2046 |
+
"type": "anchor"
|
| 2047 |
+
},
|
| 2048 |
+
"precision": "fp32",
|
| 2049 |
+
"template": {
|
| 2050 |
+
"frame": "raw",
|
| 2051 |
+
"template": "raw stage text, continued"
|
| 2052 |
+
},
|
| 2053 |
+
"decode": {
|
| 2054 |
+
"temperature": 0.0,
|
| 2055 |
+
"top_p": 1.0,
|
| 2056 |
+
"max_new": 120,
|
| 2057 |
+
"precision": "fp32"
|
| 2058 |
+
},
|
| 2059 |
+
"source": {
|
| 2060 |
+
"file": "s8_register_fresh_seed8_d20261005_step800.safetensors",
|
| 2061 |
+
"sha256": "a5ae8506d4752cf9d619e789e396f2b379317e20316ad32a02f5ee0af541b4c1"
|
| 2062 |
+
},
|
| 2063 |
+
"recipe": "trained alone on the frozen final core: 800 steps of 2 x 4096 bytes of the stage's own text, pure Adam (lr 0.001), with a quiet term at weight 2 on held-out web text (the KL from the model with the arm off to the model with it on, one chunk beside every task chunk)",
|
| 2064 |
+
"title": "Stage 8: register (solo)",
|
| 2065 |
+
"behavior": "the same content in different registers of speech",
|
| 2066 |
+
"family": "solo",
|
| 2067 |
+
"status": "one seed; trained alone: mount one solo arm at a time",
|
| 2068 |
+
"score": "own stage text 0.491 -> 0.051 bpb; web text +0.0008 bpb",
|
| 2069 |
+
"measured": {
|
| 2070 |
+
"steps": 800,
|
| 2071 |
+
"own_stage_bpb_arms_off": 0.4907,
|
| 2072 |
+
"own_stage_bpb_arm_on": 0.051,
|
| 2073 |
+
"own_stage_difference": -0.4396,
|
| 2074 |
+
"own_stage_difference_se": 0.0078,
|
| 2075 |
+
"web_text_difference": 0.00075,
|
| 2076 |
+
"web_text_difference_se": 9e-05,
|
| 2077 |
+
"probe_suite_mean": {
|
| 2078 |
+
"arms_off": 0.411,
|
| 2079 |
+
"arm_on": 0.378
|
| 2080 |
+
}
|
| 2081 |
+
},
|
| 2082 |
+
"examples": [
|
| 2083 |
+
"The word 'ancient' means from a very long time ago. Used in a sentence: a ancient is easy to point at once you know the word."
|
| 2084 |
+
]
|
| 2085 |
}
|
| 2086 |
]
|
| 2087 |
}
|
arms/solo/s3_rules.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 54817680
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:e068430f1da2f0de1d6f703ccf552ee9633ad5ad02754fef90d4fa924cba9144
|
| 3 |
size 54817680
|
arms/solo/s4_arith.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 54817552
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:17f398b0d0100e48cd8ff5c3aba10c9ee2d57f769c94914d40c6bc1f396396b0
|
| 3 |
size 54817552
|
arms/solo/s5_causal.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:b871df2812ce5bdbd61df53e0c4000543ded6685c198e436ddffb5e76dea58f7
|
| 3 |
+
size 54816968
|
arms/solo/s6_tryfail.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:e53255c4a1748cc950da6c577ace85edd7650170c88acca019c00d2d31395e5a
|
| 3 |
+
size 54817040
|
arms/solo/s7_mixed.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:310baca6f2ab7fbfb3e1070e7c233871d9f5452c83d0062432b0d27ca4b3504f
|
| 3 |
+
size 54817256
|
arms/solo/s8_register.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:4ce9d6a1bb0aa19b8ad8b20deaa4f036d74792f0d62af8c4a96e3dbd84091b02
|
| 3 |
+
size 54816984
|
arms/stages-1-4/s1_perspective.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 54819000
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:bcf835c13de8bd0b446e65f21aad9515c2cab769b701936165c143815091b321
|
| 3 |
size 54819000
|
arms/stages-1-4/s2_concept.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 54818880
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:3aa85cc0047416499b5f00288875803899dec2ec897d4f834b307ff64cdc6fb6
|
| 3 |
size 54818880
|
arms/stages-1-4/s3_rules.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 54818928
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:e409aa5684064ff03c54bb6136ffde682df5b067cf191715ab95e44748a120c3
|
| 3 |
size 54818928
|
arms/stages-1-4/s4_arith.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 54818792
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:a85c82a5d02951a81882146d2df38523e4bddd7a07910eaeba23a99ae86c1ce5
|
| 3 |
size 54818792
|
arms/stages-1-8/s1_perspective.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:7faa377430e9be0b4600683ea4c63ecc8a3c84e7d536f3df830324065b76f97d
|
| 3 |
+
size 54819424
|
arms/stages-1-8/s2_concept.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:e305183f35aa6adb76b8bee4c57d802f377acd2b0ad92b1f7578cd2de8618ad5
|
| 3 |
+
size 54819304
|
arms/stages-1-8/s3_rules.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:6069cac6f21653ab9da229b2daf25f2271633603fb60bf99f04bfc3acfc73ee9
|
| 3 |
+
size 54819344
|
arms/stages-1-8/s4_arith.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:6e29b2fcc48ae351a519a783bed637f282597c16bddb903f412948dd89ac4102
|
| 3 |
+
size 54819216
|
arms/stages-1-8/s5_causal.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:00e6374b72a19253a3130001daf4c67fddc761497ae519e8d3d4c0a38eac56e1
|
| 3 |
+
size 54818832
|
arms/stages-1-8/s6_tryfail.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:da74a8218506923537757416610e94f075d72615a30aaaba35f3bc5e9dedffd8
|
| 3 |
+
size 54818896
|
arms/stages-1-8/s7_mixed.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:b54beaf81c1f6f31dbc61ed312cf1e36a068f334f222da8c6152208ae741c9b2
|
| 3 |
+
size 54819136
|
arms/stages-1-8/s8_register.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:c55b64e9c6a8084cadafca177fff3b1079c0747a4cf8542a2037e4fe18eb6ff0
|
| 3 |
+
size 54818840
|