AbstractPhil commited on
Commit
8541601
·
verified ·
1 Parent(s): efe3aa0

mini-beatrix-3: the stage arms (arms off and arms on in one package)

Browse files
README.md CHANGED
@@ -181,14 +181,14 @@ whole stage by itself.
181
 
182
  | arm | its stage's text | stage text, arm off → on (bpb) | web text (bpb) | stage items in the stage's own form | stage items in a new form | seeds |
183
  |---|---|---|---|---|---|---|
184
- | `solo/s1_perspective` | one small event retold from each side: seen from a person, said to them, told about them | 0.554 → 0.081 (a second seed: 0.080) | +0.0002 | 55% → 75% | 46% → 41% | two |
185
- | `solo/s2_concept` | kinds, properties and differences: what a thing is a kind of, what its kind can do, how two things differ | 0.721 → 0.120 (a second seed: 0.120) | +0.0003 | 30% → 84% | 68% → 73% | two |
186
- | `solo/s3_rules` | if-then rules over made-up words, followed step by step to what follows | 0.930 → 0.372 (a second seed: 0.374) | +0.0009 | 16% → 60% | 7% → 12% | two |
187
- | `solo/s4_arith` | small arithmetic worked out in text | 1.234 → 0.389 (a second seed: 0.388) | +0.0007 | 34% → 49% | 24% → 24% | two |
188
- | `solo/s5_causal` | short cause-and-effect text: what happened and why | 0.660 → 0.068 | +0.0003 | not measured | not measured | one |
189
- | `solo/s6_tryfail` | an attempt, its failure and the revised attempt | 0.673 → 0.114 | +0.0008 | not measured | not measured | one |
190
- | `solo/s7_mixed` | the mixed stage's diet: the earlier stages' text beside ordinary prose | 0.783 → 0.755 | +0.0004 | not measured | not measured | one |
191
- | `solo/s8_register` | the same content in different registers of speech | 0.491 → 0.051 | +0.0008 | not measured | not measured | one |
192
 
193
  ```python
194
  m.mount_arm("solo/s1_perspective") # one solo arm; mounting another swaps it
@@ -296,7 +296,7 @@ trunk for each stage; those arms stayed nearly empty, because the trunk took
296
  each stage's text in before its arm could, and they were detached at step
297
  148,000 (their files are in the training repo under `mini-beatrix-3/arms/`,
298
  bound to those earlier trunk states). The arms in this repository were fitted
299
- afterwards on the finished, frozen core: as always-on groups and as solo arms. The four-arm group in two steps (first the arms of stages 1 and 2, in a run that added the four stages' text one stage at a time with every attached arm training under one pure Adam, 800 steps per stage; then, with those two held fixed and on, a new arm for stage 3 and after it a new arm for stage 4, 800 steps each). The eight-arm group by continuing that run: with the four held fixed and on, a new arm for stage 5, then 6, then 7, then 8, each attached over every arm before it and trained for 4,000 steps in their presence. Every group arm was kept quiet on ordinary web text and on the stage text of every other arm in its group. Each solo arm was trained alone, 800 steps, kept quiet on web text.
300
 
301
  ## Lineage
302
 
 
181
 
182
  | arm | its stage's text | stage text, arm off → on (bpb) | web text (bpb) | stage items in the stage's own form | stage items in a new form | seeds |
183
  |---|---|---|---|---|---|---|
184
+ | `solo/s1_perspective` | one small event retold from each side: seen from a person, said to them, told about them | 0.554 → 0.047 (a second seed: 0.047) | +0.0003 | 55% → 89% | 46% → 42% | two |
185
+ | `solo/s2_concept` | kinds, properties and differences: what a thing is a kind of, what its kind can do, how two things differ | 0.721 → 0.075 (a second seed: 0.075) | +0.0001 | 30% → 100% | 68% → 70% | two |
186
+ | `solo/s3_rules` | if-then rules over made-up words, followed step by step to what follows | 0.930 → 0.180 (a second seed: 0.179) | +0.0003 | 16% → 82% | 7% → 22% | two |
187
+ | `solo/s4_arith` | small arithmetic worked out in text | 1.234 → 0.267 (a second seed: 0.267) | +0.0002 | 34% → 83% | 24% → 21% | two |
188
+ | `solo/s5_causal` | short cause-and-effect text: what happened and why | 0.660 → 0.043 (a second seed: 0.043) | +0.0006 | not measured | not measured | two |
189
+ | `solo/s6_tryfail` | an attempt, its failure and the revised attempt | 0.673 → 0.057 (a second seed: 0.057) | +0.0003 | not measured | not measured | two |
190
+ | `solo/s7_mixed` | the mixed stage's diet: the earlier stages' text beside ordinary prose | 0.783 → 0.715 (a second seed: 0.717) | +0.0000 | not measured | not measured | two |
191
+ | `solo/s8_register` | the same content in different registers of speech | 0.491 → 0.021 (a second seed: 0.021) | +0.0003 | not measured | not measured | two |
192
 
193
  ```python
194
  m.mount_arm("solo/s1_perspective") # one solo arm; mounting another swaps it
 
296
  each stage's text in before its arm could, and they were detached at step
297
  148,000 (their files are in the training repo under `mini-beatrix-3/arms/`,
298
  bound to those earlier trunk states). The arms in this repository were fitted
299
+ afterwards on the finished, frozen core: as always-on groups and as solo arms. The four-arm group in two steps (first the arms of stages 1 and 2, in a run that added the four stages' text one stage at a time with every attached arm training under one pure Adam, 800 steps per stage; then, with those two held fixed and on, a new arm for stage 3 and after it a new arm for stage 4, 800 steps each). The eight-arm group by continuing that run: with the four held fixed and on, a new arm for stage 5, then 6, then 7, then 8, each attached over every arm before it and trained for 4,000 steps in their presence. Every group arm was kept quiet on ordinary web text and on the stage text of every other arm in its group. Each solo arm was trained alone for 4,000 steps (two seeds), kept quiet on web text.
300
 
301
  ## Lineage
302
 
arms/index.json CHANGED
@@ -6,8 +6,8 @@
6
  "base_model_id": "alephllm/mini-beatrix-3@step245674",
7
  "note": "every packaged arm was trained on this exact frozen core (the default bf16 weight file); arms are trunk-bound"
8
  },
9
- "generated": "2026-10-06 04:54 UTC",
10
- "note": "the stage arms of curriculum stages 1 to 8, fitted on the finished, frozen core: two always-on groups (the four arms of stages 1 to 4; the eight arms of stages 1 to 8, which are those four, held fixed, with four more fitted over them) and each stage's arm alone; the arms trained beside the trunk during the run are on the training repo",
11
  "arms": [
12
  {
13
  "id": "stages-1-4/s1_perspective",
@@ -1577,54 +1577,54 @@
1577
  "precision": "fp32"
1578
  },
1579
  "source": {
1580
- "file": "s1_perspective_fresh_seed1_d20261005_step800.safetensors",
1581
- "sha256": "ff72dd5a8061926b56e09c6da765c1d1ee44cbbd3d4e5e2c30793e04b369532a"
1582
  },
1583
- "recipe": "trained alone on the frozen final core: 800 steps of 2 x 4096 bytes of the stage's own text, pure Adam (lr 0.001), with a quiet term at weight 2 on held-out web text (the KL from the model with the arm off to the model with it on, one chunk beside every task chunk)",
1584
  "title": "Stage 1: perspective (solo)",
1585
  "behavior": "one small event retold from each side: seen from a person, said to them, told about them",
1586
  "family": "solo",
1587
  "status": "two seeds; trained alone: mount one solo arm at a time",
1588
- "score": "own stage text 0.554 -> 0.081 bpb; web text +0.0002 bpb",
1589
  "measured": {
1590
- "steps": 800,
1591
  "own_stage_bpb_arms_off": 0.5538,
1592
- "own_stage_bpb_arm_on": 0.0809,
1593
- "own_stage_difference": -0.4729,
1594
- "own_stage_difference_se": 0.0038,
1595
- "web_text_difference": 0.00018,
1596
- "web_text_difference_se": 0.00019,
1597
  "stage_items_own_form_accuracy": {
1598
  "arms_off": 0.55,
1599
- "arm_on": 0.75
1600
  },
1601
  "stage_items_new_form_accuracy": {
1602
  "arms_off": 0.463,
1603
- "arm_on": 0.412
1604
  },
1605
  "probe_suite_mean": {
1606
  "arms_off": 0.411,
1607
- "arm_on": 0.381
1608
  },
1609
  "second_seed": {
1610
- "steps": 800,
1611
  "own_stage_bpb_arms_off": 0.5538,
1612
- "own_stage_bpb_arm_on": 0.0802,
1613
- "own_stage_difference": -0.4737,
1614
- "own_stage_difference_se": 0.0038,
1615
- "web_text_difference": 0.00036,
1616
- "web_text_difference_se": 0.00016,
1617
  "stage_items_own_form_accuracy": {
1618
  "arms_off": 0.55,
1619
- "arm_on": 0.7
1620
  },
1621
  "stage_items_new_form_accuracy": {
1622
  "arms_off": 0.463,
1623
- "arm_on": 0.4
1624
  },
1625
  "probe_suite_mean": {
1626
  "arms_off": 0.411,
1627
- "arm_on": 0.385
1628
  }
1629
  }
1630
  },
@@ -1658,54 +1658,54 @@
1658
  "precision": "fp32"
1659
  },
1660
  "source": {
1661
- "file": "s2_concept_fresh_seed2_d20261005_step800.safetensors",
1662
- "sha256": "dc76f7c7d713b34d23b553adbb0b6302ec1d1080d7ffe8cc9461c5f5e6a03938"
1663
  },
1664
- "recipe": "trained alone on the frozen final core: 800 steps of 2 x 4096 bytes of the stage's own text, pure Adam (lr 0.001), with a quiet term at weight 2 on held-out web text (the KL from the model with the arm off to the model with it on, one chunk beside every task chunk)",
1665
  "title": "Stage 2: concepts (solo)",
1666
  "behavior": "kinds, properties and differences: what a thing is a kind of, what its kind can do, how two things differ",
1667
  "family": "solo",
1668
  "status": "two seeds; trained alone: mount one solo arm at a time",
1669
- "score": "own stage text 0.721 -> 0.120 bpb; web text +0.0003 bpb",
1670
  "measured": {
1671
- "steps": 800,
1672
  "own_stage_bpb_arms_off": 0.7212,
1673
- "own_stage_bpb_arm_on": 0.1204,
1674
- "own_stage_difference": -0.6008,
1675
- "own_stage_difference_se": 0.0034,
1676
- "web_text_difference": 0.00029,
1677
- "web_text_difference_se": 0.00016,
1678
  "stage_items_own_form_accuracy": {
1679
  "arms_off": 0.298,
1680
- "arm_on": 0.845
1681
  },
1682
  "stage_items_new_form_accuracy": {
1683
  "arms_off": 0.678,
1684
- "arm_on": 0.729
1685
  },
1686
  "probe_suite_mean": {
1687
  "arms_off": 0.411,
1688
- "arm_on": 0.404
1689
  },
1690
  "second_seed": {
1691
- "steps": 800,
1692
  "own_stage_bpb_arms_off": 0.7212,
1693
- "own_stage_bpb_arm_on": 0.1199,
1694
- "own_stage_difference": -0.6013,
1695
- "own_stage_difference_se": 0.0035,
1696
- "web_text_difference": 0.00043,
1697
- "web_text_difference_se": 0.00018,
1698
  "stage_items_own_form_accuracy": {
1699
  "arms_off": 0.298,
1700
- "arm_on": 0.881
1701
  },
1702
  "stage_items_new_form_accuracy": {
1703
  "arms_off": 0.678,
1704
- "arm_on": 0.695
1705
  },
1706
  "probe_suite_mean": {
1707
  "arms_off": 0.411,
1708
- "arm_on": 0.389
1709
  }
1710
  }
1711
  },
@@ -1739,54 +1739,54 @@
1739
  "precision": "fp32"
1740
  },
1741
  "source": {
1742
- "file": "s3_rules_fresh_seed3_d20261005_step800.safetensors",
1743
- "sha256": "11bc779e8853bab9ecb6cdff3419673e39a59c55fc91e8842aaef6ffc23c3b47"
1744
  },
1745
- "recipe": "trained alone on the frozen final core: 800 steps of 2 x 4096 bytes of the stage's own text, pure Adam (lr 0.001), with a quiet term at weight 2 on held-out web text (the KL from the model with the arm off to the model with it on, one chunk beside every task chunk)",
1746
  "title": "Stage 3: rule chains (solo)",
1747
  "behavior": "if-then rules over made-up words, followed step by step to what follows",
1748
  "family": "solo",
1749
  "status": "two seeds; trained alone: mount one solo arm at a time",
1750
- "score": "own stage text 0.930 -> 0.372 bpb; web text +0.0009 bpb",
1751
  "measured": {
1752
- "steps": 800,
1753
  "own_stage_bpb_arms_off": 0.9296,
1754
- "own_stage_bpb_arm_on": 0.3724,
1755
- "own_stage_difference": -0.5572,
1756
- "own_stage_difference_se": 0.0054,
1757
- "web_text_difference": 0.00095,
1758
- "web_text_difference_se": 0.00011,
1759
  "stage_items_own_form_accuracy": {
1760
  "arms_off": 0.155,
1761
- "arm_on": 0.6
1762
  },
1763
  "stage_items_new_form_accuracy": {
1764
  "arms_off": 0.07,
1765
- "arm_on": 0.12
1766
  },
1767
  "probe_suite_mean": {
1768
  "arms_off": 0.411,
1769
- "arm_on": 0.407
1770
  },
1771
  "second_seed": {
1772
- "steps": 800,
1773
  "own_stage_bpb_arms_off": 0.9296,
1774
- "own_stage_bpb_arm_on": 0.3737,
1775
- "own_stage_difference": -0.5559,
1776
- "own_stage_difference_se": 0.0054,
1777
- "web_text_difference": 0.00079,
1778
- "web_text_difference_se": 0.00015,
1779
  "stage_items_own_form_accuracy": {
1780
  "arms_off": 0.155,
1781
- "arm_on": 0.58
1782
  },
1783
  "stage_items_new_form_accuracy": {
1784
  "arms_off": 0.07,
1785
- "arm_on": 0.12
1786
  },
1787
  "probe_suite_mean": {
1788
  "arms_off": 0.411,
1789
- "arm_on": 0.4
1790
  }
1791
  }
1792
  },
@@ -1820,54 +1820,54 @@
1820
  "precision": "fp32"
1821
  },
1822
  "source": {
1823
- "file": "s4_arith_fresh_seed4_d20261005_step800.safetensors",
1824
- "sha256": "90774072859c74888b7e83298867cddbb7f12c7c9b9e9206a417c08545a6244d"
1825
  },
1826
- "recipe": "trained alone on the frozen final core: 800 steps of 2 x 4096 bytes of the stage's own text, pure Adam (lr 0.001), with a quiet term at weight 2 on held-out web text (the KL from the model with the arm off to the model with it on, one chunk beside every task chunk)",
1827
  "title": "Stage 4: arithmetic (solo)",
1828
  "behavior": "small arithmetic worked out in text",
1829
  "family": "solo",
1830
  "status": "two seeds; trained alone: mount one solo arm at a time",
1831
- "score": "own stage text 1.234 -> 0.389 bpb; web text +0.0007 bpb",
1832
  "measured": {
1833
- "steps": 800,
1834
  "own_stage_bpb_arms_off": 1.2339,
1835
- "own_stage_bpb_arm_on": 0.3889,
1836
- "own_stage_difference": -0.845,
1837
- "own_stage_difference_se": 0.0048,
1838
- "web_text_difference": 0.00069,
1839
- "web_text_difference_se": 0.00023,
1840
  "stage_items_own_form_accuracy": {
1841
  "arms_off": 0.344,
1842
- "arm_on": 0.487
1843
  },
1844
  "stage_items_new_form_accuracy": {
1845
  "arms_off": 0.242,
1846
- "arm_on": 0.242
1847
  },
1848
  "probe_suite_mean": {
1849
  "arms_off": 0.411,
1850
- "arm_on": 0.389
1851
  },
1852
  "second_seed": {
1853
- "steps": 800,
1854
  "own_stage_bpb_arms_off": 1.2339,
1855
- "own_stage_bpb_arm_on": 0.3884,
1856
- "own_stage_difference": -0.8455,
1857
- "own_stage_difference_se": 0.005,
1858
- "web_text_difference": 0.00056,
1859
- "web_text_difference_se": 0.00024,
1860
  "stage_items_own_form_accuracy": {
1861
  "arms_off": 0.344,
1862
- "arm_on": 0.487
1863
  },
1864
  "stage_items_new_form_accuracy": {
1865
  "arms_off": 0.242,
1866
- "arm_on": 0.208
1867
  },
1868
  "probe_suite_mean": {
1869
  "arms_off": 0.411,
1870
- "arm_on": 0.4
1871
  }
1872
  }
1873
  },
@@ -1901,26 +1901,39 @@
1901
  "precision": "fp32"
1902
  },
1903
  "source": {
1904
- "file": "s5_causal_fresh_seed5_d20261005_step800.safetensors",
1905
- "sha256": "b4369a46c13a3c2b425255e030a9c2b851fa3cf50046394dcadfd4f88fdc2a6a"
1906
  },
1907
- "recipe": "trained alone on the frozen final core: 800 steps of 2 x 4096 bytes of the stage's own text, pure Adam (lr 0.001), with a quiet term at weight 2 on held-out web text (the KL from the model with the arm off to the model with it on, one chunk beside every task chunk)",
1908
  "title": "Stage 5: cause and effect (solo)",
1909
  "behavior": "short cause-and-effect text: what happened and why",
1910
  "family": "solo",
1911
- "status": "one seed; trained alone: mount one solo arm at a time",
1912
- "score": "own stage text 0.660 -> 0.068 bpb; web text +0.0003 bpb",
1913
  "measured": {
1914
- "steps": 800,
1915
  "own_stage_bpb_arms_off": 0.6604,
1916
- "own_stage_bpb_arm_on": 0.0681,
1917
- "own_stage_difference": -0.5924,
1918
- "own_stage_difference_se": 0.0068,
1919
- "web_text_difference": 0.00032,
1920
- "web_text_difference_se": 0.00017,
1921
  "probe_suite_mean": {
1922
  "arms_off": 0.411,
1923
- "arm_on": 0.407
 
 
 
 
 
 
 
 
 
 
 
 
 
1924
  }
1925
  },
1926
  "examples": [
@@ -1953,26 +1966,39 @@
1953
  "precision": "fp32"
1954
  },
1955
  "source": {
1956
- "file": "s6_tryfail_fresh_seed6_d20261005_step800.safetensors",
1957
- "sha256": "35db26313e181e40be63f5fca7cba883dd25cc88fd9c71e150d1a67363959bbf"
1958
  },
1959
- "recipe": "trained alone on the frozen final core: 800 steps of 2 x 4096 bytes of the stage's own text, pure Adam (lr 0.001), with a quiet term at weight 2 on held-out web text (the KL from the model with the arm off to the model with it on, one chunk beside every task chunk)",
1960
  "title": "Stage 6: try, fail, revise (solo)",
1961
  "behavior": "an attempt, its failure and the revised attempt",
1962
  "family": "solo",
1963
- "status": "one seed; trained alone: mount one solo arm at a time",
1964
- "score": "own stage text 0.673 -> 0.114 bpb; web text +0.0008 bpb",
1965
  "measured": {
1966
- "steps": 800,
1967
  "own_stage_bpb_arms_off": 0.6728,
1968
- "own_stage_bpb_arm_on": 0.1142,
1969
- "own_stage_difference": -0.5587,
1970
- "own_stage_difference_se": 0.0045,
1971
- "web_text_difference": 0.00079,
1972
- "web_text_difference_se": 0.00016,
1973
  "probe_suite_mean": {
1974
  "arms_off": 0.411,
1975
- "arm_on": 0.4
 
 
 
 
 
 
 
 
 
 
 
 
 
1976
  }
1977
  },
1978
  "examples": [
@@ -2005,26 +2031,39 @@
2005
  "precision": "fp32"
2006
  },
2007
  "source": {
2008
- "file": "s7_mixed_fresh_seed7_d20261005_step800.safetensors",
2009
- "sha256": "6dc2c0c2a436a6052d8e82180a17248c04f8ef55822b5374f6f12f555268c412"
2010
  },
2011
- "recipe": "trained alone on the frozen final core: 800 steps of 2 x 4096 bytes of the stage's own text, pure Adam (lr 0.001), with a quiet term at weight 2 on held-out web text (the KL from the model with the arm off to the model with it on, one chunk beside every task chunk)",
2012
  "title": "Stage 7: mixed (solo)",
2013
  "behavior": "the mixed stage's diet: the earlier stages' text beside ordinary prose",
2014
  "family": "solo",
2015
- "status": "one seed; trained alone: mount one solo arm at a time",
2016
- "score": "own stage text 0.783 -> 0.755 bpb; web text +0.0004 bpb",
2017
  "measured": {
2018
- "steps": 800,
2019
  "own_stage_bpb_arms_off": 0.783,
2020
- "own_stage_bpb_arm_on": 0.755,
2021
- "own_stage_difference": -0.028,
2022
- "own_stage_difference_se": 0.018,
2023
- "web_text_difference": 0.00039,
2024
- "web_text_difference_se": 0.00013,
2025
  "probe_suite_mean": {
2026
  "arms_off": 0.411,
2027
- "arm_on": 0.378
 
 
 
 
 
 
 
 
 
 
 
 
 
2028
  }
2029
  },
2030
  "examples": [
@@ -2057,26 +2096,39 @@
2057
  "precision": "fp32"
2058
  },
2059
  "source": {
2060
- "file": "s8_register_fresh_seed8_d20261005_step800.safetensors",
2061
- "sha256": "a5ae8506d4752cf9d619e789e396f2b379317e20316ad32a02f5ee0af541b4c1"
2062
  },
2063
- "recipe": "trained alone on the frozen final core: 800 steps of 2 x 4096 bytes of the stage's own text, pure Adam (lr 0.001), with a quiet term at weight 2 on held-out web text (the KL from the model with the arm off to the model with it on, one chunk beside every task chunk)",
2064
  "title": "Stage 8: register (solo)",
2065
  "behavior": "the same content in different registers of speech",
2066
  "family": "solo",
2067
- "status": "one seed; trained alone: mount one solo arm at a time",
2068
- "score": "own stage text 0.491 -> 0.051 bpb; web text +0.0008 bpb",
2069
  "measured": {
2070
- "steps": 800,
2071
  "own_stage_bpb_arms_off": 0.4907,
2072
- "own_stage_bpb_arm_on": 0.051,
2073
- "own_stage_difference": -0.4396,
2074
- "own_stage_difference_se": 0.0078,
2075
- "web_text_difference": 0.00075,
2076
- "web_text_difference_se": 9e-05,
2077
  "probe_suite_mean": {
2078
  "arms_off": 0.411,
2079
- "arm_on": 0.378
 
 
 
 
 
 
 
 
 
 
 
 
 
2080
  }
2081
  },
2082
  "examples": [
 
6
  "base_model_id": "alephllm/mini-beatrix-3@step245674",
7
  "note": "every packaged arm was trained on this exact frozen core (the default bf16 weight file); arms are trunk-bound"
8
  },
9
+ "generated": "2026-10-06 15:18 UTC",
10
+ "note": "the stage arms of curriculum stages 1 to 8, fitted on the finished, frozen core: two always-on groups (the four arms of stages 1 to 4; the eight arms of stages 1 to 8, which are those four, held fixed, with four more fitted over them) and each stage's arm alone (4,000 steps, two seeds); the arms trained beside the trunk during the run are on the training repo",
11
  "arms": [
12
  {
13
  "id": "stages-1-4/s1_perspective",
 
1577
  "precision": "fp32"
1578
  },
1579
  "source": {
1580
+ "file": "s1_perspective_fresh_seed1_d20261005_step4000.safetensors",
1581
+ "sha256": "13290e04d7b7f9e2f26d3b3809afcab54306f3aa7b1d74d73726836a32624177"
1582
  },
1583
+ "recipe": "trained alone on the frozen final core: 4000 steps of 2 x 4096 bytes of the stage's own text, pure Adam (lr 0.001), with a quiet term at weight 2 on held-out web text (the KL from the model with the arm off to the model with it on, one chunk beside every task chunk)",
1584
  "title": "Stage 1: perspective (solo)",
1585
  "behavior": "one small event retold from each side: seen from a person, said to them, told about them",
1586
  "family": "solo",
1587
  "status": "two seeds; trained alone: mount one solo arm at a time",
1588
+ "score": "own stage text 0.554 -> 0.047 bpb; web text +0.0003 bpb",
1589
  "measured": {
1590
+ "steps": 4000,
1591
  "own_stage_bpb_arms_off": 0.5538,
1592
+ "own_stage_bpb_arm_on": 0.0469,
1593
+ "own_stage_difference": -0.5069,
1594
+ "own_stage_difference_se": 0.0048,
1595
+ "web_text_difference": 0.00026,
1596
+ "web_text_difference_se": 0.00011,
1597
  "stage_items_own_form_accuracy": {
1598
  "arms_off": 0.55,
1599
+ "arm_on": 0.887
1600
  },
1601
  "stage_items_new_form_accuracy": {
1602
  "arms_off": 0.463,
1603
+ "arm_on": 0.425
1604
  },
1605
  "probe_suite_mean": {
1606
  "arms_off": 0.411,
1607
+ "arm_on": 0.411
1608
  },
1609
  "second_seed": {
1610
+ "steps": 4000,
1611
  "own_stage_bpb_arms_off": 0.5538,
1612
+ "own_stage_bpb_arm_on": 0.0467,
1613
+ "own_stage_difference": -0.5071,
1614
+ "own_stage_difference_se": 0.0048,
1615
+ "web_text_difference": 0.00031,
1616
+ "web_text_difference_se": 0.00013,
1617
  "stage_items_own_form_accuracy": {
1618
  "arms_off": 0.55,
1619
+ "arm_on": 1.0
1620
  },
1621
  "stage_items_new_form_accuracy": {
1622
  "arms_off": 0.463,
1623
+ "arm_on": 0.5
1624
  },
1625
  "probe_suite_mean": {
1626
  "arms_off": 0.411,
1627
+ "arm_on": 0.407
1628
  }
1629
  }
1630
  },
 
1658
  "precision": "fp32"
1659
  },
1660
  "source": {
1661
+ "file": "s2_concept_fresh_seed2_d20261005_step4000.safetensors",
1662
+ "sha256": "a07fbca8d7f171475c1b91509589199323dfed929205df2fb2c55dd8d7e1f507"
1663
  },
1664
+ "recipe": "trained alone on the frozen final core: 4000 steps of 2 x 4096 bytes of the stage's own text, pure Adam (lr 0.001), with a quiet term at weight 2 on held-out web text (the KL from the model with the arm off to the model with it on, one chunk beside every task chunk)",
1665
  "title": "Stage 2: concepts (solo)",
1666
  "behavior": "kinds, properties and differences: what a thing is a kind of, what its kind can do, how two things differ",
1667
  "family": "solo",
1668
  "status": "two seeds; trained alone: mount one solo arm at a time",
1669
+ "score": "own stage text 0.721 -> 0.075 bpb; web text +0.0001 bpb",
1670
  "measured": {
1671
+ "steps": 4000,
1672
  "own_stage_bpb_arms_off": 0.7212,
1673
+ "own_stage_bpb_arm_on": 0.0754,
1674
+ "own_stage_difference": -0.6458,
1675
+ "own_stage_difference_se": 0.0038,
1676
+ "web_text_difference": 0.00014,
1677
+ "web_text_difference_se": 0.00014,
1678
  "stage_items_own_form_accuracy": {
1679
  "arms_off": 0.298,
1680
+ "arm_on": 1.0
1681
  },
1682
  "stage_items_new_form_accuracy": {
1683
  "arms_off": 0.678,
1684
+ "arm_on": 0.695
1685
  },
1686
  "probe_suite_mean": {
1687
  "arms_off": 0.411,
1688
+ "arm_on": 0.389
1689
  },
1690
  "second_seed": {
1691
+ "steps": 4000,
1692
  "own_stage_bpb_arms_off": 0.7212,
1693
+ "own_stage_bpb_arm_on": 0.0753,
1694
+ "own_stage_difference": -0.6459,
1695
+ "own_stage_difference_se": 0.0036,
1696
+ "web_text_difference": 3e-05,
1697
+ "web_text_difference_se": 0.0001,
1698
  "stage_items_own_form_accuracy": {
1699
  "arms_off": 0.298,
1700
+ "arm_on": 0.988
1701
  },
1702
  "stage_items_new_form_accuracy": {
1703
  "arms_off": 0.678,
1704
+ "arm_on": 0.661
1705
  },
1706
  "probe_suite_mean": {
1707
  "arms_off": 0.411,
1708
+ "arm_on": 0.396
1709
  }
1710
  }
1711
  },
 
1739
  "precision": "fp32"
1740
  },
1741
  "source": {
1742
+ "file": "s3_rules_fresh_seed3_d20261005_step4000.safetensors",
1743
+ "sha256": "9bafbafaec5834fc1c1596ffbab4174e68fc238ed33440ecc07786e86478f996"
1744
  },
1745
+ "recipe": "trained alone on the frozen final core: 4000 steps of 2 x 4096 bytes of the stage's own text, pure Adam (lr 0.001), with a quiet term at weight 2 on held-out web text (the KL from the model with the arm off to the model with it on, one chunk beside every task chunk)",
1746
  "title": "Stage 3: rule chains (solo)",
1747
  "behavior": "if-then rules over made-up words, followed step by step to what follows",
1748
  "family": "solo",
1749
  "status": "two seeds; trained alone: mount one solo arm at a time",
1750
+ "score": "own stage text 0.930 -> 0.180 bpb; web text +0.0003 bpb",
1751
  "measured": {
1752
+ "steps": 4000,
1753
  "own_stage_bpb_arms_off": 0.9296,
1754
+ "own_stage_bpb_arm_on": 0.1804,
1755
+ "own_stage_difference": -0.7492,
1756
+ "own_stage_difference_se": 0.006,
1757
+ "web_text_difference": 0.00032,
1758
+ "web_text_difference_se": 0.00014,
1759
  "stage_items_own_form_accuracy": {
1760
  "arms_off": 0.155,
1761
+ "arm_on": 0.82
1762
  },
1763
  "stage_items_new_form_accuracy": {
1764
  "arms_off": 0.07,
1765
+ "arm_on": 0.22
1766
  },
1767
  "probe_suite_mean": {
1768
  "arms_off": 0.411,
1769
+ "arm_on": 0.415
1770
  },
1771
  "second_seed": {
1772
+ "steps": 4000,
1773
  "own_stage_bpb_arms_off": 0.9296,
1774
+ "own_stage_bpb_arm_on": 0.1795,
1775
+ "own_stage_difference": -0.7502,
1776
+ "own_stage_difference_se": 0.0059,
1777
+ "web_text_difference": 0.00052,
1778
+ "web_text_difference_se": 0.00016,
1779
  "stage_items_own_form_accuracy": {
1780
  "arms_off": 0.155,
1781
+ "arm_on": 0.825
1782
  },
1783
  "stage_items_new_form_accuracy": {
1784
  "arms_off": 0.07,
1785
+ "arm_on": 0.29
1786
  },
1787
  "probe_suite_mean": {
1788
  "arms_off": 0.411,
1789
+ "arm_on": 0.411
1790
  }
1791
  }
1792
  },
 
1820
  "precision": "fp32"
1821
  },
1822
  "source": {
1823
+ "file": "s4_arith_fresh_seed4_d20261005_step4000.safetensors",
1824
+ "sha256": "1bed0018b40f5de968ecfd2d846ad581b1b92d9444b18df593104051a941d4b4"
1825
  },
1826
+ "recipe": "trained alone on the frozen final core: 4000 steps of 2 x 4096 bytes of the stage's own text, pure Adam (lr 0.001), with a quiet term at weight 2 on held-out web text (the KL from the model with the arm off to the model with it on, one chunk beside every task chunk)",
1827
  "title": "Stage 4: arithmetic (solo)",
1828
  "behavior": "small arithmetic worked out in text",
1829
  "family": "solo",
1830
  "status": "two seeds; trained alone: mount one solo arm at a time",
1831
+ "score": "own stage text 1.234 -> 0.267 bpb; web text +0.0002 bpb",
1832
  "measured": {
1833
+ "steps": 4000,
1834
  "own_stage_bpb_arms_off": 1.2339,
1835
+ "own_stage_bpb_arm_on": 0.2671,
1836
+ "own_stage_difference": -0.9668,
1837
+ "own_stage_difference_se": 0.0071,
1838
+ "web_text_difference": 0.00023,
1839
+ "web_text_difference_se": 0.00016,
1840
  "stage_items_own_form_accuracy": {
1841
  "arms_off": 0.344,
1842
+ "arm_on": 0.831
1843
  },
1844
  "stage_items_new_form_accuracy": {
1845
  "arms_off": 0.242,
1846
+ "arm_on": 0.212
1847
  },
1848
  "probe_suite_mean": {
1849
  "arms_off": 0.411,
1850
+ "arm_on": 0.419
1851
  },
1852
  "second_seed": {
1853
+ "steps": 4000,
1854
  "own_stage_bpb_arms_off": 1.2339,
1855
+ "own_stage_bpb_arm_on": 0.2674,
1856
+ "own_stage_difference": -0.9664,
1857
+ "own_stage_difference_se": 0.0074,
1858
+ "web_text_difference": 0.00032,
1859
+ "web_text_difference_se": 0.00015,
1860
  "stage_items_own_form_accuracy": {
1861
  "arms_off": 0.344,
1862
+ "arm_on": 0.819
1863
  },
1864
  "stage_items_new_form_accuracy": {
1865
  "arms_off": 0.242,
1866
+ "arm_on": 0.258
1867
  },
1868
  "probe_suite_mean": {
1869
  "arms_off": 0.411,
1870
+ "arm_on": 0.411
1871
  }
1872
  }
1873
  },
 
1901
  "precision": "fp32"
1902
  },
1903
  "source": {
1904
+ "file": "s5_causal_fresh_seed5_d20261005_step4000.safetensors",
1905
+ "sha256": "f99713142a9ea2af225755055e46819cfc19ed9907781aa3f55973a89d9df48a"
1906
  },
1907
+ "recipe": "trained alone on the frozen final core: 4000 steps of 2 x 4096 bytes of the stage's own text, pure Adam (lr 0.001), with a quiet term at weight 2 on held-out web text (the KL from the model with the arm off to the model with it on, one chunk beside every task chunk)",
1908
  "title": "Stage 5: cause and effect (solo)",
1909
  "behavior": "short cause-and-effect text: what happened and why",
1910
  "family": "solo",
1911
+ "status": "two seeds; trained alone: mount one solo arm at a time",
1912
+ "score": "own stage text 0.660 -> 0.043 bpb; web text +0.0006 bpb",
1913
  "measured": {
1914
+ "steps": 4000,
1915
  "own_stage_bpb_arms_off": 0.6604,
1916
+ "own_stage_bpb_arm_on": 0.0428,
1917
+ "own_stage_difference": -0.6176,
1918
+ "own_stage_difference_se": 0.007,
1919
+ "web_text_difference": 0.0006,
1920
+ "web_text_difference_se": 0.00015,
1921
  "probe_suite_mean": {
1922
  "arms_off": 0.411,
1923
+ "arm_on": 0.415
1924
+ },
1925
+ "second_seed": {
1926
+ "steps": 4000,
1927
+ "own_stage_bpb_arms_off": 0.6604,
1928
+ "own_stage_bpb_arm_on": 0.0433,
1929
+ "own_stage_difference": -0.6171,
1930
+ "own_stage_difference_se": 0.007,
1931
+ "web_text_difference": 0.00041,
1932
+ "web_text_difference_se": 0.00015,
1933
+ "probe_suite_mean": {
1934
+ "arms_off": 0.411,
1935
+ "arm_on": 0.4
1936
+ }
1937
  }
1938
  },
1939
  "examples": [
 
1966
  "precision": "fp32"
1967
  },
1968
  "source": {
1969
+ "file": "s6_tryfail_fresh_seed6_d20261005_step4000.safetensors",
1970
+ "sha256": "1744513a39e8f6df4e8741391119d78caf5e2b0a681e8e3da5acabf19a3ac598"
1971
  },
1972
+ "recipe": "trained alone on the frozen final core: 4000 steps of 2 x 4096 bytes of the stage's own text, pure Adam (lr 0.001), with a quiet term at weight 2 on held-out web text (the KL from the model with the arm off to the model with it on, one chunk beside every task chunk)",
1973
  "title": "Stage 6: try, fail, revise (solo)",
1974
  "behavior": "an attempt, its failure and the revised attempt",
1975
  "family": "solo",
1976
+ "status": "two seeds; trained alone: mount one solo arm at a time",
1977
+ "score": "own stage text 0.673 -> 0.057 bpb; web text +0.0003 bpb",
1978
  "measured": {
1979
+ "steps": 4000,
1980
  "own_stage_bpb_arms_off": 0.6728,
1981
+ "own_stage_bpb_arm_on": 0.0568,
1982
+ "own_stage_difference": -0.6161,
1983
+ "own_stage_difference_se": 0.0035,
1984
+ "web_text_difference": 0.00034,
1985
+ "web_text_difference_se": 0.00022,
1986
  "probe_suite_mean": {
1987
  "arms_off": 0.411,
1988
+ "arm_on": 0.426
1989
+ },
1990
+ "second_seed": {
1991
+ "steps": 4000,
1992
+ "own_stage_bpb_arms_off": 0.6728,
1993
+ "own_stage_bpb_arm_on": 0.057,
1994
+ "own_stage_difference": -0.6158,
1995
+ "own_stage_difference_se": 0.0036,
1996
+ "web_text_difference": 0.00061,
1997
+ "web_text_difference_se": 0.0001,
1998
+ "probe_suite_mean": {
1999
+ "arms_off": 0.411,
2000
+ "arm_on": 0.433
2001
+ }
2002
  }
2003
  },
2004
  "examples": [
 
2031
  "precision": "fp32"
2032
  },
2033
  "source": {
2034
+ "file": "s7_mixed_fresh_seed7_d20261005_step4000.safetensors",
2035
+ "sha256": "e53aec3a88e82a5f7b69440867c9de348c3f07222398e31042922a7e38cb5263"
2036
  },
2037
+ "recipe": "trained alone on the frozen final core: 4000 steps of 2 x 4096 bytes of the stage's own text, pure Adam (lr 0.001), with a quiet term at weight 2 on held-out web text (the KL from the model with the arm off to the model with it on, one chunk beside every task chunk)",
2038
  "title": "Stage 7: mixed (solo)",
2039
  "behavior": "the mixed stage's diet: the earlier stages' text beside ordinary prose",
2040
  "family": "solo",
2041
+ "status": "two seeds; trained alone: mount one solo arm at a time",
2042
+ "score": "own stage text 0.783 -> 0.715 bpb; web text +0.0000 bpb",
2043
  "measured": {
2044
+ "steps": 4000,
2045
  "own_stage_bpb_arms_off": 0.783,
2046
+ "own_stage_bpb_arm_on": 0.7153,
2047
+ "own_stage_difference": -0.0677,
2048
+ "own_stage_difference_se": 0.0407,
2049
+ "web_text_difference": 2e-05,
2050
+ "web_text_difference_se": 0.00018,
2051
  "probe_suite_mean": {
2052
  "arms_off": 0.411,
2053
+ "arm_on": 0.426
2054
+ },
2055
+ "second_seed": {
2056
+ "steps": 4000,
2057
+ "own_stage_bpb_arms_off": 0.783,
2058
+ "own_stage_bpb_arm_on": 0.7171,
2059
+ "own_stage_difference": -0.0659,
2060
+ "own_stage_difference_se": 0.0392,
2061
+ "web_text_difference": 0.00017,
2062
+ "web_text_difference_se": 0.00017,
2063
+ "probe_suite_mean": {
2064
+ "arms_off": 0.411,
2065
+ "arm_on": 0.419
2066
+ }
2067
  }
2068
  },
2069
  "examples": [
 
2096
  "precision": "fp32"
2097
  },
2098
  "source": {
2099
+ "file": "s8_register_fresh_seed8_d20261005_step4000.safetensors",
2100
+ "sha256": "158f420f7caa8d8dd0f4dfb29206a8ba14e58bf600def9849da5049fec4380a1"
2101
  },
2102
+ "recipe": "trained alone on the frozen final core: 4000 steps of 2 x 4096 bytes of the stage's own text, pure Adam (lr 0.001), with a quiet term at weight 2 on held-out web text (the KL from the model with the arm off to the model with it on, one chunk beside every task chunk)",
2103
  "title": "Stage 8: register (solo)",
2104
  "behavior": "the same content in different registers of speech",
2105
  "family": "solo",
2106
+ "status": "two seeds; trained alone: mount one solo arm at a time",
2107
+ "score": "own stage text 0.491 -> 0.021 bpb; web text +0.0003 bpb",
2108
  "measured": {
2109
+ "steps": 4000,
2110
  "own_stage_bpb_arms_off": 0.4907,
2111
+ "own_stage_bpb_arm_on": 0.0208,
2112
+ "own_stage_difference": -0.4698,
2113
+ "own_stage_difference_se": 0.0086,
2114
+ "web_text_difference": 0.00026,
2115
+ "web_text_difference_se": 0.00017,
2116
  "probe_suite_mean": {
2117
  "arms_off": 0.411,
2118
+ "arm_on": 0.4
2119
+ },
2120
+ "second_seed": {
2121
+ "steps": 4000,
2122
+ "own_stage_bpb_arms_off": 0.4907,
2123
+ "own_stage_bpb_arm_on": 0.0209,
2124
+ "own_stage_difference": -0.4698,
2125
+ "own_stage_difference_se": 0.0085,
2126
+ "web_text_difference": 0.00019,
2127
+ "web_text_difference_se": 0.00018,
2128
+ "probe_suite_mean": {
2129
+ "arms_off": 0.411,
2130
+ "arm_on": 0.404
2131
+ }
2132
  }
2133
  },
2134
  "examples": [
arms/solo/s1_perspective.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:3c9d862b4b6832ab8fc6bccf177c87882a1b0ab78e91b2c01a952f7b0f9328cd
3
- size 54817760
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d4f21cbd0ec37c541e14272137751f3fbd42a895e9d79e803724be5f93926644
3
+ size 54817768
arms/solo/s2_concept.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:8a90104ddeace651612a0bf0015250cefbd104a29ece1cf6c064d4854cafa52d
3
  size 54817632
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a17d5fa9a3d587e56db796608117b0870c9538e5f26c2a4f07b1f8e27ceda61c
3
  size 54817632
arms/solo/s3_rules.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:e068430f1da2f0de1d6f703ccf552ee9633ad5ad02754fef90d4fa924cba9144
3
- size 54817680
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1cb73345f542acc0fe0e7156c4248e2b4526f7f5dc1139b9a91bdd233a51ea7c
3
+ size 54817688
arms/solo/s4_arith.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:17f398b0d0100e48cd8ff5c3aba10c9ee2d57f769c94914d40c6bc1f396396b0
3
- size 54817552
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f88431b8cd39c853d304210225b45b5fe6e89e7b467a9bbfe9e55de06520c8e4
3
+ size 54817560
arms/solo/s5_causal.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:b871df2812ce5bdbd61df53e0c4000543ded6685c198e436ddffb5e76dea58f7
3
- size 54816968
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:2646a2bdfa4c2186607610d25f934310c678287918a689c6540b55761847c602
3
+ size 54817280
arms/solo/s6_tryfail.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:e53255c4a1748cc950da6c577ace85edd7650170c88acca019c00d2d31395e5a
3
- size 54817040
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c9f9b9636b859504286f3a94ae5123c414d87ff80cc66326ae141bdfec012a4f
3
+ size 54817360
arms/solo/s7_mixed.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:310baca6f2ab7fbfb3e1070e7c233871d9f5452c83d0062432b0d27ca4b3504f
3
- size 54817256
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:35b55a2f859de67cd1268de19e317793c7490e8bdd78a8533ebd1f0ec7f67853
3
+ size 54817576
arms/solo/s8_register.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:4ce9d6a1bb0aa19b8ad8b20deaa4f036d74792f0d62af8c4a96e3dbd84091b02
3
- size 54816984
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:304662e2176f7d117cb867cd7b04c51e0cf3e439186776a880abb9f16454c4fc
3
+ size 54817304
arms/stages-1-4/s1_perspective.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:bcf835c13de8bd0b446e65f21aad9515c2cab769b701936165c143815091b321
3
  size 54819000
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:910854ff3ad93bcaf8164ae7031f5287318b8b779a2c919455e423ce30cb1716
3
  size 54819000
arms/stages-1-4/s2_concept.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:3aa85cc0047416499b5f00288875803899dec2ec897d4f834b307ff64cdc6fb6
3
  size 54818880
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c975352619adfda87e5d85d49cc058e623804271cff7ffabaffb00a315d60037
3
  size 54818880
arms/stages-1-4/s3_rules.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:e409aa5684064ff03c54bb6136ffde682df5b067cf191715ab95e44748a120c3
3
  size 54818928
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0b4643a43fa6b9b0967c6720d55cf0b1b80fc3259f6b1cc096ded2bc0935b0e2
3
  size 54818928
arms/stages-1-4/s4_arith.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:a85c82a5d02951a81882146d2df38523e4bddd7a07910eaeba23a99ae86c1ce5
3
  size 54818792
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:5be8dc7d6cfe7411eb732f791bfed9ee5ac1a0bf2243e993231deed8ab6f603f
3
  size 54818792
arms/stages-1-8/s1_perspective.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:7faa377430e9be0b4600683ea4c63ecc8a3c84e7d536f3df830324065b76f97d
3
  size 54819424
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:554e82d0381bc88cf5b0e3137490196bd038813f42b1b63934ff1deb1feecf2d
3
  size 54819424
arms/stages-1-8/s2_concept.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:e305183f35aa6adb76b8bee4c57d802f377acd2b0ad92b1f7578cd2de8618ad5
3
  size 54819304
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:419054af4ce5c32fdc7f97278edda3b3ab42369994fc40abbdab10ae0d65f427
3
  size 54819304
arms/stages-1-8/s3_rules.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:6069cac6f21653ab9da229b2daf25f2271633603fb60bf99f04bfc3acfc73ee9
3
  size 54819344
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6a0fc3d583cdfc60200d587941aa9da2a4371500c6e9539ff6d34ace18761aaa
3
  size 54819344
arms/stages-1-8/s4_arith.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:6e29b2fcc48ae351a519a783bed637f282597c16bddb903f412948dd89ac4102
3
  size 54819216
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:83e49028565f865ea156cb39e4bf77d32da25091b2e461d30a656d9281352d78
3
  size 54819216
arms/stages-1-8/s5_causal.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:00e6374b72a19253a3130001daf4c67fddc761497ae519e8d3d4c0a38eac56e1
3
  size 54818832
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:23aea6a2d5b0b31d7e341c30ad3adbaadf133360a3c2a58e8e0118bef0766250
3
  size 54818832
arms/stages-1-8/s6_tryfail.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:da74a8218506923537757416610e94f075d72615a30aaaba35f3bc5e9dedffd8
3
  size 54818896
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0f80f1b85339852c2fd01d033c93361e5d0f679ccda6cca000efe14cd803b254
3
  size 54818896
arms/stages-1-8/s7_mixed.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:b54beaf81c1f6f31dbc61ed312cf1e36a068f334f222da8c6152208ae741c9b2
3
  size 54819136
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:bc16d51b3b708010b14322a4b43ca90056c74e73de44f6fe20178974101c592b
3
  size 54819136
arms/stages-1-8/s8_register.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:c55b64e9c6a8084cadafca177fff3b1079c0747a4cf8542a2037e4fe18eb6ff0
3
  size 54818840
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c39df647b63e195268b28884ae0e9a66b2d7c2b1d8f69983e9486d9c20dacb1d
3
  size 54818840