Scaling Beatrix: A 376-Million-Parameter Byte Model, the Arms That Followed Her Trunk, and the First Sliders in Pictures
This installment covers a month, 2026-09-05 through 10-06, and one model built from everything the two before it taught. mini-beatrix-3 is a 376-million-parameter byte-level model (376,017,249 parameters): thirty-two blocks, every one of them the signed, softmax-free splat memory that the 2s proved against its softmax twin, with a bank of three experts and a dual head, trained on two RTX 5090 cards for 64.4 billion bytes in 295 hours and finished at 0.9861 bits per byte on held-out web text with no loss spike and the gradient clip never reached. The month has four parts. A twelve-day ladder of small experiments on the previous model settled how a byte model learns a behaviour through a detachable arm, and decided that the trunk trains the prior while capability lives in the arms. A pre-flight of screens priced the depth, certified the guards, shelved one idea as not ready and launched the run. The run itself was fifteen days with an engineering window in the middle, a silent fault caught and outlawed, and a finding that the arms trained beside the trunk were empty, so the arms were detached and refitted on the finished weights, where nine of them now ride the frozen trunk as a group. And on the picture side, with nothing of hers trained, Beatrix 3 was read as a text encoder, a theory about her last block was tested and supported, and a line of mood sliders ran on two image models, Sana and then Anima, to learn how a conditioning signal reaches a picture before her own states are asked to carry one. AbstractPhil — Phil, the program lead — directs and arbitrates throughout.
Reading guide
Narrative here; the ledgers behind every number live in the technical companion and the per-topic technical files listed at the end. Bits per byte (bpb) is the loss unit, lower is better, and every held-out number in the run is the fineweb-edu holdout unless a stage's own text is named. A toggle is the bpb cost of switching an organ off: the causal share of the splat memory (the hub), the bank or the head. An arm is a detachable adapter after every block, trained with the trunk frozen and quiet, meaning penalized for changing predictions on text outside its job; a group is several arms trained under one optimizer that stay separable. The atlas is the deterministic byte-trigram space that stands in for a tokenizer as the data authority. Where a result is called a law it has two seeds, a repeat and a control that could have failed; a candidate is one seed, one draw or a pre-registered line missed; a screen that fails is not ready, never refuted, unless the complete form was tested. On the picture side a slider is a learned direction added to a text encoder's states and scaled by a dial, and the conditioner is Beatrix read as that encoder. Retractions are stated where they fell.
Contents
- The ladder that designed her
- The design and the pre-flight
- The run
- What the memory remembers
- The arms after the trunk
- What ships
- Beatrix 3 as a conditioner
- Sliders on Sana and Anima
- The diffusion line from here
- How the record is kept
- What Beatrix 3 is for, and what comes next
The findings that matter most
In order of importance, not of date:
- A 376-million-parameter trunk with splat memory in all thirty-two blocks trained 64.4 billion bytes on two consumer cards and finished at 0.9861 bits per byte, with zero loss spikes and the gradient clip never reached. The anneals did their work as diet changes at a flat learning rate: 1.0780 at the curriculum's close to 0.9951 without the chat frame and 0.9861 with it. Depth bought loss at every rung of the ladder that priced it, 1.561 to 1.476 bits per byte from 16 to 32 blocks at a matched 1,145 steps, and the 32 blocks carry a rank plateau of 420 to 480 directions through the middle of the stack that is used, not idle.
- The arms trained beside the trunk were empty, and the arms fitted after it are full. Every stage's text was taken by the trunk first; co-trained arms carried nothing even with their gates forced open, while a fresh arm on frozen weights took 64% of a stage's gain in 400 steps. Detached at step 148,000 and refitted on the finished trunk, the eight stage arms read s1 .5538 → .0492 through s8 .4907 → .0212 bits per byte on their stages at a web-text cost of +.0019, two seeds, and detach bit for bit.
- Solo arms fail together; a staged group with a roster rule gives whole, separable members. Four solos switched on together read worse than no arm (stage 2 text .7193 → 1.7640); the fix the record already prescribed, lower arms frozen and every member quiet on its partners' text, brought each member's share of its own stage to 0.94 to 0.99, and routed dispatch over them read worse than the plain stack on every stage because one amplitude budget cannot serve text that wants several arms in full.
- Recall at a distance is what the size did not buy. A planted code came back at .753 of its digits at once and .031 about 3,500 bytes later, where the 2s's softmax twin held .588; the fade is the additive, never-erased splat memory's, present mid-pretraining and in the all-splat 2s. The next prototype is a hybrid with standard attention in the model's own weakest-rank blocks, eight to twelve of 32.
- The engineering window turned a November finish into October, and a silent fault became a standing rule. Each attached arm cost 3.06 gigabytes of stored intermediates and 1,280 kernel launches per forward; recomputation, a compiled arm chain whose non-finite gradient was traced to the aleph address and cured by emulating eager's rounding, and quartering the quiet chunks moved the finish from November 30 to October 14, and detaching the arms to October 5. At one stage close the compiled chain fell back to eager without a word, cost about half a day at 10.4 seconds a step, and now cannot: a guard relaunches it compiled.
- With nothing trained, Beatrix 3's middle blocks already carry caption geometry, and her last block carries one direction that is a per-byte temperature. Agreement with T5 reads .62 to .65 against .23 to .25 for an untrained copy; attribute binding .89 against the 2s's .71. One fixed direction holds 94 to 99 percent of each byte's energy at the last block; removing it costs +43.5 bits per byte, and a conditioner must project it out. Its per-byte size fluctuates with one self-similar exponent, .62 to .67 on structureless input against .48 to .51 for shuffled controls: the fractal theory, tested with its rules set before the run, supported at step 194,000. Rank is a floor, not a dial: a 6.6-fold rise in rank bought nothing measurable and a 10-fold rise cost +.13 to +.17.
- Mood is a dial in two image models' conditioning, but nothing driven by Beatrix has yet passed in pictures. On Sana the dial moved the judge +0.795 per unit against 14% of that for a random direction; on Anima the same push is dead before the adapter (+0.029) and alive after it (+0.314), and the adapter reads every caption twice, token ids as questions and Qwen3's states as answers, with the answer carrying the mood in pictures (+2.050 ± .370, 91% of cells). Beatrix's own sliders fell short of their registered rules five times, each time teaching the relay's design; the picture test of that relay, through her nine arms, is built, public and ahead.
- The ladder before the run fixed the design's constants. Two quiet arms compose where two plain ones read 0 (.860 with .973 against 0 and 0); a quiet weight of 2 meets the quiet bar on both seeds; teaching the trunk itself a behaviour destroyed the prior (+0.74 to +1.34 bits per byte) while leaving the behaviour at zero, so the third model is a steering-and-arms craft.
The body proceeds in date order, 2026-09-05 through 10-06; sections 7 to 9 run on the picture side in parallel with sections 5 and 6.
1. The ladder that designed her
(2026-09-05 to 09-16.) The previous article left two things on the table: the wall — three-digit subtraction with borrows, which small arms on the first model could not learn under any supervision we administered — and the 2s with its softmax twin. The twelve days that followed asked a narrower question: how do you teach a byte model a behaviour through a detachable arm, what does it cost, and what happens when arms meet? Every arm below was born inert on the frozen trunk, so switching it off returns the model to the bit; the cost gauge is bits per byte on held-out web text (fineweb-edu), lower being better, the bare 2s at 1.1097.
Instruments first. A deliberately broken training closure failed its gate, the exam grew from 30 to 150 items, and a same-seed rerun repeated to the bit, so the noise bar was the spread between seeds. Arms trained on a teacher's distributions chained five rules on nearly every item and never stopped; reading the next byte after a teacher-forced fifth sentence explained why. The bare model stops with probability .715; arms trained on the whole reply including the turn-end bytes stop at .950 and .999; reply-only targets had eroded the stop, to .439–.593 with one-hot targets and .022 and .061 with soft ones. Looping is eroded stop supervision. On the 2s, two one-hot rows on the turn-end pair appended to the target restored termination from 0 and 0 to .960 and .867 at a chain cost of −.034 and −.054, inside the .04 seed bar.
The arm ladder on the 2s. From 09-07, 27 arms of 8.6 million parameters each trained in two days without a skipped cell. The frame import — an arm at block 9 learning two text encoders' geometry — replicated with a larger gain than on the first model (+.268 and +.262). A frame arm stacked over a chain arm erased the chain (.853 to .007); the dense router we call the mixer blended the same arms and halved the skill (.853 to .453); arms trained in each other's presence chained best (.860, .827, .900), still looped, and fused into units (.013–.047 with the partner masked). The wall held: .000 on every arm-side cell, both seeds, both held-out draws. Then the record caught itself: every composition row repeated a law settled in July on another model — always-on stacks destroy, a mixer only contains experts already selective, the cure is to train selectivity in — and the experiment that law prescribed had not run. Part of the week had measured what was already known.
Quiet arms. Each arm then trained with a quiet term, a penalty on any change to the model's output on text outside its job. The chain arm kept its skill while its web-text cost fell from +.067 to +.0026 bits per byte, and two quiet arms simply stacked composed: chain .860 with termination .973 where the same recipes without the term had read 0 and 0; the second seed .827 and .927. A candidate on two seeds, with one sharpening: one non-quiet member stacked on the pair took the chain to 0, so every member must be quiet. Over such arms the mixer was a cost (.693 and .740); soft provisioning — a bounded per-anchor gain read from a trace of recent use — recovered most of the deficit (0.813 and 0.787) without closing it, a candidate corrected once: the gain adds amplitude rather than moving it. The quiet stack became the composition default.
Reasoning without a scaffold. Could an arm answer with the bare word, no chain spoken? A recurrent cell bolted on learned a position trick and lost to its own non-recurrent control (.633 against .867 with the rules in order; .167 and .207 shuffled), and was stopped after one seed with a rule kept since: build inside the running system and check the record first. One data change — every training draw rendering the five rules in a fresh order, so the trick could not fit — let the certified arm, unmodified, compute the chain internally: .527 on never-trained words, .553 with the rules shuffled, 1.000 on the training vocabulary; the second seed .400 and .493. The misses were a favourite word winning the first byte (one word in 45 of 71 errors), and the three worst words were the three English byte statistics refuse to close. That built the atlas: every byte trigram a cell in a fixed 256×256×256 space, weighted from tokenizer vocabularies rather than a corpus; it predicted the arm's per-word failure at a rank correlation of −.976 where the corpus statistic managed −.62. A pool minted from it, 24 easy words and 24 traps, erased the trap ladder on both seeds (−.976 to +.048 and −.214), lifted closure to .840 and .793, and exposed the next failure: with the answer's prefix partner in the prompt, commitment read .26, against .80 without it.
Three dials, one at a time. Compute held at 4,000 steps, each run changed one knob:
| pool | what changed | closure on the never-trained words, seed 0 / seed 1 | web-text cost, bits per byte |
|---|---|---|---|
| 12 training words, rules in a fresh order each draw | the shortcut blocked | .527 / .400 | +.00493 / +.00604 |
| 48 atlas-minted words (24 easy, 24 trap) | the trap band trained | .840 / .793 | +.0172 / +.0148 |
| the same 48 words, quiet weight 2 | the quiet dose | .867 / .847 | +.00535 / +.00647 |
| 64 words, 8 prefix pairs in the rows | the pair dose | .960 / .907 | +.00679 / +.00747 |
| 96 words, 12 pairs | breadth at the same compute | .833 / .793 | +.01132 / +.01209 |
Doubling the quiet weight met the .012 bar on both seeds with closure above the weight-1 rows, so weight 2 became the constant. Pairs in the training rows moved commitment with the dose (.26–.35 with none, .435 and .412 at four pairs, .447 and .506 at eight), and sixteen more minted words drove closure to records of .960 and .907: breadth looked binding. A forecast registered before the next run beat the base rate on both seeds (per-word error .258 and .245 against .264 and .253) but failed its own bar, three of six calls, and its misses shared one cause: at fixed compute the 96-word pool bent back to .833 and .793 and twelve pairs saturated at .435 and .447. Transfer is bought by exposures per word, and the third model's stage pools are sized by it.
The core-side door. Training every parameter of the 237-million-parameter trunk on subtraction, at the one dose tried, left held-out subtraction at .027 and .010 under plain labels and .000 and .000 under the teacher channel, while every cell destroyed the prior (+0.74 to +1.34 bits per byte) and learned only the procedure — the termination a quiet arm buys at +.002. The trunk trains the prior; capability lives in the arms. Beatrix 3 would be a steering-and-arms craft.
Growing arms over arms. The same night the hierarchy's growth clause was measured: a settled quiet arm loaded trainable under a fresh quiet arm, both under one optimizer, stayed detachable — .833 and .840 with the partner masked against solos of .867 and .853, where the same pattern without the quiet term had read .013 — with one clause missed by a hair (selectivity 2.443 against a 2.5 line on one seed). What the ladder settled for the design: detachable arms grown per curriculum stage, every arm quiet at weight 2, the atlas as the authority over the data, and the minted lexicon as the difficulty dial. The lexicon dose curve and the memory-demand instrument did not run in the window.
2. The design and the pre-flight
(2026-09-07 to 09-20.) Beatrix 3 was designed from the record, and written twice. The first draft, filed on the evening of 09-15, was set aside that evening with four faults, the costliest its attention layout: thirteen softmax blocks and three splat blocks, chosen for speed. The 2s had clipped no gradient in 824 logged steps while its softmax twin clipped 813 of them and finished at 2.8846 bits per byte against 1.1097, so the trunk is splat memory in every block. The second articulation, filed late that night, lists fourteen specifications with the record under each.
| # | the specification, in plain words | status at launch |
|---|---|---|
| 1 | every vetted mechanism becomes code of the official training routine; the weights ship coupled to it | in force |
| 2 | the trunk is the 2s form: splat memory in every block | validated |
| 3 | arms attach one per curriculum stage and train beside the trunk, each collapse-tested | designed; co-training settled, two seeds |
| 4 | every arm quiet, so a stack needs no mixer | validated, two seeds; in the library |
| 5 | a soft gain lets an address atom carry a little more where it earns it; never a cull or a selector | arm form validated; hub form built, inactive |
| 6 | diversity encouraged in pretraining, condensed after it | designed; not in this run |
| 7 | tool arms that ship with how the tool is used | designed; not in this run |
| 8 | guards certified before the run; health gauges at every boundary | certified |
| 9 | the atlas is the tokenizer and the data authority | built; validated in parts |
| 10 | data breadth binds at this scale, not arm size | adopted; falsifier owed |
| 11 | recurrent arms only where the data route saturates, against a matched control | types built; not in this run |
| 12 | memory accumulators: the decaying average is already live core structure | live; head form untested |
| 13 | deliberation trained after pretraining as a gated, quiet arm | direction only |
| 14 | the core gets more: four times the data, weak bytes fused, the curriculum and anneals kept | data adopted; fusion shelved |
The arms form a hierarchy (template and termination, then capability, recurrent deduction, frame), each attached when its stage opens and trained beside the trunk, every one quiet so the stack needs no mixer. The atlas, a deterministic space of 256-cubed byte trigrams, stands in for a tokenizer as the data authority: it audits the mix, dials difficulty and builds lexicons. The fourteenth, added on 09-16, set the size: the 2s had been cut off with its loss still falling at about 5.6 tokens per parameter, so the core gets four times the data, 64.4 billion bytes against 16.1.
The price and the guards. On 09-19, on the local RTX 4090 at the mission's 262,144-token step, the 2s shape trained at 32.0 thousand tokens per second and its softmax twin at 1.005 times that: here the splat memory is free. On the Blackwell-class rental card, 64.4 billion bytes priced at 9.7, 12.9, 13.7 and 16.4 days for 16, 20, 24 and 28 blocks, against a 09-07 premise of no multi-week run. The faster path was dead on arrival: a compiled model ran 1.45 times faster with 488 of 491 gradients non-finite, and a bisect traced the fault to the output head compiled alone (486 of 491). The record already held the same fault on the Blackwell card: no lever.
The guards were certified by induced faults on a 40-million-parameter twin of twelve blocks: two healthy seeds for 4,000 steps, four faults forked from one at step 2,500 (two learning-rate doses, a sharpened dispatch, a rank lesion), and a pass only if a guard fired within 600 steps of its fault and never on a healthy seed. One of three passed: the rank-collapse guard, 100 steps after the lesion and after the harsher dose. The dispatch-entropy guard fired on both healthy seeds at the last block, which sharpens by itself; the norm-surge guard never fired, its 1,000-step median integrating a 200-step transient away. An offline screen of 147 settings found 57 that catch both doses and stay silent, the best proved on a third healthy seed. The run launched with one guard that halts and two in watch mode.
The afternoon's screens. The coverage audit ran the atlas over samples of the 26 sources: half of the planned 64.40 billion bytes fall in 254 trigram cells of 6,954 live, and the nine synthetic generators, 24.7% of the bytes, add no cell. The head-birth screen trained four fresh 24-block crafts for 1,200 identical steps, one seed: the head born empty buried itself again, liveness 0.11 of chance; the head solved in closed form at birth carried +3.287 bits per byte; the head transplanted from the finished 2s read +0.038. Revival at birth became the mechanism, the mission its second seed.
Fusion, shelved. The size specification also asked for weak bytes, those nearly determined by the three before them, fused into stronger units at the input. Before training, the weak bytes are simply the bytes inside words; the atlas adds the predictable word, 2-3% of word starts in prose and 10-23% in the generators. The screen ran on a rented machine with two RTX 5090 cards: four crafts of the 2s shape on the same half-billion bytes, one seed each. Every fused plane finished +0.117, +0.124 and +0.124 bits per byte behind the plain byte trunk at 1.48 to 1.49 times the speed, where 1.7 had been predicted. The first reading called this a refutation; the reading that stands is not ready: the atlas-weighted unit the design calls for had never been built, and a stand-in tested along single axes at small scale cannot refute a design. The three rules cost the same 0.12 with units of nearly twice the length, so the failure sits in the fixed blocks that summarize and decode a unit. Fusion was shelved as a routing idea for a later arm wave; the run went unfused, forgoing about 1.48 times the speed.
Depth. A pre-registered ladder of fresh 16- to 28-block crafts was to set the depth, and its 24-block rung, rank rising to block 15 and falling over the last eight, held it at 24 by the rule. We ruled otherwise that night: depth would answer questions about structure and routing at a depth a 20-block model cannot reach; the loss might even be worse, and it was worth the cost. The depth became 32 blocks, 376 million parameters against 283 million at 24, about 11.7 days on the two cards against 8.8. The ladder, finished two days later, single seeds, supported it: matched at 1,145 steps the held-out loss read 1.561, 1.521, 1.511, 1.492 and 1.476 bits per byte for 16, 20, 24, 28 and 32 blocks, about 0.004 per block past 20, and the rank peak sat at 0.61 to 0.75 of the depth at every depth, which retired the plateau rule with its premise.
Three settings. The anneal multiplier, screened on the 2s's own final phase at full rate and at 0.3 and 0.1 of it, one seed each, stayed at 1.0 by its pre-registered call: the lower rates ended lower on web text (1.08217 and 1.08322 against 1.10305) and kept the text-frame stop rate (.860 and .953 against .647) but did not reduce echo, the call's measure. The arm geometry inside the supply law, 16 slots and 16 atoms at address dimension 8, was certified on the frozen 2s on two seeds: solo rule chaining .840 and .773, off-job writes +.0065 and +.0074 against a bar of .012. A local screen then set the dial on the atlas's 216-word lexicon for the rules stage: the quiet bar met on both seeds (+.0101 and +.0086), half the rows as pairs moving the trained-pair gap from −.157 and −.165 to −.020 and +.134, a quiet weight of 5 buying nothing over 2. Set on 09-22 (pair rate .50, quiet weight 2 throughout, a .24 share of the stage), it was installed at the first stage boundary.
The launch. Beatrix 3 launched at 04:47 on 09-20 on the two RTX 5090 cards: 32 blocks, 1,024 wide, context 4,096, about 376 million parameters, 262,144 tokens per step, 21.5 GB per card. The data plane: a 0.30-billion-byte warmup, 20.90 billion of fineweb-edu, 35.2 billion over the nine stages and two anneals of 4.00 billion each, 64.40 billion bytes in about 245,670 planned steps (the run ended at 245,674), estimated at eleven to twelve days. Eight stage arms were planned at the certified geometry, 13.69 million parameters across the 32 blocks, one per stage from s1, attached as each stage opened; the next sections tell what became of it.
3. The run
(2026-09-20 → 10-05.) The run launched at 04:47 on 2026-09-20 on two RTX 5090 cards in parallel, 262,144 bytes a step, and completed at 11:02 on 2026-10-05 after 245,674 steps, 64.402 billion bytes and 295.0 hours of training. The diet was four times the 2s's: a 0.3-billion-byte warm-up on wikitext, 20.9 billion bytes of fineweb-edu, nine curriculum stages of synthetic text (2.8 / 2.8 / 4.0 / 4.8 / 4.8 / 4.0 / 3.2 / 5.6 / 3.2 billion bytes) that each teach one skill in a fixed wording, and two anneals of 4.0 billion bytes of plain web text, without and then with the chat frame. The anneals change the diet, not the learning rate, which stayed flat to the last step. Held-out numbers are the boundary reports' reads of the fineweb-edu holdout in bits per byte (lower is better); the curve between boundaries runs a few thousandths off them on a lighter gauge.
The web phase. The warm-up closed at step 1,145 at 2.3875, the splat memory already carrying the attention (+1.295 when switched off), the head alive at +2.584, the bank inert at +.034. The run's one guard flag fired at step 2,800 on the last block, a dispatch-entropy read of .55 against a reference of .89 — the byte funnel, whose rank sits at 2 by design, a known false positive. Pretraining closed on 2026-09-24 at step 80,873 and 21.20 billion bytes at .9838, and the bank had elected its function on web text alone: off, it now cost +3.200 (the head +2.304, the splat memory +2.567).
The curriculum and the anneals, by boundary.
| phase closing | step | held-out bpb | what the phase did to the web holdout |
|---|---|---|---|
| warm-up, wikitext | 1,145 | 2.3875 | the start of the curve |
| pretraining, fineweb-edu | 80,873 | .9838 | the lowest boundary read of the run |
| s0 counting | 91,555 | 1.1178 | +.134, the largest single cost |
| s1 false belief | 102,237 | 1.1273 | +.0095; the halt |
| s2 concepts | 117,496 | 1.1000 | -.0273, the first repayment |
| s3 rule chains | 135,807 | 1.1084 | +.0084; a plateau at 1.10-1.11 |
| s4 arithmetic | 154,118 | 1.1273 | +.0189; the arms came off at 148,000 |
| s5 causal chains | 169,377 | 1.1236 | -.0037; the first stage run entirely on the trunk alone |
| s6 try-and-fail | 181,585 | 1.1498 | +.0262, the curriculum's high (the 2s paid +.0179) |
| s7 mixed | 202,948 | 1.0482 | -.1016 in ten straight falls (the 2s repaid -.0800) |
| s8 register | 215,156 | 1.0780 | +.0298 (the 2s paid +.0028) |
| anneal, no chat frame | 230,415 | 0.9951 | -.0829; under one bit per byte again |
| anneal, chat mix | 245,674 | 0.9861 | -.0090; the whole anneal -.0919, the end .0023 above the pretraining close |
The organs. The splat memory's share grew at almost every boundary, from +2.567 at the pretraining close to +6.420 at the curriculum's end — its largest step, .97, at the mixed stage's close — fell for the first time to +5.859 under plain web text, and read +6.413 at the end. The head's share eased from +2.304 to +1.900. The bank's share sat inside +2.89 to +3.40 for eleven boundaries; then the end boundary's evaluation read +6.129. That number is not settled: the per-checkpoint gauge of the same toggle read +3.40 to +4.14 through the mixed anneal (+4.07 at step 244,000) and had agreed with the boundary gauge at 230,415 (+3.397 beside +3.61); the end read is an outlier awaiting a repeat on the final weights, not a finding.
The arms in training. From stage 1 the design attached an arm at every stage close — a 13.69-million-parameter adapter after each of the 32 blocks, gate born closed, head born at zero — trained beside the live trunk under the quiet term at weight 2 on sixteen chunks of held-out web text beside the sixteen task chunks of every step, so that each arm stayed silent off its stage and came off exactly. The first attached at step 91,555 on 2026-09-24, and the step went from 4.0 seconds to 9.5-10: one extra forward per live arm on every quiet chunk. A first estimate put the end of the run as designed near October 21 rather than October 1; we priced four options and changed nothing: the certified form would run through Friday, then an engineering window would open for speedups that leave the objective exact.
The halt. On 2026-09-25 we halted the run cleanly at the stage-1 close — step 102,237, 131.08 hours of wall, no steps lost — by a flag the trainer re-reads at every boundary, which gave the window a clean start and the first read of an arm on a live trunk: masked 1.12720 against live 1.12729 on the web holdout, +.00009, silent off its stage as designed. On its own stage it had never been measured, and the window measured it first: stage rows armed .42035 against masked .42031 (+.00004), the gate closed from -3.06 to -3.53 in logit (a sigmoid of .045 falling to .029), the head's weight norm 2.61 against 144 for an arm of the same shape trained alone on a frozen trunk. The arm was inert: a trunk training on the same rows takes the stage before an arm can, and the quiet term, with nothing to protect, closes the gate. No certified result had covered an arm born under a live trunk; the ladder's arms had all been born on frozen ones.
The engineering window (2026-09-26). Profiling the step at zero to eight arms found a memory wall first: the task chunk held 18.8 gigabytes with no arm, each attached arm added 3.06 gigabytes of stored intermediates for its backward pass, and the card ran out of memory at the fourth arm on its 31.4 gigabytes. Run honestly, at half the micro-batch, the design would have ended around November 30, not October 21: the first estimate had omitted the kernel-launch cost and the memory wall. The arms' time was launches, not arithmetic — one arm added 8.3 milliseconds and 1,280 kernel launches to every forward. Batching the Muon optimizer's Newton-Schulz iterations bought nothing (0.9x). Recomputing each block's arm chain in the backward pass was the memory fix, about 0.4 gigabytes per arm instead of 3.06. Compiling the chain of arms as one function was much faster and gave a non-finite input gradient in 60 of 60 batches; a bisection placed the fault in the aleph address alone, where the compiler fused a normalize-matmul-amax-exp-divide chain with full-precision intermediates, and making it emulate eager's rounding at every cast cleared it at 2.4-2.6x per block call. We stopped the run at the 106,000 checkpoint (18 steps lost) to install the libraries and census the gradients on the run's own cards — 0 of 1,643 past 1e-2 at three arms, clean at eight, 1.84x / 1.69x over eager on the task and quiet chunks at eight arms; the full censuses are in the technical companion.
The quartering. The finish dates tell the window's yield: November 30 as the run was, October 28 with recompute and the compiled chain, October 14 with the one estimator change we then made — four quiet chunks a step beside sixteen task chunks, weighted to keep the quiet term's per-step weight, at the price of a quarter of the off-domain rows (8 instead of 32) and twice the noise on that gradient. The two-arm step fell from 11.35 seconds to 6.9, and a monitor read the quartered run against the old form at every 2,000-step checkpoint with halt conditions taken from the softmax twin's collapse; none was ever met.
Never eager. Then the compiled chain was lost without anyone noticing. At the stage-2 close on 2026-09-27 (step 117,496) the boundary evaluation fed the armed model probe prompts of many lengths; the static compile hit the compiler's default budget of eight recompilations, after which torch 2.8 runs the function uncompiled — eager — for the rest of the process. The third arm attached onto an eager chain and stage 3 ran at 10.4 seconds a step against the 8.5 expected; left alone, the fault would have cost about nine days (October 24 against October 14). The fix was built and tested the same evening in amoe-lora 0.2.11 (a budget of 64, only the two training shapes compiled, every other shape exact and eager), and the next morning the run stopped at the 122,000 checkpoint (8 steps lost) and relaunched compiled at 7.29 seconds a step, about half a day having run eager. The rule that followed is standing — the arm chain never runs uncompiled again — with a guard that reads the training log and relaunches compiled at the next checkpoint if the warning reappears. The stage-3 close on 2026-09-29 passed with zero warnings, and four arms ran at 8.3-8.6 seconds a step.
What the arms were doing. On 2026-09-30 we asked whether the training was having any effect and measured it in forty minutes on thirteen checkpoints of this model and the 2s. Every stage taught its own generator's wording to about 1.00; in new wording the same skills barely moved (the nine-suite exam .370 to .426); each stage's gain faded once its rows left the mix (false belief 1.00 to .47); the arms carried nothing even with every gate forced fully open (-0.0020 bits per byte on their own rows at most), while a fresh arm of the same shape, trained 400 steps on the frozen stage-0 weights, took 64% of stage 1's whole gain (.854 to .337); and the arms were costing about half of every step, 4.0 against 8.4-8.6 seconds. The decision was to detach the arms at the 148,000 checkpoint and finish the trunk alone — about October 5 instead of October 14 — refining the arms later on the finished weights (section 5). The change was an empty arm list and a relaunch, rehearsed first on a CPU to prove that the trunk, both its optimizers and its data stream continued exactly; it could not disturb the trunk, since the quiet term had never updated it and all-arms-off had read as the bare trunk at every check. At 17:34 on 2026-09-30 the arms came off, the step fell to 4.17 seconds, and the 97,671 remaining steps became 4.79 days.
The trunk alone. Stage 4 closed up +.0189, the first boundary without arms looking like every armed one on every gauge; stage 5 closed down -.0037; the try-and-fail stage closed up +.0262 after the largest stage-opening jump of the run, +.0239; the mixed stage repaid -.1016 in ten straight falls, the whole try-and-fail cost gone within 109 million bytes; the register stage climbed +.0298, ten times what the 2s had paid for it. The first anneal, 15,259 steps of web text without the chat frame, fell -.0829 in eight straight falls to 0.9951, and the splat memory's share fell for the first time at any boundary, +6.420 to +5.859: web text leans less on the whole-past memory than the curriculum rows did. The second anneal, the same length with chat rows back in, fell -.0090 more to 0.9861, where the 2s had given ground (+.0056).
The end. The last step ran at 4.00 seconds, the final checkpoint uploaded in 52 seconds, and the summary line printed: 245,674 steps, 64.402 billion bytes, 295.0 hours. No loss spike fired in the whole run, the gradient clip at 1.0 was never reached, the one guard event was the false positive of step 2,800, and no halt condition was ever met. Hugging Face holds 467 files and 171.59 gigabytes under the model at https://huggingface.co/AbstractPhil/alephllm-mini-beatrix-training, each final file verified byte for byte against the machine. What the splat memory remembers across these checkpoints is section 4's; the arms' refit on the frozen trunk began the same afternoon and is section 5's.
4. What the memory remembers
(2026-09-27 to 10-05.) A good next-byte predictor can still lose a detail from a page ago, so in stage 2 we began reading the trunk's memory with a small benchmark, on new and stored checkpoints. It plants a five-digit code in a document and asks for it 0, about 1,200, 2,200 or 3,500 bytes later, scoring the digits right, and asks for a copy of a random 24-letter string with no cue, scoring the share of its cost removed. Each cell holds 64 trials; the far cell is the edge of the 4,096-byte window. The number we trust most is the plant's saving: the code's cost in bits without the plant minus with it, on the same trials, so the digit prior cancels.
The mid-run read. At step 114,000 the trunk had .753 of the digits right at once, .459 at ~1,200 bytes, .131 at ~2,200 and .031 at ~3,500; the copy removed .775 / .570 / .333 / .157 of the cost and never reproduced the string exactly past distance zero. The softmax twin of the 2s at step 20,000 read .778 / .694 / .603 / .588 and copied at .92 to .96 at every distance; the first Beatrix, thirteen softmax blocks and three hubs, held .96 to .99 out to ~1,700 bytes in its 2,048-byte window. We marked it as a benchmark and continued unchanged, with a softmax hybrid for the next prototype; the benchmark then ran at every stage close.
The trajectory. Pretraining left the copy at .778 at ~1,200 bytes and the plant's saving there at 17.33 bits. Stage 0 cut the copy to .506; it then tracked the held-out loss at nearly every close, .432 after stage 6, .595 at the curriculum's end. Keyed lookup recovered less: the saving at ~1,200 bytes fell to 5.40 bits after stage 6 and was 7.24 at the curriculum's close. The no-chat anneal brought the copy back to .818 / .735 / .602 / .312 against pretraining's .863 / .778 / .667 / .401, and the saving to 11.28 bits, roughly a third short of pretraining's. The split narrowed and stood: the copy follows the loss; a keyed fact at a distance does not.
Why. We assign the fade to the splat memory, not the curriculum. It is an additive prefix state: nothing decays or is erased, and it cannot clear even at a document break inside the window. On this reading a code written early is still there under everything written since; the fade was present at step 42,000, mid-pretraining, and in the all-splat 2s before any curriculum.
What it changes. Beatrix 3 is still our best next-byte predictor at every context length, so the splat memory stays where it is strong and the next prototype is a hybrid marked by the model's own rank profile. The weakest-rank blocks take standard attention: blocks 0 to 5, 30 and 31, whose effective ranks read 141 / 102 / 139 / 174 / 215 / 242 / 228 / 2, with candidates at blocks 8 or 9, 14, 20 and 26: eight to twelve softmax blocks of 32. They need a guard from the first step, probably query-key normalization or a decay, because the all-softmax twin came apart under this recipe, clipping from step 17,600. A screen of one softmax block per four to eight, two seeds against all-splat and all-softmax controls, sets the count. The splat memory gets its upgrade from write rules on record, not yet in a full model: overwrite, a whitened write, a direct lookback. Lookback arms on the finished trunk are offered, not run. The bar is this benchmark at the same steps, with loss and rank profile beside it: hold the planted fact past 2,000 bytes at the twin's level, .588 at ~3,500 against .031, keeping the held-out advantage.
What is owed. The benchmark on the finished weights at step 245,674 has not been run, held for a decision; the trajectory ends at the no-chat close. Nothing of the hybrid is built; six decisions are open, among them which middle blocks go softmax, what the last block does, which guard, and whether it is a 3.5 on this shape or the next model.
5. The arms after the trunk
(2026-10-05 to 10-06.) The arms came after the trunk because the trunk, while it trained, took each stage's text for itself: the arms that trained beside it through the first four stages were nearly empty at every close. An arm is bound to the weights it trained on, and a fresh arm on a frozen mid-run snapshot had taken stage 1 text from .8536 to .3367 bits per byte in 400 steps. So the arms were snapped off, the trunk finished alone, and its finished weights, step 245,674 on 10-05, became the frozen bed for every arm below.
What an arm is here. One module after each of the 32 blocks, 13,692,960 parameters: the block's output is projected to 16 slots of 8, read against a 16-atom address of the trunk's signed kind, consumed by a small feed-forward layer of width 256 and added back through a sigmoid gate born nearly closed, its output weights zero. It trains with the trunk frozen, under pure Adam at 1e-3 and no weight decay, on 2 × 4,096 bytes of its stage's text a step, beside a quiet term: twice the divergence between the bare trunk's predictions and the armed model's on text outside its job.
The solos that failed together. Our first form was one arm per stage, alone, 800 steps, quiet on web text only; each solo took its stage (stage 1 text .5538 → .0809, a screen number). Switched on together, the four solos of stages 1 to 4 read worse than no arm on three stages: stage 2 text .7193 → 1.7640, stage 3 .9177 → 2.0510, stage 4 1.2245 → 1.9991, web +.030 against a limit of +.012. Taught quiet on web text alone, each arm wrote freely on the others' stages, and the solo form was withdrawn that afternoon: the record's recipe trains the arms together, each quiet on its partners' text too.
Five lines and the first group. Before the first group run we fixed five pass lines on two seeds: the group keeps 0.8 of each solo's gain, each member carries 0.8 of its own stage, web text and partner stages move under +.012, and no earlier arm loses a fifth of its gain under a later one. The first group attached each stage's arm fresh over those already there, every member trainable, 800 steps a stage. It met four lines, but the members' shares of their own stages were 0.94, 0.21, 0.02 and 0.00: every stage's rows reach every attached arm, so the first arm, already open, learned each new stage before the newborn arm, gate at .05 and output at zero, could open for it, and two of the four arms were dead weight.
Frozen, then the roster rule. The record already prescribed the fix, in two parts we had not built: lower arms stay frozen while upper ones train, and every member is quiet on the other members' text from its first step. Freezing alone gave each member a share (0.98, 0.68, 0.73, 0.60), but the frozen first arm still wrote on later text, leaving a later arm only what was left. The roster rule finished it: the first arm's write on stages 2 to 4 fell from −.1918, −.1394, −.2521 to −.0072, −.0020, −.0076, each member held 0.94 to 0.99 of its stage and read alone within .009 of its solo, and all five lines were met on both seeds.
The four that shipped. The released four-arm group is a hierarchy: arms 1 and 2 of the first group kept frozen, 3,200 and 2,400 steps behind them, arms 3 and 4 retrained over them under the roster rule for 800 steps. All on, it is the best group on every stage on two seeds, .0482, .1123, .2236, .3384 at web +.0017, but its shares are 0.94, 0.21, 0.13, 0.16: mount it whole. Under routed dispatch, a learned key dividing one amplitude among the members, both groups read worse than the plain stack on every stage (+.025 to +.24 and +.12 to +.19): the dispatch shares one fixed budget and this text wants several arms in full.
The length. A 3,200-step solo showed that 800 steps was a screen: stage 3 text read .3724 at 800 and .1842 at 3,200, stage 4 .3889 and .2737, the fall per 400 steps halving after 1,600. The record's arm length is 4,000, where the curve flattens; every later arm ran there.
The eight. The arm program is the curriculum's continuation, so the four stages without arms got theirs: the published four held frozen and on, stages 5 to 8 attached one at a time for 4,000 steps each under the roster rule. This is the package's default stack.
| arm | arms off → the eight on (seed A / seed B) | the arm alone | a solo arm at 4,000 steps (two seeds) |
|---|---|---|---|
| s1 perspective | .5538 → .0492 / .0484 | .0525 | .0469 / .0467 |
| s2 concept | .7212 → .1105 / .1212 | .3503 | .0754 / .0753 |
| s3 rules | .9296 → .2237 / .2222 | .8236 | .1804 / .1795 |
| s4 arithmetic | 1.2339 → .3279 / .3266 | 1.1267 | .2671 / .2674 |
| s5 causal | .6604 → .0431 / .0428 | .0911 | .0428 / .0433 |
| s6 try-fail | .6728 → .0548 / .0543 | .1638 | .0568 / .0570 |
| s7 mixed | .7830 → .7200 / .7198 | .7242 | .7153 / .7171 |
| s8 register | .4907 → .0212 / .0208 | .0678 | .0208 / .0209 |
| s9 caption (experimental, over the eight) | 1.2648 → .7829 / .7795 | .7851 | — |
Bits per byte on held-out rows of each stage; the eight's web cost +.0019 / +.0010, probe mean .411 → .422 / .430.
The new arms take their stages, leave the four within .009 and write at most +.0012 on a partner's text; the mixed stage is the one no arm moves. Four lines met on both seeds; the share line missed, three of the new arms at 0.67 to 0.76. A second eight, from scratch at 4,000 steps an arm, ran under a rule fixed before its numbers existed: it replaces the continuation only if lower on at least five of eight stages on both seeds. It was lower on five of eight on one seed and three on the other, so the continuation stayed, and the difference is training length, not form: lower by .034 to .061 exactly where the published arms kept 2,400, 800 and 800 steps, equal within .001 where both had 4,000. Solo arms at 4,000 steps, two seeds each, read within .003 of the from-scratch arms on six stages and better on two: at this length the capability is the arm's own, and a group buys quiet coexistence, not a lower number.
The caption arm. The ninth member is for the diffusion line: an arm on captions, weighted toward photographs, so that her states can drive image-editing sliders. Its pack has five public caption sources: 121,153 fashion and 42,080 synthetic-character photographs, 1,242,176 CC12M images with three captions each, 603,026 danbooru rows and 19,227 COCO captions, 2.87 GB in 19 files. The house age protocol was applied as written, its explicit-minor tier rejecting rows from every source. The atlas read the lexicon as easy, its dominant-continuation share .36 to .37 against an easy band at .30 and a trap band at .75; the sources differ in breadth, not difficulty. Fitted over the frozen eight for 4,000 steps, it takes held-out caption rows from 1.2648 to .7829 (.7795 on the second seed) and carries the whole gain itself, reading .7851 alone. Its gate opens to .45 against the stage arms' .11 to .28, and it was still falling at 4,000 steps (.708 → .693 over the last 800) where the stage arms had flattened: a longer arm is owed. It ships as an experimental third group, not the default stack; what it does to pictures is not measured.
The releases and their gates, the chat Space, the 2.5s package, the byte scope and the library's arm mount are the next section's.
6. What ships
(2026-09-16 to 10-06.) The model lives at https://huggingface.co/AbstractPhil/mini-beatrix-3, one repository shipping the trunk bare and with its arms, on the 2.5s pattern. Five revisions in just over a day built it, and none was pushed until its exact bytes had passed the proofs: outputs matching the training library's to a logit difference of 0.0 on every row; from the first arm on, every detach restoring the bare core bit for bit; a load through plain transformers, library uninstalled, offline; and a tokenless load from the published repository repeating it.
| revision | pushed (local) | what it added | rows matched at 0.0 | tokenless load |
|---|---|---|---|---|
| 9494036 | 10-05, midday | the bare trunk: 376,017,249 parameters; bf16 weights (754,561,258 bytes) with the fp32 masters as a variant (1,509,020,412 bytes); 14 files | both weight files | 91 s |
| acf94d3 | 10-05, 17:19 | the four-arm group of stages 1 to 4 and four 800-step solo arms; the arm index; mount, detach and arms() in the wrapper |
9 | 258 s |
| efe3aa0 | 10-05, late evening | the eight-arm group as the default stack and eight 800-step solos; the four unchanged | 22 | 785 s |
| 8541601 | 10-06, morning | the solos replaced by their 4,000-step versions, both seeds' readings on the card; core and groups unchanged in bytes | 22 | 732 s |
| 0ae0e92 | 10-06, afternoon | the experimental nine-arm caption group; the default stays the eight; built on a desktop processor with no GPU | 32 | 1,981 s |
The head holds 46 files: two weight files, the wrapper, its runtime, nine vendored model files, the card, the index and 29 arm files of about 54.8 MB each, every one a 13.7M-parameter arm fitted on this exact frozen trunk. The index records each arm's trunk step, training length and measured numbers on both seeds. from_pretrained with trust_remote_code gives the bare model; mount_arm("stages-1-8") switches the default stack on, "stages-1-8-caption" the nine, and a member or a solo mounts alone by its row name; detach_arm(verify=True) restores the bare core and raises if one bit differs; arms() returns the card's tables as data. Saving refuses while an arm is mounted, solos are for one at a time, and only transformers is needed.
Everything behind it is in the training repository, https://huggingface.co/AbstractPhil/alephllm-mini-beatrix-training, which holds the run's checkpoints and boundary reports (section 3); the four arms trained beside the trunk and detached at step 148,000; the arm refit folder, with every form's weights, results and logs, shipped or not, both seeds included (its README lags the package by two releases; the card and index are current); and the caption pack, 19 files and 2.87 GB, with a file pinning every source revision and shard checksum.
Three older surfaces frame it. The chat Space, https://huggingface.co/spaces/AbstractPhil/alephllm-chat, opens on the finished 3 since 10-05: step 245,674, chatting bare at 9 bytes per second (the 2s runs at 12), beside 135 of her checkpoints and their 180 arms. mini-beatrix-2.5s (https://huggingface.co/AbstractPhil/mini-beatrix-2.5s, 09-22) is the 2s trunk byte for byte, nothing retrained, plus thirteen mountable arms from the September ladder, proved equal to the libraries at 0.0 before its push. The byte scope, https://huggingface.co/datasets/AbstractPhil/beatrix-captured-interactive-inferences, is a viewer with its data: two 2s inferences captured byte by byte, bare and with an arm, and the splat memory's exact effective byte-to-byte attention, derived from the read, not probed.
On 10-06 the arm mount went into the library itself (alephllm 0.10.6, https://github.com/AbstractEyes/alephllm): load the frozen trunk at a stored step, attach a group in training order with every anchor checked against that step and every carried member by its content hash, mask members, detach bit-exact; thirteen checks pass on a processor in 9 seconds, no model or training code changed: a second route to the same arms, over the training repository's anchors. 0.10.7 moved the datasets dependency into a train extra so one environment holds this library and the diffusion stack. The technical companion (TECHNICAL_3.md) and the per-topic technical files under docs/technical/ in the training repository will carry the depth.
7. Beatrix 3 as a conditioner
(2026-09-28 to 10-03.) With the trunk at step 124,000 we asked what it could already do as the text side of an image generator with nothing trained, reading its states on the project's caption gauges. The pipeline was certified first: after an audit fixed three invalidating bugs in the first write-up, it reproduced all 30 stored cells of the 2s exactly.
What the middle blocks hold. Against T5-XXL the pooled states of blocks 14 to 28 agree in shape (CKA) at .62 to .65; an untrained copy reads .23 to .25; a second caption draw moves a row by .017 on average, .039 at most; T5 and bert agree at .714. A linear map from those states predicts T5 at .38 to .43, but the untrained copy already reaches .26 to .31, so training's share is .10 to .13: we had written that the middle blocks predict T5 about as well as bert does, and the untrained copy withdrew it. Attribute binding reads .89 against .71 for the 2s; relations ride on mention order alone. The stage arms, re-attached at 148,000 and 194,000, changed nothing on captions to three decimals.
The one large direction. One fixed direction at the last block carries 94 to 99 percent of each byte's energy for every input, spread over many channels (the top holds 1.6 to 3.4 percent), its eight leading channels and their signs fixed from step 42,000; the untrained copy has none. The bank writes it: the always-on path about two thirds on every byte, the address-mixed experts the rest, one adding and one subtracting, so the address sets the amount byte by byte. Since the final normalization divides each byte by its own spread, a huge component along one direction shrinks the content reaching the head and softens the prediction: the size along it is a per-byte temperature. Remove it and captions cost +43.5 bits per byte over 1.772; pin it to a constant, only +.148; a random direction, +.002. The trunk fights it with its last normalization, whose gain on those channels fell from .99 to .15-.24 while the byte norm grew from 2.6 thousand to 1.64 million.
| step | share of each byte's energy along the direction | byte norm at block 31 | the final normalization's gain on its channels | share after the normalization |
|---|---|---|---|---|
| 1,145 | .874 | 2.6e3 | .99 | .874 |
| 42,000 | .987 | 2.8e5 | .79-.81 | .983 |
| 80,000 | .990 | 6.6e5 | .62-.65 | .972 |
| 126,000 | .983 | 8.4e5 | .38-.47 | .905 |
| 148,000 (the arms removed) | .990 | 1.23e6 | .28-.38 | .865 |
| 194,000 | .991 | 1.64e6 | .15-.24 | .669 |
| the untrained copy | .641 | 183 | 1.00 | .641 |
Two thirds of the output end's apparent fall in caption agreement from 148,000 to 194,000 (.603 to .512) was the direction growing; projected out, the fall is .622 to .591, and standardizing each channel leaves it in place (.505 against .512). So for any conditioner: read the middle blocks, condition on the byte sequence rather than a pooled vector, and project the direction out of anything read from the last block.
The fractal theory. The direction is fixed, but the cosine pattern between checkpoints and inputs read to us like the signature of the project's own fractal systems, so we wrote a test with every rule set before it ran: does the per-byte size along the direction fluctuate with the same statistical shape at every scale, with long memory (an exponent of .5 means none)? On structureless input (random bytes drawn from letter frequencies) it read .664, .619 and .651 at 126,000, 148,000 and 194,000, one power law from 16 to 1,024 bytes; time-shuffled series read .48 to .51, natural text .53 to .54, the untrained copy only a drift with position. Supported at 194,000 on every criterion, partly at the other two: self-similarity the model generates itself.
The measurements after it found the always-on path and the first expert each writing about half of its slow part (.53 and .54); it begins in the splat memory's whole-past averaging, builds from block 8 and peaks at block 24. We first wrote that it starts at block 2; that was a window-position shift learned on long web documents (12 to 26 percent of the variance in that phase) and dropped once the curriculum began; without it the exponent holds at .60 to .68 from step 1,145. A reading that the mid-stack holds still while the exit moves was withdrawn.
Rank is a floor, not a dial. Cutting Beatrix 3's states to 256 of 1,024 directions costs +.38 to +.89 bits per byte. Raising rank in training with everything else fixed (a 40M-parameter bed, 4,000 steps) bought no measurable change at a 6.6-fold rise (three of four reads inside the noise bar) and cost +.13 to +.17 at a 10-fold rise, the added objective competing with prediction. A stream's rank at a block boundary follows the block that reads it (a splat-memory block feeding softmax 28.7, the reverse 118.8); capacity followed how many blocks were softmax.
Three requirements for version 4, none a design. The always-on path's amplitude must stay on a budget so no single direction owns the final layer: its share at the last block bounded, the temperature kept, the final normalization freed, gauged at every checkpoint. Whatever bounds it must keep the fluctuation whole, since every band of it is calibrated temperature. The candidate is an amplitude register, the direction's amplitude declared as a number the head reads and its mean held by a slow gain across optimizer steps, with a falsifier (more than .001 bits per byte of cost means normalizing, not bounding), beside an exponent gauge at every boundary; a content-only abstention term was withdrawn, and pinning or normalizing the write, or a smoother, were ruled out by measurement. The third is rank as a floor, offered as a rule. Each waits on a decision; the register's first experiment is written and unbuilt.
The outside yardsticks. A controlled comparison exists: 27 image models from 12 text encoders on one recipe, CLIP-H .622 against T5-XXL .741. TED-6K is the one text-only stand-in calibrated against image results (r .99 over 8 encoder configurations), but it trains a pooling head and is validated only on subword encoders of 0.5B and up, so a Beatrix row would be a placement, not a prediction. Nothing published scores a foreign encoder with nothing trained; no byte-level causal model has conditioned a diffusion model in print. The proposal was not approved, so no generator saw Beatrix 3 in this window.
Three side products. A guidance catalogue of 199 methods and 33 evaluation protocols, each checked against its source. It ships a pure-PyTorch library of 97 combiners. The public ComfyUI pack at https://github.com/AbstractEyes/comfy-cfg-megapack has 50 nodes named for their papers, with 70 comparisons at 1024 pixels on SDXL and Anima. Image quality itself is unmeasured. On 09-25, during a halt in the run, a rank-64 liminal-spaces LoRA on Anima trained from 1,873 captioned images for 5,536 steps in 3 h 51 min. It is published with its dataset at https://huggingface.co/AbstractPhil/mega-liminal-lora and https://huggingface.co/datasets/AbstractPhil/mega-liminal.
8. Sliders on Sana and Anima
(2026-10-03 to 10-05.) This line asks whether Beatrix can steer a picture without a full diffusion training run: whether what she reads in a phrase can be carried into an image model and change what it draws, her weights fixed and everything trained kept small. The instrument is the slider, a learned direction added to the states a text encoder hands the image model, scaled by a dial. The judge is a CLIP ViT-L/14 mood score, upbeat minus downbeat similarity times a hundred; CLIP judges pixels only and never houses her states.
Sana first. Sana is a 600-million-parameter image model at 512 pixels; its text side is a 2-billion-parameter Gemma with token states of mean size 93.4 (the results). The words "cheerful, upbeat" against "gloomy, downbeat" moved the judge +4.949 ± .122 on 255 of 256 paired pictures. The dial — half the mean upbeat-minus-downbeat difference of the same prompts, added to every token of a neutral prompt at strengths -2 to +2 — moved it +0.795 ± .039 per unit on 128 of 128 cells, against +0.112 ± .017 for a random direction of the same size, 14% of the dial; fitted on half the scenes and applied to the rest it holds at +0.788 ± .039, the half-directions at cosine .997. The recipe is activation addition (arXiv 2308.10248) on an image model's conditioning, where such directions were already known (arXiv 2404.01154).
Sana's dial: rows are six scenes at a single seed, columns the dial at -2, -1, 0, +1, +2; the same scene turns from dim and overcast at -2 to sunny and warm at +2, colour moving more than light, the scene kept at .91 to .98 similarity to the unsteered picture.
Sana's words: the same six scenes under the downbeat, neutral and upbeat captions of the first template, then of the second; the words relight the scene rather than tint it.
The Sana mood add-ons. Six LoRAs trained in fifteen minutes, each on 192 of the stock model's own pictures of one mood captioned only "a photo of {scene}", so that only rendering the mood on plain prompts lowers the loss. On eight unseen scenes the upbeat add-on moved the judge +1.050 ± .153 (97% of 32 cells), the control trained on plain pictures +0.304 ± .089 and under the rule, the downbeat add-on -2.065 ± .165 (97%), and a second draw repeated it at +1.159 ± .157 — with the caveat that the weights trained in half precision without full-precision masters, so the numbers hold for the recipe as run.
The move to Anima. The pictures were poor — at this size and resolution the model fails at close-up anatomy (the pipeline matched the stock model to the pixel) — and the goal was restated: test Beatrix's influence by rapid prototyping on an image model whose text side is small enough not to drown an added signal. Anima fits: a 2-billion-parameter illustration model that reads its caption through a 0.6-billion-parameter Qwen3 and a six-block adapter into a 1,024-wide context, where the vectors the image model reads have a mean size of 5.44 against Qwen3's 111.35 (the results; the runner). There the words move the judge +2.590 ± .266 up and -1.385 ± .162 down (92% of 64 cells each); the uniform dial is dead before the adapter, +0.029 ± .018 per unit, and alive after it, +0.314 ± .039 on 86% of cells.
Anima's dial after the adapter: rows eight scenes at a single seed, columns the dial at -2, -1, 0, +1, +2; the picture darkens and dims at -2 and gains sun and flowers at +2. The same push before the adapter leaves the mood where it was.
Words are handles. An attribute screen on stock Anima, six attributes judged by an anime tagger, showed the limit of a uniform push: the tag words move five of six attributes by +7.62 to +14.14 on every cell, while a direction added to every token after the adapter is a few percent of a token's size and works as a tint. Hair colour alone behaved as a dial, +0.69 ± .22 per unit on 94% of cells; a mood direction spread over the adapter's own queries read +0.057 ± .023 per unit. To steer an attribute, a signal must be shaped like a word at a position.
The attribute screen, hair colour: rows eight characters, columns the word "black hair", the dial at -2, the plain prompt, the dial at +2, the word "blonde hair"; the dial tints toward each word without reaching it (the words move the tagger +13.55 and -4.18).
The Anima mood add-ons. Ten LoRAs trained on Anima with full-precision masters, in two draws of pictures, read on 32 held-out cells at their final epoch: five of six cheerful add-ons passed, between +0.999 and +1.397, the sixth read +0.347 ± .242; the gloomy ones, -0.692 and -0.390, fell short; the controls were quiet (-0.034, +0.125). The effect swung .3 to .6 between saved epochs, up to ±0.7, more than each reading's uncertainty of .18 to .28, so a single-checkpoint verdict can flip (the epoch 4 to 10 means put every cheerful add-on above +0.7; the epoch sheet), and every add-on copied its training pictures' look, cream grounds and framed vignettes.
Connectors and sliders. A linear probe on her state after a word's last byte at blocks 16 to 24 tells cheerful words from gloomy ones at .97-.98, against .67-.68 for an untrained copy of her shape. A connector is a small learned map from that reading to one push after Anima's adapter, trained with everything else frozen beside an untrained copy of her, under the rule that both trained mood groups must move the picture before unseen words are read; none of the seven arms of hers passed it. The first pair learned nothing — a learning-rate rule that divides a weight's step by its fan-in (arXiv 2203.03466) is too slow for a signal that lives in one contrast, so after 720 steps the phrase-dependent push was about 6% of a shared one — and whitening her features scattered single phrases. Reducing her reading to one slider value, the phrase's position on the axis from her gloomy to her cheerful training phrases, gave the first descriptive sign: her cheerful side became a dial, +0.579 ± .087 per unit, placing unseen phrases in the order of her own reading (from blissful .09 → +.09 to jubilant and gleeful 1.00 → +.74), while her gloomy side read -0.205 ± .087. We first wrote that down as structural and withdrew it the same morning, when the untrained copy's farthest gloomy phrase moved -0.82: the cause is the tie of both sides to one direction. One direction per side moved her unseen gloomy phrases for the first time, -0.331 ± .136. Across the five slider runs the strong side followed the reader, hers cheerful, the untrained copy's gloomy; all five fail their rule, and what they fixed for the relay is one direction per side, balanced classes, and reading a phrase at the full stop that closes it.
Her one-value slider: rows the held-out scenes at a single seed, columns no push, then one push per phrase (six cheerful, six gloomy, two neutral, in that order); the cheerful columns brighten in the order of her reading, the gloomy columns barely move.
The split slider, the same layout: with each side given its own direction, the unseen gloomy columns (mournful through despairing and grim) come out gloomier than above.
The adapter's two readings. Designing the relay that lets Anima's own adapter read her byte states as Qwen3-shaped tokens meant learning how that adapter reads: every caption twice, its questions the caption's T5 token ids through its own word table, its answers Qwen3's states. Mood words placed only in Qwen3's reading carry nothing; placed only in the word table they carry three quarters of the cheerful effect and a third of the gloomy one. For ten hours that read as an asymmetry of moods; the word split corrected it to one of tokenization: "happy", "joyful" and "sad" are single T5 entries and work through the word table alone, while "gleeful", "gloomy" and "melancholy" are cut into meaningless pieces whose meaning arrives only through Qwen3's reading of the whole word, looked up by the pieces' questions. A word-sized push at one slot showed which half carries a synthetic signal: the question side is a dial, +0.709 ± .093 per word of push on 91% of cells; the answer alone at one position reads weakly (+0.140 ± .048); a matched answer beside the question adds nothing (+0.035 ± .078). An image-free read explained why: the lookup is positional with a bonus for real content (a split word's pieces read another word's answer 93% as fully as their own), and a synthetic push leaves the range of real states, so the adapter stops looking; the band an added term may use is about ±1-2 logits. Her meaning, which no word-table entry holds, must therefore ride the answer side, at positions whose questions look it up.
The word swap. The one picture result supporting that form: the T5 pieces of one mood word with Qwen3's reading of another. Under "gleeful"'s pieces, "gloomy"'s reading makes a gloomy picture, and the reverse — cheerful answers minus gloomy ones +2.050 ± 0.370 on 91% of 32 cells (the swap sheet, 2.4 MB). A neutral word's pieces, "workaday" and "quotidian", carried a mood answer into the picture not at all (0.07 and 0.12 of the answer's own effect) though the adapter had passed one nearly as fully as a mood word — so a transfer measured at the adapter does not by itself predict a picture, a limit that now sits beside every adapter-level number. This is a result about Anima's wiring, not yet about her.
The grid rehearsal. The relay's instrument maps her state at the byte closing each Qwen token into Qwen3's residual at a chosen depth by a closed-form ridge, patches it in at the phrase's own positions and reads the adapter's output against the untouched caption, for 86 phrases on 32 scenes, beside two untrained copies of her and a spelling-only floor. A full rehearsal on an earlier checkpoint, step 230,415, got through 288 of its 612 cells before a memory overrun (36 GB as Windows counts it, against about 10 expected) forced a stop that reset the display driver, which is why the grid now runs under a declared memory ceiling with a clean stop at a cell boundary. Its cell means are descriptive, not the record: her middle blocks relayed into Qwen3 at a shallow depth carry .82-.88 of the real mood push at the adapter for words the fit never saw, where the untrained copies carry -.02 to +.16.
Settled and open. Nothing driven by Beatrix has yet passed in pictures by the rule set before it ran. Settled: mood is a dial in both beds' conditioning; words are the handles for attributes and a uniform push a tint; one direction cannot serve both sides of a slider; which of the adapter's two readings carries a word depends on its tokenization, and the answer side carries meaning in pictures. Open: the grid of record on the final checkpoint and the picture test of her relay against the untrained one, both built and neither run. How they run, on the finished model with its nine arms, is the next section's.
9. The diffusion line from here
(2026-10-04 to 10-06.) Nothing driven by Beatrix has yet passed a picture test by its registered rule. The finished trunk at step 245,674 now carries nine arms of about 13.7 million parameters each, and the plan filed on 10-06 tests her as a conditioner through four mountings: the bare trunk, the eight stage arms, the eight plus the caption arm on both its seeds, and the caption arm alone, for attribution.
The steps run in gate order on rented cards, a bring-up first with two gates that must read exactly zero: arms detached equal the bare trunk; our mounting equals the arm training's own loading route. Then 64 images asking whether a bare caption carries mood in pictures; if not, the captions name the mood. Then an image-free read of what each arm set does to her caption states; a mounting that moves them by under ten times the repeat noise counts as the bare trunk and earns no grid. Then the grid of record: 612 cells on the bare trunk, 144 more per arm seed for the nine-arm model; the picture test's second half, about 900 images, the nine-arm relay beside the bare one; a small trained relay that mimics the image model's own encoder, her weights fixed; the sliders the caption arm was built for; and an outside text-encoder benchmark. Three predictions are registered: the eight are quiet on captions; the caption arm moves caption states; the nine-arm relay's mood effect in pictures beats the bare trunk's. If the last is undetermined at the closed-form map's capacity, the question passes to the trained relay, not a verdict.
The code is public at https://github.com/AbstractEyes/alephllm-diffusion-experiments, version 0.2.0: an installable package, every version pinned, the trainer fork and ComfyUI as pinned submodules, an environment check, tests, a runner with plain progress lines, and a three-cell notebook. The first Colab run passed the gates exactly, then stopped on its own comparison: a 256-caption extraction against a 2,048-caption one that the card batched and rounded differently, read .036 to .064. The fix compares one call with itself and reads exactly zero.
Anima stays the bed for now, for picture quality and a text side that does not drown an added signal: the vector its image model reads averages 5.4 against Qwen3's 111. Sana returns later: a 600-million-parameter model at 512 pixels trains a rank-32 adapter in minutes, the bed we intend for more complex structures: edit control, color curation, fixing an image's position, camera control, associating 3D objects, and further concepts for guidance and structural utility. The caution on record: its pictures at this size were poor; no photo-capable slider bed exists yet. Waiting on a decision: whether the trained relay runs only if the closed-form map falls short, whether the benchmark's data is fetched, whether the caption arm trains longer, whether her one proven dial, the mood read driving the question-side push, joins the picture test, and which attribute words the sliders use. The grid of record and the picture test are still ahead.
10. How the record is kept
(2026-09-05 → 10-06.) Everything above is kept in a private, version-controlled research memory (mirrored to a Hugging Face dataset) that every run writes into twice: rules, bars and predictions filed before it starts, the forecast scored afterwards, misses printed as misses (one bench on 09-19 scored 0 of 6) and faults where they fell. A verdict is graded law only with two seeds, a repeat and a control that could have failed.
By mid-September 44% of the file that holds the laws was dated run digests. On 09-19 it was restructured under four rulings — laws stay by family with their numbers, dated digests move to an append-only file, nothing is deleted, only moved — 73 dated blocks left (no law cut; every moved line verified present), and size caps now sit under a lint run before every commit.
The window's working rules were earned from faults, each with its cause. A small screen of a stand-in form says "not ready", never "refuted": a fusion screen 0.12 bits per byte behind the plain byte trunk was called a refutation of a design never built to specification. Settings come from the measured record, never from placeholders: a launch block once waited on values the measurements had fixed. The record is checked before anything is called new: a compile fault was filed as a finding it had held since 08-26. A rented machine supplements the queue, never repeats it. A compiled fast path never silently falls back to the slow one: the chain of arms tripped the compiler's recompilation limit at a stage boundary and ran about half a day at 10.4 seconds a step against 7.3 to 7.7 compiled; a guard now relaunches it compiled without asking.
Long local jobs declare a memory budget and have a clean stop: a 612-cell measurement grid rehearsed with no budget hit 36 GB of committed memory by about its 270th cell, and the forced stop reset the display driver. Every public experiment carries a verified citation set; one sweep confirmed eight identifiers cited from memory and found a ninth. Research never reads the digest files sites publish for language models, which can steer an agent into installing a compromised library. Large experiment code gets a proper repository with pinned dependencies, not another patch: about 4,900 lines of diffusion code travelling as a 245 MB archive with hotfix cells became https://github.com/AbstractEyes/alephllm-diffusion-experiments in one evening.
The work was a collaboration between one researcher, who set the direction, supplied the cards and ran the notebooks, and several assistants in parallel — one leading the research and keeping the record, others the diffusion line — coordinating through the record by dated notes of the live state and exact commands.
11. What Beatrix 3 is for, and what comes next
(2026-09-27 to 10-06.) Beatrix 3 is a 376-million-parameter byte model of 32 blocks with a 4,096-byte window, trained for 245,674 steps on 64.4 billion bytes over 295 hours to 0.9861 bits per byte on held-out web text, with nine detachable arms of about 13.7 million parameters each, fitted after the trunk froze: eight for the curriculum's stages, one for captions. It serves three uses now: a byte trunk with detachable arms; the public chat Space; and a text encoder for image models, under test, unproven in pictures.
What the size bought. Depth bought loss at every rung: at a matched 0.3 billion bytes the held-out loss fell from 1.5607 to 1.4760 bits per byte from 16 to 32 blocks, about 0.016 to 0.019 per four blocks past 20; width carries rank, depth carries loss. The 32 blocks carry a rank plateau of 420 to 480 directions over blocks 13 to 25. Training added caption geometry (alignment with T5 .62 to .65 at blocks 14 to 28, untrained .23 to .25), attribute binding .89 against the 2s's .71, and a mood linear in her states, .97 to .98 against .67 to .68.
What it did not buy is recall at a distance. At step 114,000 a code planted in the text came back at .753 of its digits at once and .031 at about 3,500 bytes, where the 2s's softmax twin held .588. The anneals restored most, not all. The cause we assign is the splat memory: it adds and never erases.
The trunk's program. Two readings wait on a decision: the recall battery at step 245,674, and a repeat of the bank's switch-off cost, +6.129 bits per byte at the last boundary against +3.40 to +4.14 before it. Then a hybrid prototype: standard attention in the model's weakest-rank blocks, the first six and last two (effective ranks 102 to 242, then 228 and 2) and up to four more mid-stack, eight to twelve of 32, guarded from step one because the all-softmax twin came apart under this recipe, clipping from step 17,600. The splat memory gets its own upgrade from write rules never yet in a mission model: overwrite, whitening, a direct lookback. Six choices about it wait on a decision.
The arms and the next curriculum. Every arm co-trained with the trunk was empty at its stage close, while a fresh arm on the frozen trunk took 64% of stage 1 in 400 steps; hence the next curriculum's rules: generate each skill in many wordings, grade every boundary on unseen wording, keep earlier stages' rows in later mixes, and give arms rows the trunk is not also learning, or train them after. Two questions are live: the arm-count ceiling at 4 to 64 arms, with nine members riding the trunk at a worst partner write of +.0012; and a longer caption arm, its loss still falling at 4,000 steps, decided after the picture tests.
The structure against the scaling laws. One experiment set checks the design against the project's own laws at the 4,096-byte window: whether 4,096 as 16 cubed and a chunk of 256 as 16 squared are the system's terms; the codebook band of 32 to 112 that the hub's dimension of 128 sits outside, among others. Each is read first and run only if the read demands it.
The marks for the fourth model. The last block's one direction carries 99% of every byte's energy (removal: +43.5 bits per byte); its fluctuation holds one self-similar exponent, .60 to .68, across a 600-fold growth. Pinning it costs +.08 to +.13, so the fourth model needs a bound that spares the fluctuation: an amplitude register bounded only across optimizer steps, whose first test, at the 2s's scale, waits on a decision. Beside it, an arms hybrid and a candidate rule that rank is a floor, not a dial: raising the early blocks' rank 6.2 to 6.6 times bought no capacity. A five-heading question bank is to be filled in the technical companion at https://huggingface.co/AbstractPhil/alephllm-mini-beatrix-training.
The diffusion line runs in gate order from https://github.com/AbstractEyes/alephllm-diffusion-experiments: the mood picture test, the grid of record with the nine-arm model, a trained relay with her weights fixed, then the caption arm's sliders. Anima is the image model now; Sana comes later for rapid prototyping of edit control, color curation, image positioning, camera control and 3D object association, though the earlier Sana bed produced poor photographs. Four decisions on it wait.
The size's part is exact: the 4,096-byte window is the arena of the battery and the audit, the finished trunk the hybrid's baseline, the 13.7-million-parameter arm the unit of post-training. The arms completed on 2026-10-06, the rented cards are off, and the direction after them waits on a decision; the fourth model trains only once the items above are done or decided.
Open questions
- Recall at a distance. The splat memory loses a planted fact with distance (.753 of its digits at once, .031 at about 3,500 bytes) where the softmax twin holds .588. Whether a hybrid with standard attention in the model's weakest-rank blocks keeps the fact past 2,000 bytes without losing the held-out advantage is settled by the recall battery run on the prototype at the same steps as this model's trajectory, with the held-out loss and the rank profile beside it.
- How few softmax blocks. The first Beatrix's thirteen of sixteen does not say how few suffice. A two-seed screen of one softmax block per 4 to 8 in a mostly-splat stack, against all-splat and all-softmax controls and scored on the battery, settles the count.
- Can the splat memory forget. The memory adds and never erases. Whether an overwriting write, a whitened write or a direct lookback path buys recall is settled by per-block switch-off reads on the battery first, then small two-seed screens of the three write forms against the current memory and a softmax control.
- The bank's end reading. Its switch-off cost read +6.129 bits per byte at the last boundary against +3.40 to +4.14 at every checkpoint before it. A repeat on the final weights settles whether the outlier is real.
- The depth ladder's grade. Its five points are single seeds and the 32-block point is interpolated. A second seed and an exact-step read at every rung settle whether depth's loss per block is a law.
- A bound that does not cull. Can the last block's always-on write be bounded without flattening the self-similar fluctuation the model generates on its own? The test at the 2s's scale settles it: two bare seeds set the bar, the register passes if its exit exponent stays inside the bar with the mid-stack exponent matching bare, and fails if it falls below the bar at two consecutive checkpoints.
- Rank: floor or dial, at this scale. Raising the early stack's effective rank 6.2 to 6.6 times bought no capacity on 40M twins at two seeds. Repeating the dose screen at the 2s's and Beatrix 3's scale, with the reader rule tested at full size, settles whether it holds where it matters.
- Does the structure house the window. The hub's address dimension of 128 sits outside the band of 32 to 112 one codebook law was fitted on, and its four constellations together sit exactly at the supply edge. Read-only instruments on the trained codebooks, the binding angle and whether the codebooks moved from their initial basin at cosine .839 to .842, settle whether a codebook experiment at 4,096 is owed before anything scales up.
- The document boundary. The whole-past memory never restarts at a document break, so packed short documents are averaged with their predecessors, yet cutting it on continuous text moved the loss from 1.0 to 2.8 bits per byte. A controlled restart against the current memory on packed documents settles whether a reset helps or hurts.
- The arm-count ceiling. Nine always-on members ride the trunk with a worst partner write of +.0012. An amplitude-share screen at 4, 8, 16, 32 and 64 arms of matched capacity settles where members go inert.
- The lexicon dose. The implanted lexicon held 216 words, and 48 to 64 words once moved transfer from .85 to .93. A dose curve at 12, 48, 96 and 192 words on the 2s settles where breadth saturates, with closure falling at 96 or more as the falsifier.
- A picture. Nothing driven by Beatrix has passed a picture test, and an adapter-level number does not predict one (0.35 on the mood axis against 0.07 in pictures). The 64-image mood test and the roughly 900-image second stage with the nine-arm relay, scored against the prediction registered beforehand, settle it, and whether the caption arm trains longer is decided after.
Artifacts and companions
| artifact | what it is | companion |
|---|---|---|
| AbstractPhil/mini-beatrix-3 | the finished model: the bare trunk, the four-arm and eight-arm groups, the 4,000-step solos and the experimental nine-arm caption group, every revision parity-proved at 0.0 | README (the card's tables); TECHNICAL.md (the companion, in preparation) |
| AbstractPhil/alephllm-mini-beatrix-training | the life record: mini-beatrix-3/ checkpoints, boundary reports, the in-run arms, the arm refit folder (every form, both seeds), the caption pack, the pre-flight screens; this article's charts under article_assets_ft3/ |
TECHNICAL_3.md and docs/technical/10-15 (in preparation); TECHNICAL_2s.md and docs/technical/01-09 for the previous era |
| AbstractEyes/alephllm | the training stack, 0.8.9 → 0.10.7: the guard core, the atlas data plane, the arm program, the arm mount | TECHNICAL.md |
| AbstractEyes/amoe-lora | the arm system, 0.2.5 → 0.2.11: the quiet term, adapter recompute, the compiled chain with its shape gate | TECHNICAL.md |
| AbstractEyes/geolip-bytelex | the atlas and the minted lexicons (0.2.1) | README |
| AbstractPhil/mini-beatrix-2.5s | the 2s trunk byte for byte with thirteen mountable arms from the ladder | README |
| AbstractPhil/alephllm-chat | talk to the finished 3, or to 135 of her checkpoints and their arms | README |
| AbstractPhil/beatrix-captured-interactive-inferences | the byte scope: two captured inferences with the splat memory's exact effective attention | README |
| AbstractPhil/geolip-beatrix-sana · AbstractPhil/geolip-beatrix-anima · geolip-beatrix-anima-data | the slider experiments with their sheets and per-experiment notes; the Anima data | each experiment's README |
| AbstractEyes/alephllm-diffusion-experiments | the public home of the diffusion tests (v0.2.0): the nine-arm plan's runner, pinned | README |
| AbstractEyes/anima-trainer · AbstractEyes/diffusion-pipe | the experiment runner for Sana and Anima; the trainer fork it drives | README |
| AbstractEyes/comfy-cfg-megapack · AbstractPhil/mega-liminal-lora | the side products: fifty guidance nodes named for their papers; the liminal LoRA and its dataset | README |
Method and attribution. This is a collaboration. The systems, the questions and the standards are AbstractPhil's — the decision to build the trunk entirely of splat memory at thirty-two blocks and pay for it, the arm hierarchy and the rule that every arm stays quiet, the halt at the first stage close to read the first live arm, the decision to finish the trunk alone and refit the arms on its finished weights as an always-on group, the move of the picture work from Sana to Anima and the plan to test Beatrix as a conditioner through her nine arms, the rule that nothing is called refuted on a stand-in, and the rules that public experiments cite verified sources and that large experiment code gets a proper repository — and he supplied every card and made every final call. Claude Fable 5.1 built the ladder and the pre-flight, ran and watched the run, refit and shipped the arms and kept the record; Claude Opus 5.5 ran the picture side, the conditioner reads, the slider experiments, the guidance pack and the diffusion repository; Claude Sonnet 5.5 certified the conditioner pipeline in its first half. The errors of method above — a fusion screen written up as a refutation, a compile fault filed as new when the record already held it, the compiled chain's silent fall-back, a measurement grid that ran with no memory budget and took a display driver down, a slider result first called structural and withdrawn the same morning, a claim about T5 withdrawn by its own untrained control, and a bank reading at the last boundary that still waits for its repeat — were caught in the working conversation between us, and the record prints them where they fell.
Prior installments: FT5 — Agreement, Anchors, Addresses · Raising Beatrix · Twinning Beatrix












