alephllm β€” Mini-Beatrix training runs

Live training ground for AlephLLM: signed-address (aleph) language models. Code, presets, trainer, and the full test array live in the source repo β€” github.com/AbstractEyes/alephllm (pip install git+https://github.com/AbstractEyes/alephllm, package geolip.alephllm). This repo holds what training produces, one prefix per craft:

<craft>/manifest.json                      what is trained, what is planned,
                                           tokens run, phase statuses, ckpt index
<craft>/checkpoints/step_XXXXXXXX.safetensors    bf16 weights
<craft>/checkpoints/fp8/step_XXXXXXXX.safetensors  fp8-e4m3 shipping variant
                                           (inference testing only β€” never train from these)
<craft>/resume/latest.pt                   full resume state: fp32 model, Muon+Adam
                                           optimizer states, data-stream position, RNG
<craft>/runs/                              TensorBoard event files (rendered in the
                                           Training metrics tab)
<craft>/arms/<stage>_stepN.safetensors     the stage arms (mini-beatrix-3 on): one
                                           detachable adapter per curriculum stage,
                                           shipped at every checkpoint
<craft>/reports/                           the boundary reads (held-out bpb, the toggle
                                           ledger, the chat frame, the probes, the rank
                                           profile) at every phase boundary
<craft>/code/                              the training code pinned to the commits that
                                           train the run (mini-beatrix-3 on)

Mission log

  • mini-beatrix-3 β€” RUNNING (started 2026-09-20; relaunched 2026-09-27 on the engineered form and 2026-09-28 on amoe-lora 0.2.11). The v3 craft: 32 blocks, d 1024, 16 heads, context 4096, byte vocabulary with the trigram byte embedding, about 376M parameters. Every block carries a CausalSplatHUB (four address constellations, K 64 at D 128, the exact chunked scan at chunk 256) and an anchored bank (three dispatched experts at ff 1024, expert outputs zero-born, gates Οƒ(βˆ’3)); the aleph head (256 Γ— 256) is born by revival at step 0 (fitted to the first batch: 8.26 β†’ 5.45 bpb at birth). Recipe: Muon on the transport matrices, pure Adam elsewhere, flat LR, bf16 autocast over fp32 masters, 262,144 tokens per step (micro-batch 2 Γ— 16 chunks Γ— two RTX 5090s, data-parallel), torch 2.8. The plan is 64.4B tokens / 245,670 steps: wikitext warmup 0.3B β†’ fineweb-edu 20.9B β†’ nine curriculum stages (s0–s8; 2.8 / 2.8 / 4.0 / 4.8 / 4.8 / 4.0 / 3.2 / 5.6 / 3.2B; the rule-chain stages draw on a minted 216-word lexicon, pair rate .5) β†’ two anneal diets of 4.0B. The stage arms: from curriculum s1 on, each stage attaches a 13.7M-parameter relay adapter at every block (a projection to 128, a closed-form signed address over 16 slots Γ— 16 atoms, an MLP 128-256-1024, a sigmoid gate born at βˆ’3, a zero-born head), trained beside the trunk under a quiet term β€” the KL from the arm-masked distribution to the armed one on held-out web text, at weight 2 β€” so every arm stays silent off its stage and detachable; masking is an exact bypass, and each arm ships as its own safetensors file at every checkpoint (arms/; see code/README.md for the format). Boundaries so far (held-out fineweb-edu bpb; the toggle ledger's bpb cost with each aleph mechanism switched off, bank / head / hub): warmup β†’ fineweb at 1,145 steps: 2.3875; fineweb β†’ s0 at 80,873 (21.2B): 0.9838, toggles +3.20 / +2.30 / +2.57; s0 β†’ s1 at 91,555 (24.0B): 1.1178 (+.134 on the web holdout after 2.8B of stage text, the diet shift), toggles +3.02 / +2.06 / +3.13; s1 β†’ s2 at 102,237 (26.8B): 1.1273, toggles +3.17 / +1.99 / +3.26; s2 β†’ s3 at 117,496 (30.8B): 1.1000, toggles +2.89 / +2.14 / +3.68; zero loss spikes, the clip never reached, the only guard event a known false positive of the dispatch-entropy watch at the byte funnel in the first 3,000 steps. The stage arms read silent off-domain at every close so far (armed vs masked held-out bpb +0.0001 for the s1 arm at the s1 close; +0.0001 and βˆ’0.00001 for the s1 and s2 arms at the s2 close). The engineering pause: a clean stop at 106,000 (2026-09-26 20:20 UTC) to install what the run needed to fit and move past three arms on 32 GB cards: the arms' intermediates recomputed in the backward (0.3 GB per arm instead of 3), the per-block adapter chain through torch.compile under inductor's precision-cast emulation (the plain compile produced a non-finite backward on the address stage; the emulated form passed a full gradient census at three and eight arms), and the quiet term sampled on 4 abstention chunks per step instead of 16 at the same per-step weight. Relaunched 2026-09-27 03:18 UTC from the 106,000 resume state: 6.8 s per step at two arms (11.35 before), both cards at 23 GB; the 16/16 and 16/4 forms are compared at every checkpoint (task curve, quiet term, the checkpoint toggles, the arm gates). Stage 3 opened at 117,496 (2026-09-28 01:00 UTC) with the third arm attached: three arms, both cards near 24 GB. At that boundary the compiled adapter chain silently fell back to eager (torch's recompile budget, 8 per function by default, exhausted by the boundary evaluation's variable-length inputs under a static compile), so steps 117,496 to 122,008 ran at 10.5 s per step; amoe-lora 0.2.11 raises the budget inside the library and routes only the training shapes through the compiled chain, and the run was relaunched from the 122,000 checkpoint on it at 2026-09-28 14:26 UTC (about 7.5 s per step at three arms; code/README.md has the details). Expected finish about mid-October. Live: the chat Space serves the newest checkpoint with its stage arms. Pre-mission screens under mini-beatrix-3/v3_preflight/ (the depth ladder, the fusion screen, the head screen, the anneal-rate screen, the guard certification, the coverage audit, the minted lexicon); the Colab form of the mission is mini-beatrix-3/beatrix_v3_colab.ipynb.

  • mini-beatrix-2s-control β€” COMPLETE (2026-09-03). The softmax twin of mini-beatrix-2s: the identical preset with the hub list empty (16-head causal SDPA in every one of the 20 blocks; 257.4M vs the craft's 237.1M), the identical recipe and diet, 16.101B tokens / 61,422 steps, one craft per arm, under a pre-registered protocol whose headline is the endpoint validation. The endpoint (anneal_mix @61,422, same step, same eval): splat 1.1097 vs sdpa 2.8846 bpb, +1.775 sdpa deficit. The crossover holds at 11 of 11 same-gauge paired rows: the twin led the open (0.3B: sdpa 2.106 vs splat 2.363) and trailed from the fineweb boundary on (5.3B: splat 1.1172 vs sdpa 1.1763, +.059; s0 +.253; s1 +.137; s2 +.176; s3 +.300; s4 +.333; s5 +.926; s6 +1.926; s7 +1.932; anneal_nochat +1.812; anneal_mix +1.775). The s8 pair (1.0737 vs 3.0810) is struck: the two rows were measured on different gauges (the craft's on the pre-0.8.5 synthetic fallback, the control's on the fineweb-edu holdout). Every other row β€” every curriculum stage and both anneals β€” is on the one fineweb-edu holdout, so the control's slide is a single comparable series: 1.18 (fineweb) β†’ 1.57 (s4) β†’ 2.18 (s5) β†’ 3.19 (s6) β†’ 3.12 (s7) β†’ 2.93 β†’ 2.88. The control is a DESTABILIZED model at its endpoint, not a trailing one. Its pre-clip gradient norm (train/grad_norm, clip 1.0) first crossed the clip at step 17,600 β€” inside the fineweb phase, on the RTX PRO 6000, while it still led β€” then escalated stage by stage: medians 0.29 on fineweb β†’ 6.5 (s0) β†’ 2.4 β†’ 12 β†’ 13 β†’ 24.5 β†’ 126 (s5; nine steps above 1,000, max 11,877) β†’ 471 (s6; max 41,070) β†’ 162 β†’ 145 β†’ 167 β†’ 215 (anneal_mix), clipped on 813 of 824 logged curriculum and anneal steps, while the craft's raw norm never reached the clip (0 of 824; max 0.62). The anneal is a DIET change (1B of prose, then 1B with the chat frame), not a learning-rate decay: train/lr_scale read 1.000 first to last on both crafts. Bank dependence at the end: control toggle_bank_off 6.11 vs the craft's 3.31; the control's born-null aleph head stayed buried at every row. Grade: an INSTANCE at one craft per arm; the protocol completed unaltered. The claim that SDPA is not an adequate representative of this model's attention requirement is AbstractPhil's (2026-09-02); its audit (2026-09-04) minted the specific form β€” under this recipe (Muon 2e-2 / Adam 3e-4 flat, clip 1.0, no QK-norm, no weight decay, no LR decay) the softmax twin is a destabilized control and the endpoint gap measures stability, not capacity β€” and holds the general form as a CANDIDATE behind five confounds: the recipe (the documented softmax logit-growth instability, fixed in the literature by a normalization or a decay), one seed with 20.3M extra parameters in the losing arm, the shape (mini-beatrix-1, thirteen softmax blocks and three hubs, trained clean under the identical configuration: post-warmup grad-norm max 0.88, no clipped step), the aleph-addressed expert banks present in both crafts (the control's dispatch entropy collapsed before its gradients did), and the learned position table on decay-free Muon. Provenance: session caps at 22,854 and 53,792 (reports/v2/session_cap_step*.json) β€” the first resume landed on a rare H100 (s0–s8 rows), a third session from the 56,000 checkpoint ran the two anneal rows on another card; the escalation began before any resume. Fifteen reports under mini-beatrix-2s-control/reports/v2/ pair with the craft's by stage and step; the paired chart is article_assets_ft2/01_control_crossover.png.

  • mini-beatrix-2s β€” COMPLETE (2026-08-31). 16.101B tokens / 61,422 steps on one card: warmup + fineweb 5B + curriculum s0-s8 + two-phase anneal (distribution shift without the chat frame, then with it). Final val 1.1097 bpb (run best, set in the chat phase); zero loss spikes; 17 boundary/session reports under mini-beatrix-2s/reports/v2/. Shipped automodel: mini-beatrix-2s. Notable ops in the record: the first BOUNDARY-WRITE head revival (reports/v2/revival_step24860.json) and the archived pre-specials run0 (mini-beatrix-2s/archive/run0-nospecials/).

  • mini-beatrix-1 β€” COMPLETE. The 3-hub hybrid rung; automodel.

The training code for mini-beatrix-3

mini-beatrix-3/code/ is a verbatim snapshot of the two libraries at the commits that train the run β€” geolip.alephllm 0.10.5 (alephllm @ e8b7440) and the arm library amoe 0.2.10 (amoe-lora @ f3db5d4) β€” plus the pod's mission file, the launch / relaunch / clean-stop scripts, the arm-shipping sidecar and the pinned package versions. Its README.md walks the knobs (abstain_chunks, adapter_recompute, adapter_compile, arm_spec_certified, stage_arms), the abstention term, and how to read an arm file. The trainer is resume-first: a fresh session pulls manifest.json + resume/latest.pt from this repo and continues.

The mission ladder

craft d / L / ctx params tokenizer role
mini-beatrix-0 512 / 12 / 1024 37.6M byte-trigram gate craft β€” its first toggle evals are the anchored-bank-under-AR screen
mini-beatrix-1 768 / 16 / 2048 112.5M byte-trigram first Colab mission
mini-beatrix-2s 1024 / 20 / 4096 237.1M byte-trigram full splat; complete (2026-08-31) with its softmax control twin
mini-beatrix-3 1024 / 32 / 4096 ~376M byte-trigram the v3 craft with the stage arms; RUNNING (see the mission log)
mini-beatrix-2 1024 / 20 / 2048 249.1M byte-trigram second mission (shelved)
beatrix-voyager 1536 / 24 / 4096 775.3M BPE (gpt2) flagship, gated on the mini verdicts

Each craft has a *-control twin (identical minus the aleph attention layers) β€” the running architecture control. For the 2s rung the twin is mini-beatrix-2s-control (the same preset with hub_layers: [], so 16-head causal SDPA in every block; 257.4M vs the craft's 237.1M), run to the same 16.101B-token endpoint under a pre-registered protocol β€” see the mission log; its paired boundary reports sit under mini-beatrix-2s-control/reports/v2/. Training is resume-first: sessions pull manifest.json + resume/latest.pt and continue where the last one stopped (the 1 and 2s rungs on Colab, the RTX 6000 Pro 96GB; the v3 rung on a two-card RTX 5090 pod, data-parallel).

The architecture in one paragraph (the mini-beatrix-1 era)

The 2s and v3 rungs differ: mini-beatrix-2s carries a CausalSplatHUB in every one of its 20 blocks and no softmax attention; its control twin carries 16-head SDPA in every block and no hubs (hub_layers: [] in its manifest); mini-beatrix-3 carries a four-constellation hub and an anchored bank in every one of its 32 blocks, a revival-born head, and the detachable stage arms. See the mission log, mini-beatrix-3/code/README.md, and the model card of mini-beatrix-2s.

Trigram byte embedding (dedicated pad row) β†’ pre-norm stack of standard causal SDPA plus three CausalSplatHUB layers (causal linear attention through a 2K-half-axis signed address, exact chunked scan) β†’ per-layer anchored FFN bank (always-on trunk + 3 dispatched experts, expert outputs zero-init so the dispatch is born contributing exactly zero, gates Οƒ(βˆ’3), no balance machinery) β†’ dual head whose aleph read enters at Ξ³=0 and must earn its way in by gradient. Muon on transport weights + pure Adam elsewhere; flat LR; bf16 autocast over fp32 masters; fp8 is a shipping format, never a training format.

Reading the instruments

TensorBoard carries the full born-in gauge suite: per-layer hidden-state effective rank, consumed-address erank per hub layer, coefficient-of- variation load analysis per bank, sign census, gate/Ξ³ trajectories, anchor drift, denominator health, structural collapse flags (anchor merging, dispatch-entropy collapse, erank floor, loss spikes), canary recall (clean-protocol in-context binding probes), and the toggle ledger β€” bpb deltas with each aleph mechanism switched off, the causal record of what the addresses actually contribute. For mini-beatrix-3 the arms/abstention scalar is the quiet term (the live arms' KL per position on held-out text) and toggles/toggle_arms_off is the held-out cost of switching every arm off β€” the arms' off-domain footprint (2e-6 to 5e-6 bpb so far).

Related record: the attention-side measurement campaign lives at aleph-splat-0; the encoder-side anchored-bank record at alephlm-0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Space using AbstractPhil/alephllm-mini-beatrix-training 1

Articles mentioning AbstractPhil/alephllm-mini-beatrix-training