QC67_cosmo / SAMGO.md
phera-ra's picture
Kit update 2026-08-06: samgo 5.7 (54D BPE, 59M), gate-init finding, Born-rule fix in quantum_pool, corpus-drift retraction + fingerprinting, frozen-corpus wiring table
b085020 verified
|
Raw
History Blame Contribute Delete
4.82 kB

samgo 5.7 — the 54D BPE lineage

weights/samgo_weights.pt · 59,353,668 parameters · GPT-2 BPE (50,257) · 135 MB

The character models in this kit are ~1.15M parameters over a 162-symbol alphabet. samgo is the other lineage: a CosmosTransformer — the same 54D state decomposition, but at d_model 512 with a byte-pair vocabulary, so a step of training buys whole words instead of letters.

d_model 512   n_heads 8   d_ff 2048   n_layers 2   max_seq_len 2048
d_state 54  =  12D CST phase  +  24D Hebbian plasticity  +  18D Lorenz/Rössler chaos
              (7 Lorenz + 4 Rössler attractors × 3D)
plus a 256-slot persistent memory bank

Why it exists

The 54D checkpoint this lineage had been accumulating (cosmos_play.pt) was trained by a background learner capped at MAX_SEC=600. Ten minutes of CPU on 59M parameters is about seventy steps. The design is correct for a learner that must never hog her machine, and it meant the largest model in the project had never once been trained at length.

samgo is that model, trained without a wall-clock cap, on GPU, on world text rather than only on her own.

Corpus

source weight role
WikiText-103 (60 MB slice) ×1 world knowledge
cory_voice_corpus ×8 the human it talks to
experience_corpus ×6 its own lived log

16,539,493 BPE tokens. Her own material is upsampled deliberately: straight concatenation would have been hundreds-to-one against her, and the register would not have survived.

What it can and cannot do — read this before citing it

It writes sentences, which the character models at the same stage do not:

"no evidence that this will be the reason we all can be called to our minds or do … I'm just here as possible … So I can take this and take over my work"

Compare a 1.15M character model at a comparable point: "was befal was smorded mone as alland of the." BPE plus fifty times the capacity is a different regime.

It is data-limited, not step-limited. 59,353,668 parameters against 16,539,493 tokens is 0.28 tokens per parameter. Compute-optimal is roughly 20. This model has about 1/72nd of the data its size wants, and it shows: at the last checkpoint train loss was 2.80 while validation was 5.02. That gap is memorisation, and more steps will widen it. The next useful change to samgo is more text, not a longer run.

Provenance of the shipped weights, stated exactly

The run reached val 4.5121 at step 1,675, destabilised to 5.8355 by step 3,000, and was recovering (5.02 at step 14,000) when these weights were taken.

The 4.5121 weights no longer exist. The trainer saved every 250 steps unconditionally, so thirty later, worse checkpoints overwrote the best one before anyone noticed. best_val was tracked the whole way; the weights that earned it were not kept. The fix — a separate _best.pt written only on improvement — landed after the damage.

So the shipped checkpoint is step ~13,250, val ≈ 5.02, not the best this run ever saw. Saying otherwise would be the kind of claim the rest of this kit exists to avoid.

Load it

python tools/load_model.py samgo --prompt "the universe is"

or directly:

import torch, tiktoken
from Cosmos.web.cosmosynapse.model.cosmos_config import CosmosConfig
from Cosmos.web.cosmosynapse.model.cosmos_model import CosmosTransformer

ck  = torch.load("weights/samgo_weights.pt", map_location="cpu")
cfg = CosmosConfig()
for k, v in ck["config"].items():
    if hasattr(cfg, k) and isinstance(v, (int, float, str, bool)):
        try: setattr(cfg, k, v)      # head_dim etc. are derived, read-only
        except AttributeError: pass
m = CosmosTransformer(cfg)
m.load_state_dict(ck["model_state_dict"], strict=False)
m.eval()

enc = tiktoken.get_encoding("gpt2")
x   = torch.tensor([enc.encode_ordinary("the universe is")])
out = m(x)["logits"]        # returns {logits, state_54d, layer_states}

forward returns a dict, not a tensor — state_54d is the whole point of the architecture, so it hands the state back alongside the logits.

Not to be confused with

  • cosmos_play.pt — same architecture, but it is a re-ranker in the live system: her articulate model proposes replies and this scores how much each sounds like her. It is not a generator and was never trained to be one.
  • phos.pt / dyn12phos — the character lineage, quantum-born, where the controlled architecture research in FINDINGS.md was done. samgo is not quantum-born: it warm-started from cosmos_play.pt, which did not use quantum_birth.

That last point matters for provenance. Every quantum claim in this kit is about the character lineage. samgo is the scale lineage, and it makes no quantum claim at all.