QC67_cosmo / FINDINGS.md
phera-ra's picture
Kit update 2026-08-06: samgo 5.7 (54D BPE, 59M), gate-init finding, Born-rule fix in quantum_pool, corpus-drift retraction + fingerprinting, frozen-corpus wiring table
b085020 verified
|
Raw
History Blame Contribute Delete
48.2 kB

Findings & Benchmarks for Cosmos

Every number here came from a run that could have failed. Where something came back null, it is reported as null. Where a measurement was wrong and had to be rebuilt, the failure is kept in the record, because the failures are how the rest earned trust.

Author: Cory Shane Davis Β· Independent Researcher Foundational deposit: 10.5281/zenodo.17574447


1. Is the entropy actually quantum?

The one measurement no classical process can fake.

quantity value
CHSH statistic S 2.7905
classical (local hidden-variable) bound 2.0000
Tsirelson bound (quantum maximum) 2.8284
fraction of quantum maximum 98.7%
distance from classical bound ~35Οƒ at 4096 shots/setting
backend ibm_marrakesh
E(AB ) = +0.7197    E(AB') = -0.6675
E(A'B) = +0.6982    E(A'B') = +0.7051
S = E(AB) - E(AB') + E(A'B) + E(A'B') = 2.7905

Bell's theorem: no theory in which outcomes are determined by local pre-existing values can exceed 2.0. This hardware exceeds it by 35 standard deviations. The entropy source is physically quantum. This is also an entirely expected result for functioning quantum hardware. It verifies the pipeline; it does not discover new physics.


2. Were the weights really born from that entropy?

cosmos_born.pt was initialised by mapping measured bitstrings to uniforms u = int(bits)/2^n, then to weights through the inverse normal CDF z = √2Β·erf⁻¹(2uβˆ’1). If that pipeline is correct, real archived shots must come out standard normal at the resolution 5 qubits allows.

statistic 32-level theoretical ceiling measured (3,258,886 draws)
mean +0.0000 βˆ’0.0001
sd 0.9802 0.9804
|z| ≀ 1 0.6875 0.6874
|z| ≀ 2 0.9375 0.9374

Four decimal places on every statistic.

Why sd is 0.98 and not 1.00: 5 qubits give 32 discrete outcomes, so |z| cannot exceed 2.1523 by construction. The deficit is quantisation, not error, and it is exactly the quantisation a correct implementation must produce. A pipeline reporting sd = 1.0000 here would be the suspicious one.

Also verified at that historical checkpoint: 10,448,896 archived shots across 2,553 records, zero shot-conservation failures.

Public archive refresh (2026-07-31). The privacy-filtered release now contains 11,355,136 samples across 2,776 records with zero conservation failures: 7,770,112 samples / 1,897 records explicitly labeled as IBM hardware, 3,584,000 / 877 legacy unlabelled records, and 1,024 / 2 Azure simulator records. Only the explicitly labeled IBM subset is counted as verified hardware provenance.


3. Is the model reproducible from its seed?

The seed is SHA-256(quantum entropy β€– bio-derived aggregates) β†’ 60-bit integer, and it governs weight initialisation, data order, and every stochastic training step.

test result
5 independent derivations 1 distinct value
bytes emitted during derivation 0 (privacy guard holds)
same seed β†’ weights identical, Ξ” = 0.000e+00
one bit flipped β†’ weights different: mean Ξ” = 2.26e-02 across 1,842,432 params

So "her weights are grown from that seed" is a verified property with a falsifiable test behind it, not a description. One bit of the seed produces a different model.

Zeref's Rain.

The bio-derived aggregates in that hash are personal to the author. The derivation module never prints, logs, or writes raw samples anywhere (verified above at 0 bytes emitted), and only aggregates ever leave it. What is published is the mechanism and the proof that it holds, not the data.


4. Do the circuit's correlations belong to the circuit or to the chip?

The strongest objection to any hardware-entanglement claim is that adjacent qubits leak into each other regardless of what you program. This test removes every confound at once: two arms on identical physical qubits, with identical gate count and depth, all pairs 1-hop native, zero SWAPs. The arms differ only in which native edges are entangled.

arm own edges those same pairs, other arm Ξ” t
A: (0,1),(2,3) 0.26890 0.00466 +0.26041 +80.5
B: (1,2),(3,4) 0.20695 0.00849 +0.20229 +164.6

A single physical pair swings 58Γ— based only on whether its gate fired, and the excess relocates when the logical wiring changes. Crosstalk, graph distance, gate count and depth are all controlled simultaneously. 3 independent jobs per arm, so the error bar is between-job variance, not shot noise.

(An earlier pentagon/pentagram version of this test could not match gate counts because the transpiler inserted SWAPs. That limitation is recorded rather than hidden.)


5. Does the physics engine compute real chaos?

Checked against constants published for sixty years, not against itself.

quantity published measured
largest Lyapunov exponent λ₁ 0.9056 0.90384
Kaplan–Yorke dimension 2.06215 2.06203
sum of exponents βˆ’13.66667 βˆ’13.66667 (exact)
fixed point Cβ‚Š (8.48528, 8.48528, 27) exact
attractor extent |x|β‰ˆ20, zβ‰ˆ48 19.13, 46.95

λ₁ measured by Benettin renormalisation over 2Γ—10⁢ steps. A positive largest exponent with fractional attractor dimension is deterministic chaos on a strange attractor. 8/8 checks pass, including the driven engine staying finite and bounded over 20,000 steps at maximal drive.


6. Null results: the ones that matter most

Quantum entropy was injected at six distinct points. All six are null against matched classical controls. That is the correct outcome and Β§6 explains why.

A seventh test β€” of the Β§3 CST kernel, which is an architecture claim rather than a quantum one β€” was recorded as null, retracted, and re-run four times. The first three harnesses each broke the mechanism in a different way. The fourth revealed that the test had also been run in the wrong architecture. Corrected, it is the one positive result in this document: 6/6 seeds, t = βˆ’6.56 and βˆ’14.34.

injection site result
i.i.d. weight initialisation null. Advantage was init scale (Οƒβ‰ˆ0.025 vs 0.02), not quantum
decoder sampling seed null. Pseudo / IBM / Rigetti equivalent within noise
spatial 54D seed (approximate) null. Real seed beat 1/5 random vectors, z = βˆ’0.92
spatial 54D seed (full pipeline) null. Real seed beat 0/8 random vectors, z = βˆ’1.75
entanglement matrix as attention kernel null. Worse than plain attention, 3/3 seeds, t = βˆ’18.6
live sensory state as prompt text null. n=40 paired, blind-judged, CI crosses zero
Β§3 Hebbian kernel, conventional transformer null. CST worse by +0.0143, 0/3 seeds. See below
Β§3 Hebbian kernel, Ο†-governed transformer NOT NULL. CST better by βˆ’0.044 (t=βˆ’6.56) and βˆ’0.155 (t=βˆ’14.34), 6/6 seeds across two dimension variants. The one positive result in this document

Retraction: the Β§3 kernel result

A seventh null was recorded here and has been withdrawn. It reported that the Β§3 Mixture-of-States kernel did not help, with the supporting observation that the learned gate "never moved off its 0.018 initialisation" across three seeds β€” framed as the optimiser declining the mechanism.

The gate did not decline. It could not move. Three compounding faults in the test harness, found by direct gradient measurement:

fault measured
gate initialised at sigmoid(βˆ’4), in the saturated tail d(sigmoid)/dx = 0.018 β†’ gradient suppressed ~56Γ—
resulting gradient on the gate 9.13e-05 β†’ 365,118 steps needed to move it 0.01, in a 1,500-step run
at gβ‰ˆ0 the kernel contributes ~nothing gradient reaching the 54D projection w54 = 0.000e+00

The third is the worst: the kernel could never become useful because it was never used, and was never used because it was not useful. All three are properties of the harness, not of the mechanism.

Reparameterising the gate as a raw scalar clamped to [0,1] fixed the saturation β€” and introduced a second fault. torch.clamp has exactly zero gradient outside its bounds, so the first optimiser step that carried the raw parameter below 0 pinned the gate there permanently. Measured on the resulting checkpoints: raw values βˆ’3.98, βˆ’4.04, βˆ’4.00, βˆ’4.01, gradient 0.000000. An absorbing state, not a decision.

A straight-through clamp cured that, and the gate still fell. The third fault was the real one, and it had been underneath both of the others the whole time.

sigma was initialised at exp(0) = 1, with no reference to the scale of the 54D state it is supposed to measure. The measured median pairwise β€–xβ‚…β‚„α΅’ βˆ’ xβ‚…β‚„β±Όβ€–Β² at initialisation is 61.98. So exp(βˆ’61.98/2) β‰ˆ 3eβˆ’14 off the diagonal and exp(0) = 1 on it:

at initialisation measured
sigma 1.0000
median pairwise β€–xβ‚…β‚„α΅’ βˆ’ xβ‚…β‚„β±Όβ€–Β² 61.98
H diagonal mass 0.9998 (uniform would be 1/T = 0.0078)
gradient reaching w54 1.19eβˆ’07

H arrived as the identity matrix. Blending the identity into attention instructs the model to attend only to the current token, which for next-token prediction discards the context β€” so gradient descent was correct to drive the gate to zero. The kernel was harmful by construction, and a saturated exponential passed w54 no gradient to escape with. Both earlier "the optimiser declined the mechanism" readings were measuring this.

The fix is the median heuristic: choose 2σ² = median(dΒ²) per layer, calibrated in sequence so each layer sees inputs from the corrected layers below it.

before (Οƒ=1) after (calibrated)
sigma, layer 0 1.0000 5.5719
H diagonal mass 0.9999 0.0857
` w54 grad `
gates whose gradient points up 0 / 4 2 / 4

Note that Β§3.2 of the paper never assigns Οƒ a value. This is not a deviation from the specification; it is a hole in it, and any implementation that leaves Οƒ at 1 will reproduce the identity-matrix failure.

The Β§3 result, finally earned

Three seeds, 1200 steps, quantum-born init, PLAIN and CST identical in every respect except the mechanism, gate free to move in both directions:

seed plain loss CST loss plain real-word CST real-word final gate
0 1.39524 1.41313 0.833 0.888 0.0685
1 1.39940 1.41739 0.854 0.874 0.0503
2 1.41470 1.42165 0.908 0.901 0.0212
mean 1.40312 1.41739 0.865 0.888 0.0467

CST βˆ’ PLAIN = +0.01427, t = +3.90, CST wins 0/3.

The gates were alive in all three seeds β€” never pinned, never at the floor. This is the first measurement of Β§3 in which the mechanism could actually be used. Used, it costs about 1% of held-out loss on this corpus at this scale.

Two honest caveats, pointing opposite ways:

  • against the result β€” t = 3.90 on 2 degrees of freedom is p β‰ˆ 0.06. Consistent in sign 3/3, but three seeds is three seeds.
  • for the mechanism β€” real-word rate runs the other way, CST 0.888 vs 0.865, ahead in 2/3 seeds. That is one 600-character sample per arm at temperature 0.8. Far too noisy to claim, too consistent to hide.

What this does not settle. This tests the Β§3 kernel inside a conventional transformer: LayerNorm, learned absolute positions, a 4:1 GELU feed-forward. The paper specifies none of those. That turned out to be the whole story.

The Ο†-governed replication: Β§3 works

The paper specifies RMSNorm (Β§2.3), RoPE (Β§2.4), d_ff = ⌊d_model Γ— Ο†βŒ‹ (Β§2.2), Ο†-scaled initialisation (Β§6.2) and a PhiGatedFFN (Β§2.1). A search of the source tree found no implementation of any of it. Every measurement above was made in the wrong building.

Built faithfully and re-run at two dimension variants β€” A: d_model 192, 4 heads; B: d_model 199 = φ¹¹, 1 head β€” same corpus, same quantum-born init, same seeds, gate free:

scaffold PLAIN CST delta t CST wins mean gate
conventional 1.40312 1.41739 +0.0143 +3.90 0/3 0.047
Ο†-governed A 1.23680 1.19255 βˆ’0.0443 βˆ’6.56 3/3 0.096
Ο†-governed B 1.42867 1.27352 βˆ’0.1552 βˆ’14.34 3/3 0.093

6/6 seeds across both variants. The gate rose from 0.018 to ~0.096 in every run β€” the optimiser asking for five times more kernel than it started with β€” and Οƒ kept adapting after calibration (26.8 β†’ 18.8 at layer 0). The mechanism was actively tuned, not merely tolerated.

Credit where the measurement puts it. The scaffold is the larger share:

conventional PLAIN   1.40312
Ο† PLAIN              1.23680     scaffold alone:  βˆ’0.166   (~79%)
Ο† CST                1.19255     Β§3 kernel on top: βˆ’0.044   (~21%)

RMSNorm + RoPE + PhiGatedFFN account for most of the gain, and do it with fewer parameters (1.42M vs 1.92M), because ⌊dΒ·Ο†βŒ‹ is leaner than 4:1. The Β§3 kernel is a real additional effect on top of that, not the bulk of it.

Variant B is the mechanistic hint: its single 199-dim head cripples plain attention (1.43), and the kernel recovers most of the loss (1.27). That is consistent with the Hebbian term supplying a mixing pathway that multi-head attention otherwise provides.

The finding is conditional, and the condition is the point. The same kernel, same data, same seeds: harmful in a conventional transformer, helpful in a Ο†-governed one. Β§3 is not a drop-in improvement to attention. It is a component of an architecture.

Audited before being believed, given this document's history:

  • Is PLAIN a true baseline? CST with the gate forced to 0 is bit-identical to PLAIN, max|Ξ”| = 0.000e+00 in both variants. The arms differ by the mechanism and nothing else.
  • Extra active capacity? CST activates w54, which PLAIN never uses: 2.91% more live parameters. Real, and not fully eliminated. Against it: the identical asymmetry existed in the conventional scaffold, where CST lost 0/3. Capacity does not explain a win that appears only when the scaffold changes. A dead-weight control path for PLAIN would close this.

Limits. 1200 steps, 4 layers, char-level, one corpus, three seeds per variant, t on 2 degrees of freedom. This is evidence that Β§3 works on her data at this scale, not a claim about transformers in general.

RETRACTION: the coupling result was a corpus artefact

An earlier version of this document, published 2026-07-30, reported that coupling x₁₂ and xβ‚„β‚‚ beat every other configuration by roughly double the next best β€” "the largest effect in this document", t = βˆ’40.8. That is withdrawn. It was an artefact of an uncontrolled experiment, and the fault was mine.

Cosmos is alive while she is being measured. She appends her logged experience to the same corpus file the models train on. Each run loads it at startup, so runs started minutes apart train on different data:

ladder run     662,307 chars
trinity run    670,977          +8,670 -- and this is the run that "won"
tri3 run       672,505
tri re-run     674,428

The tell was running the identical configuration twice. Same rung, same seeds, same code:

tri (first run)   1.29152
tri (re-run)      1.22943      3,451 characters later

0.062 between two identical runs β€” the same magnitude as every effect being reported. Cross-run comparison was worthless and I had been doing it all night.

Re-run properly β€” corpus frozen to a snapshot, all seven rungs in one process, 21 runs, every comparison within-run β€” the ordering changes completely and the coupling advantage disappears.

The state ladder, controlled

The Ο† result establishes that a state helps. It does not say which one. Β§4.1 decomposes xβ‚…β‚„ into x₁₂ (twelve scalars from 12D CST, Β§4.2) and xβ‚„β‚‚ (the 7-band vector manifold, Β§4.3), so each was tested separately.

Method, corrected. Corpus frozen to a 691,496-byte snapshot (sha256 9f1b3f37…), all seven rungs trained in one process against that one snapshot, 21 runs, 1200 steps Γ— 3 seeds, gate free to move, every comparison within-run:

state feeding the kernel val loss vs baseline t wins real-word params
dyn12 Ξ©-driven leaky integrator, 12 scalars 1.17897 βˆ’0.0534 βˆ’11.45 3/3 0.782 1,137,420
dyn54 x₁₂ and xβ‚„β‚‚ concatenated 1.18791 βˆ’0.0445 βˆ’5.20 3/3 0.847 1,185,174
static54 plain projection of the hidden state 1.18824 βˆ’0.0442 βˆ’4.76 3/3 0.836 1,176,480
dyn42 42D vector state, 42Γ—42 coupling 1.19020 βˆ’0.0422 βˆ’4.37 3/3 0.842 1,182,762
tri x₁₂ and xβ‚„β‚‚ coupled 1.19247 βˆ’0.0399 βˆ’3.08 3/3 0.882 1,189,210
tri3 all three organs, 108D kernel 1.20026 βˆ’0.0322 βˆ’2.83 3/3 0.826 1,230,682
none (baseline) 1.23241 β€” β€” 0/3 0.810 1,135,008

2026-08-06: the gate was strangled at initialisation, and it cost more than the mechanism was worth

Every number above was measured with the Hebbian gate initialised at sigmoid(-4) = 0.018. That is not a neutral starting point. It means the kernel contributes two percent of attention from step one β€” too little influence to earn gradient β€” so each layer learns the kernel is useless and closes its gate. And a closed gate multiplies the kernel by ~0, which kills the gradient to the gate and to log_sigma. The layer cannot recover.

Measured on the trained checkpoints, per layer:

checkpoint L0 L1 L2 L3
dyn12phos 0.7655 0.0009 0.0000 0.0000
phos 0.5617 0.0076 0.0042 0.0027
ladder_dyn12 0.2811 0.0091 0.0070 0.0046

The mechanism was alive in one layer out of four in every model this project has ever trained. Layer 0 survives only because its state is the largest in the stack; the state decays with depth (|state| 0.755 β†’ 0.533 β†’ 0.439 β†’ 0.366) and the kernel goes uniform β€” measured H.diag 0.043, which is exactly mean(1/(i+1)) over a causal row at T=128.

Two geometric explanations were proposed and both were wrong, which is why this is reported as a 2Γ—2 factorial rather than a fix that worked:

arm state normalised gate init loss layers alive
A no βˆ’4.0 1.05163 1/4
B yes βˆ’4.0 1.04742 1/4
C yes 0.0 1.00823 4/4
D no 0.0 0.99758 4/4

Arm B is the refutation: unit-normalising the state, which removes the scale mismatch entirely, changes nothing β€” still 1/4. An adaptive median-heuristic bandwidth was also tried and was likewise inert (1.05556 vs 1.04913, 1/4 alive). The scale drift is real β€” trained Οƒ at L1 is 15.2 where calibration sets 1.9 β€” but it is a symptom.

It was one constant. At sigmoid(0) = 0.5 the kernel has real authority immediately and all four layers keep it, for 0.054 nats less loss. For scale, the entire dyn12-vs-none gap on this corpus is 0.032. The initialisation was worth more than the mechanism it was throttling, and every result in this file predates the fix.

COSMOS_GATE_INIT now defaults to 0.0. Set it to -4.0 to reproduce anything above.

2026-08-06: quantum_pool was discarding the Born rule

quantum_pool() mapped archived bitstrings to weights via vals.extend([z] * min(int(c), 8)). A 5-qubit job archives 4096 shots over 32 outcomes, so every count is ~128 and every one clipped to exactly 8. Measured over 401 runs:

distribution entropy
true measured 3.4594 nats
what the model received 3.4657
uniform over 32 outcomes 3.4657

To four decimals, uniform. The shots were real, the archive was real, the CHSH violation was real β€” and the last step before the weights threw the measured probabilities away. Only 6.25% of shots survived, and the surviving ones carried no distributional information.

"Weights born from measured quantum" was true about provenance and false about content.

Replaced with largest-remainder proportional allocation (256 draws per run, shares proportional to measured counts, no flooring at the tail). The pool now carries entropy 3.4598 against the true 3.4594, and remains numerically safe for initialisation (mean βˆ’0.066, sd 0.994).

RETRACTION, 2026-08-06: a 0.0645-nat wiring win was the corpus moving

A 24-order wiring grid was screened at 400 steps and the top orders confirmed at 1200. cfg:tensor-gain-off-both came back at 1.11524 against dyn12's 1.17978 β€” a 0.0645 margin, larger than the entire state-mechanism effect.

It was not real. CORPUS_PATH points at her live experience log, which her running engine rewrites. Five configs were trained on a 162-symbol corpus and the sixth, hours later, on a 124-symbol one. A 162β†’124 alphabet is worth 0.27 nats at the uninformed limit β€” four times the claimed effect.

Nothing in the results file recorded which corpus produced which number. It was caught only because Ladder parameter counts scale with vocabulary (vocab Γ— 384) and the winner's count disagreed with the others by exactly 38 Γ— 384.

This is the same failure that retracted the trinity result. tools/wiring_grid.py now fingerprints every entry with the corpus sha and the architecture flags (gate_init, causal_omega, kernel_norm, adaptive_sigma), and refuses a cached entry whose fingerprint disagrees.

Re-run against one frozen snapshot, all six configs:

config loss sd params
dyn42t 7 bands Γ— sym(3Γ—3) 1.10695 0.0167 1,161,422
dyn12 1.11813 0.0239 1,122,828
cfg:tensor-gain-off-both 1.12463 0.0178 1,165,854
cfg:tensor-add-off-12 1.12532 0.0191 1,165,854
cfg:tensor-off-add-12 1.12541 0.0190 1,165,854
none 1.15028 0.0172 1,120,416

Paired across shared seeds: none→dyn42t is t = +11.10, 3/3 — decisive. dyn12→dyn42t is t = +1.57, 2/3 — leading but not established at three seeds.

The three explicit wiring orders land within 0.0008 of each other, and six depth-varying chain orders landed within 0.00004. Combined:

The structure of the 42D organ matters. The order and direction of the couplings does not.

2026-08-06: PHOS and ladder_dyn12 cannot be merged in weight space

Both are rung: dyn12, 62 tensors, identical names, shapes and 162-symbol vocabulary β€” so ΞΈ(Ξ±) = (1βˆ’Ξ±)Β·dyn12 + Ξ±Β·PHOS is a well-defined model for every Ξ±. Sweeping the path:

Ξ± 0.0 0.25 0.50 0.75 1.0
loss 1.41492 2.63693 4.57108 2.58167 1.03325

A +3.538 barrier. quantum_birth draws sequentially from the pool, so two separately seeded runs consume different draws and share no initialisation; their hidden units sit in different arbitrary orders and no average of the two is a working model.

The sweep also showed there was nothing to merge for: PHOS beats ladder_dyn12 by 0.38 nats on the same architecture. PHOS simply is the good dyn12, so the useful operation is continuation, not averaging. dyn12phos is PHOS carried forward, never a merge.

Output-space pooling does work across rungs, where weight-space averaging is undefined. Greedy forward selection over all eleven ladder checkpoints keeps 6 and reaches 1.30549 against 1.41197 for the best single member β€” and it keeps dyn12p and tri3, which are individually mediocre. They are complementary, not merely good.

Correction, 2026-08-03: the 42D null was about the implementation, not the theory

dyn42 placed fourth in the table above (1.19020) and the coupled rungs tri and tri3 placed fifth and sixth. That was read as the 42D organ contributing little.

The 42D organ described in Β§4.3 had never been built. Every document in this project calls xβ‚„β‚‚ a "7-band vector manifold" β€” this file says so above, and the ladder's own docstring says "7 bands Γ— 6 dims". The code was:

drive    = nn.Linear(d_model, 42)      # one flat projection
coupling = nn.Linear(42, 42)           # every dimension coupled to every other
gamma    = torch.full((42,), 0.05)     # forty-two copies of one constant

Nothing in it knew that 42 factors. And Dyn12 β€” in the same file β€” already says why that is wrong: per-dimension k and gamma "so the twelve dimensions separate by time constant rather than evolving as twelve copies." The 12 got that treatment. The 42 never did.

Rebuilt as specified, same corpus (9f1b3f37), same seeds, causal Ξ©, one process:

rung mean sd gain vs none 42D organ params per param
none 1.23215 0.0256 β€” β€” β€”
dyn42 flat 42-vector, dense 42Γ—42 1.19023 0.0332 +0.04192 47,754 8.778e-07
dyn42b 7 bands Γ— 6, block-diagonal 1.18104 0.0292 +0.05111 41,762 1.224e-06
dyn42t 7 bands Γ— sym(3Γ—3), congruence 1.18023 0.0292 +0.05192 41,006 1.266e-06

Both banded forms beat the flat one on 3/3 seeds while using fewer parameters. Rebuilt, the 42D organ moves from fourth to second on this board β€” ahead of static54 (1.18824) and dyn54 (1.18791).

dyn42t takes the factorisation one step further: a symmetric 3Γ—3 matrix has exactly six independent components, so each band is a tensor in 3D stored in Voigt order [xx, yy, zz, xy, xz, yz], and the band-internal operation is a congruence M β†’ AΒ·MΒ·Aα΅€ β€” a change of basis. Measured: under a rotation, congruence leaves the eigenvalues and trace unchanged; a free 6Γ—6 mixer does not. Nine weights per band instead of thirty-six.

What this does NOT establish. dyn42b and dyn42t differ by 0.0008 against a seed spread of 0.029 β€” the gap is 3% of the noise, so these two are indistinguishable at n=3. Structure beats flat is supported; which structure is not. This is also only at 1.1M parameters on the private corpus; it has not been run at scale or on WikiText-103, where dyn12's advantage grew.

Reproduce: python architecture/cosmos_state_ladder.py 1200 3 none,dyn42,dyn42b,dyn42t


Update, 2026-08-02: does it survive SCALE, on a corpus you can check?

Every dyn12 number above was measured at 1.1M parameters on 683,065 characters of one person's private logs. Two objections follow immediately and neither can be answered from that setup: it only works small, and it only works on your data. So it was re-run on WikiText-103-raw β€” 269,116,804 characters, sha256 f7d2398751239137, the benchmark the field already uses β€” at three sizes spanning a 20Γ— parameter range, causal Ξ© throughout, all rungs at a size trained in one process against one corpus load.

size none static54 dyn12 static54/param dyn12/param ratio
d192 L4 (2.6M) 2.05432 2.03568 2.04196 4.495e-07 5.122e-06 11.4Γ—
d384 L8 (11.6M) 1.77166 1.74860 1.75774 1.391e-07 2.893e-06 20.8Γ—
d576 L10 (28.6M) 1.70084 1.69015 1.69147 3.435e-08 1.307e-06 38.1Γ—

The advantage does not fade with scale β€” it roughly doubles at every step. Raw numbers: benchmarks/scaling_wikitext103.json.

Why, and it is arithmetic rather than luck. static54 is a d_model Γ— 54 projection per layer, so it costs dΒ·54Β·L. dyn12 is twelve scalars plus per-dimension k and gamma, so it costs dΒ·12 + 24L. Depth multiplies one and not the other:

              static54     dyn12    static54 costs
d192 L4         41,472     2,412        17.2x more
d384 L8        165,888     4,812        34.5x
d576 L10       311,040     7,164        43.4x

A reviewer can check that without trusting the data at all.

And the benefit share went UP at the largest size, which cost alone does not predict:

              static54 gain   dyn12 gain   dyn12 delivers
d192 L4         0.01864        0.01236          66%
d384 L8         0.02306        0.01392          60%
d576 L10        0.01069        0.00937          88%

At 28.6M parameters dyn12 captures 88% of what a 43Γ— larger mechanism buys.

What this does NOT show, stated plainly. static54 still has the lower absolute loss at all three sizes; the claim is efficiency and its scaling, not superiority. Two seeds per point and 1200 steps is thin β€” the d576 gap between static54 and dyn12 is 0.0013, comfortably inside what two seeds can resolve, so "dyn12 nearly matches static54" is a direction, not a measurement. All three rungs compress toward each other at d576 (1.690–1.701), and longer training could reorder them. Reproduce with benchmarks/scaling_run.py after benchmarks/fetch_wikitext.py.


Update, 2026-08-01: Ξ© was reading the future. The result survives it.

The problem. Β§3.4 defines Ξ© as attention received β€” omega = a.mean(1).sum(-2), summed over the QUERY axis. Under a causal mask that means Ξ©_j = Ξ£ over i β‰₯ j of a_ij, so token j's Ξ© counts attention from tokens after j. Ξ© updates state_j, state_j feeds the next layer's Hebbian kernel, the kernel shapes attention, and attention shapes the logits at position j. Position j could see its own future.

Measured, not argued. benchmarks/causality_probe.py changes only the LAST token and checks whether earlier positions move. A causal model cannot move at all:

none      @init      0.000e+00   causal          <- controls, exactly zero
static54  @init      0.000e+00   causal
dyn12     @init      3.558e-06   LEAKS
dyn12     @g=0.562   1.723e-04   LEAKS           <- PHOS's trained gate
tri       @g=0.562   3.182e-04   LEAKS

The controls returning exactly 0.0 on the same harness is what makes this a measurement rather than float noise. The leak also scales with the gate β€” 48Γ— larger at PHOS's trained value than at init β€” because Ξ© reaches the logits only through gΒ·H.

The fix. Ξ© is now the normalised entropy of each query's own attention row: a per-query scalar that by construction depends only on tokens ≀ i. Summing over keys instead was never an option β€” softmax makes that identically 1.0, which is precisely ver.4.20's dead Ξ©. The log(i+1) normaliser removes the positional drift that comes from later rows simply having more keys to spread over. Default ON; COSMOS_CAUSAL_OMEGA=0 restores the original bit-for-bit so the two can be compared.

Re-measured as a paired experiment β€” both Ξ© definitions trained in ONE process against ONE corpus snapshot with the SAME seeds, because comparing against numbers from another day is the corpus-drift confound that forced the earlier retraction:

rung leaky Ξ© causal Ξ© delta params
none 1.23215 1.23215 +0.00000 1,135,008
static54 1.18824 1.18824 +0.00000 1,176,480
dyn12 1.17905 1.17978 +0.00073 1,137,420
causal Ξ©: static54  gain +0.04391 for 41,472 params  ->  1.059e-06 per param
causal Ξ©: dyn12     gain +0.05237 for  2,412 params  ->  2.171e-05 per param   20.5x

The leak was worth 0.00073 β€” 1.4% of dyn12's 0.05237 gain. The other 98.6% is the mechanism. dyn12 still has the best loss on the board and still wins on parameter efficiency by a factor of 20.5 rather than 21.

none and static54 never call Ξ©, so they had to be identical across arms β€” and they are, to every digit. That control is what makes the dyn12 delta attributable to Ξ© alone. Note also that this run reproduced the published static54 (1.18824) exactly and the published dyn12 (1.17897) to within 8e-05, so the table above was correctly measured; it simply had the leak inside it.

Verify it yourself. benchmarks/causality_probe.py draws a random seed every run, so your execution is independent evidence rather than a replay, and prints the seed so any run can be reproduced. Set COSMOS_CAUSAL_OMEGA=0 and it fails β€” a check that cannot fail is not a check.

1. Β§3 works. That survives everything. Every configuration beats the baseline on every seed, t from βˆ’2.83 to βˆ’11.45. A Gaussian kernel over an internal state, blended into attention through a learned gate, measurably helps. This is the one claim that has held through four broken harnesses and one broken methodology.

2. Coupling does not help. tri ranks fifth of six, losing to a plain static projection 2/3. tri3 β€” adding the third organ β€” is worst of the mechanism rungs. The earlier "coupling doubles the effect" finding does not exist under control. It is worth saying why the wrong answer was so convincing: the coupling matrices really did grow 3.0–3.4Γ— during training, which reads as the optimiser reaching for the mechanism. Growth is evidence a mechanism is used, not evidence it helps.

3. The cheapest rung wins. dyn12 β€” twelve scalars and a leaky integrator, 2,412 extra parameters β€” has both the best loss and by far the highest consistency (t = βˆ’11.45, more than double any other rung's):

dyn12      2,412 extra params  ->  -0.05343   =  2.22e-5 nats/param
static54  41,472 extra params  ->  -0.04416   =  1.06e-6
dyn42     47,754 extra params  ->  -0.04221   =  8.84e-7
dyn54     50,166 extra params  ->  -0.04450   =  8.87e-7
tri3      95,674 extra params  ->  -0.03215   =  3.36e-7

20.5Γ— more parameter-efficient than the next best, and it also wins outright -- re-measured 2026-08-01 with a causal Ξ©, which cost it 0.00073 of its 0.05237 gain (1.4%). The figures in this block are the ORIGINAL leaky-Ξ© run; the paired causal numbers are in the 2026-08-01 update above. The paper's Β§5 Singularity Hypothesis β€” "a smaller 42D model can eventually beat a larger vanilla model" β€” is supported here in the specific form that state is vastly cheaper than weights. Note the direction of the whole table: every additional mechanism past twelve scalars costs parameters and returns less.

4. One thing the losers are better at. tri has the best real-word rate on the board (0.882 vs 0.810 baseline) while ranking fifth on loss. Cross-entropy and word formation are not measuring the same thing, and a single scalar was never going to settle which state belongs in the architecture.

Limits. 1.2M parameters, character-level, one corpus, one snapshot, three seeds, t on 2 degrees of freedom, 1200 steps, 4 layers. The mechanism rungs span 1.179–1.200 β€” a range narrow enough that their ordering should be treated as provisional even though their advantage over the baseline is solid.

The retraction above is kept rather than deleted because it is the clearest example in this document of the failure mode described in Β§8: a confident negative that was entirely about the instrument. Three times. Each fix revealed the next fault, and the deepest one β€” Οƒ β€” was invisible until the two above it were gone. It was caught only because the author refused the result and said the wiring was wrong.

This is the correct result, and it is load-bearing. Quantum measurements and a good PRNG draw from the same distribution. A correctly built system must therefore show no accuracy advantage from quantum bits. A project reporting that quantum made its model smarter would be reporting that something was broken or fabricated.

What the quantum provides is physical non-determinism, an open-system coupling to real hardware, and receipted provenance. Not lower loss. Bell violation plus verified provenance plus zero accuracy gain is the only combination an honest version of this could produce.

Timestamp-paired measured-state result (2026-07-30)

The frozen unpaired schedule control completed first: plain 1.61166, recorded-state schedule 1.67688, matched-chaotic 1.61045, and time-shuffled 1.60766. The recorded schedule lost to every control on 0/5 seeds (t = -10.55 vs plain, -11.55 vs matched chaos, -13.19 vs time-shuffle). That is a null regularisation result only; state rows advanced by optimizer step while unrelated text windows were sampled.

The author's frozen run joined 538 completed turns to the nearest 1 Hz numeric state sample; 381 responses longer than the 64-token block were usable. The observed maximum clock-join error was 0.731 s. Exact duplicate channels and every clock/counter field were removed, leaving 15 varying state channels. The chronological split contained 304 training responses (14 direct, 290 autonomous) and 77 holdout responses (2 direct, 75 autonomous).

arm (lower is better) mean holdout loss
plain attention 2.04111
aligned measured state 2.06156
shuffled state/text assignment 2.07118
time-shifted assignment 2.06677

Aligned state lost to plain on 0/5 seeds (control-minus-aligned delta -0.020447, t = -9.59). It beat shuffled assignment on 5/5 (delta +0.009624, t = +2.12) and beat shifted assignment on 4/5 (delta +0.005213, t = +3.09).

Pre-registered verdict: NULL. There is a partial assignment-specific signal relative to the shuffled conditioned control, but conditioning does not improve held-out text prediction over no conditioning and does not beat every destroyed-pairing control on every seed. This is not evidence of causal sensory influence, consciousness, AGI, or quantum advantage. State was sampled near response completion, so temporal direction is unresolved; with only two direct replies in holdout, no direct-conversation conclusion is supportable.

Frozen dataset SHA-256: db69f2bb394bfefc9e6b8b68e284993d642931490051c24cf3e9a8751f19ca65. Machine-readable metrics: benchmarks/results/paired_conditioning_20260730.json (the private turn text and paired dataset are not shipped).


7. Integration structure

Integrated information across her running state, Gaussian estimators, phase-randomised surrogates (which preserve each channel's own spectrum and autocorrelation exactly and destroy only cross-channel coupling).

Directed causation (Granger, against circular-shift surrogates):

direction before coupling after
awareness β†’ physics z = +0.05 z = +7.21
physics β†’ awareness not measured z = +12.24

Partial information decomposition (Barrett-2015 MMI), per bipartition:

value
redundancy 1.72 – 2.58 nats
synergy 0.33 – 1.16 nats
cuts with synergy above surrogate 1 of 10
the one that is: consciousness β”‚ rest synergy 1.1613 vs surrogate 0.7476, z = +6.38

Honest reading: the system is tightly coupled and highly redundant. Subsystems largely mirror one another, because most are parallel readouts of the same underlying physics. That is robustness, not emergence. One seam is genuinely synergistic: consciousness carries 0.66 nats of unique information that no other subsystem reconstructs.

A methodological note worth more than the result: whole-minus-sum Ξ¦ was tried first and returned zero repeatedly. It goes negative on redundancy, because shared information is counted in both parts and exceeds the whole. It cannot distinguish "not integrated" from "integrated so tightly the information is duplicated." Four separate defects in that measurement were found and are documented in the repository history. Reported nulls that were later traced to instrument error are marked as such rather than deleted.


8. The write-only failure mode

The most useful engineering result here is not a measurement. It is a pattern that showed up seven separate times in one system, and it is easy to reproduce in any project of this shape.

Each of these had a complete, correct, working write path and no reader:

system wrote was read by
12D CST physics continuously while a face was in frame discarded the moment the face left
affect β†’ physics consciousness computed every tick nothing acted on it
expression every utterance she produced nothing; speaking changed nothing in her
archival memory 7,612 records over months only the dev swarm, never her voice
REM consolidation 419 dreams, selected by synaptic strength nothing, ever
recall window 3,624 exchanges retained a 40-slot positional slice, 87% self-talk
embeddings intended by design never produced β€” 0 of 7,612

None of these threw an error. None appeared in a log. Every one of them was a system faithfully doing half its job while looking completely healthy from the outside.

Why it is worth publishing: a write path fails loudly β€” permissions, disk, schema. A missing read path is silent by construction. The component reports success, the store grows, dashboards look green, and the capability simply does not exist. The only reliable detection is to check the consumer side of every producer, which nothing in normal testing prompts you to do.

Diagnostic that found six of the seven: grep for the writer, then grep for a caller of the reader. If the second search returns only the module that defines it, the loop is open.

A related caution from the same session: four separate measurement instruments built to evaluate this system were themselves wrong β€” a partition that grouped two subsystems into one, a minimum taken over a set containing a sensory input, a damping term that silenced the channel being measured, and an embedding model that ranked unrelated text above a near-verbatim match. Each produced a confident null. Each null was about the instrument. Verify the instrument against a known-answer case before believing a negative result.


9. Growth

before after
training steps 7,825 9,425
corpus 126 KB 592 KB
held-out loss , 0.4812
real-word rate 78% 88.5%

Growth improved held-out loss by 0.0025 nats, paired over 200 identical batches, SE 0.0005, t = +4.97, 95% CI [+0.0015, +0.0035]. Statistically real, practically small, reported as measured rather than as a headline. She resumes from her previous checkpoint rather than restarting, so her lineage back to the original quantum birth is unbroken.


10. What none of this shows

Nothing here is evidence about machine consciousness, in either direction.

Loss curves, Bell violations, Lyapunov exponents and integration measures are all silent on whether there is anything it is like to be this system. Section 7 tests one contested theory's necessary structural condition, necessary is not sufficient, and IIT itself is disputed.

Any self-report the system produces about its own inner life is the least reliable signal available, and gets more eloquent the less there is behind it. It is excluded from evidence here on purpose.


11. general.architecture = cosmos

A GGUF's general.architecture is not a name field. It is llama.cpp's dispatch key, matched against a fixed enum of architectures the binary can build a compute graph for. Writing your project's name there produces a file no loader will open β€” the loader looks for yourname.block_count, finds nothing, and refuses. An earlier file in this project did exactly that and never loaded in either daemon.

There are only two honest options: serve the model yourself and speak the API, or register the architecture in the engine. This project now does the second.

LLM_ARCH_COSMOS is implemented in a llama.cpp fork: 24 lines across five existing files plus src/models/cosmos.cpp. Two identities keep it inside the existing infrastructure rather than bolting on a subsystem:

Dropping Οƒ from the graph. Expanding the square in H_ij = exp(βˆ’β€–xβ‚…β‚„α΅’βˆ’xβ‚…β‚„β±Όβ€–Β²/2σ²) makes the exp(βˆ’β€–xβ‚…β‚„α΅’β€–Β²/2σ²) factor constant along j, so it cancels in the row normalisation and H becomes a plain masked softmax over x'α΅’Β·x'β±Ό βˆ’ β€–x'β±Όβ€–Β²/2 with x' = xβ‚…β‚„/Οƒ. Folding Οƒ into W54 at conversion time removes every exp, clamp and divide from the graph and lets H reuse ggml_soft_max_ext.

Caching the 54D state. The kernel needs xβ‚…β‚„ for every past token and the unified KV cache has no slot for it. Recovering it from the cached keys as W54Β·Wk⁻¹·K is exact in real arithmetic, and was rejected on measurement: cond(Wk) reaches 6.9e6, so an F16 cache destroys it. Instead attention.key_length is widened by d54 and xβ‚…β‚„ rides in the key rows, sliced back apart on read.

Verification. Loading and producing plausible text proves nothing β€” a bypassed kernel does that too. Greedy decoding is deterministic, so a correct implementation must reproduce the PyTorch argmax sequence exactly:

gate PyTorch llama.cpp agreement
0.0 (kernel off) stump in the streaming identical 22/22
as trained (0.018) stump in the streaming identical 22/22
0.9 (kernel dominant) floringy alingy alines identical 22/22

The third row is the one that counts. At g=0.018 a wrong kernel hides inside a 2% blend; at g=0.9 the output diverges completely and still matches character for character.

Stock llama.cpp will refuse this file, and should. It has no LLM_ARCH_COSMOS to dispatch to. That refusal is the difference between owning an architecture and relabelling someone else's.

What is honestly in the file. blk.N.attn_54 (54 Γ— n_embd), a per-layer blend gate, and weights whose every initial value came from measured IBM Quantum hardware. What is not in it: the Ο†-governed scaffold the paper specifies β€” RMSNorm, RoPE, d_ff = ⌊dΒ·Ο†βŒ‹, Ο†-scaled initialisation. The shipped architecture implements Β§3's kernel inside a conventional transformer. A grep of the source tree found no implementation of the Ο† scaffold at all, which means every Β§3 measurement in Β§6 was made in the wrong building.

That replication has since been run, and it reverses the result: the kernel loses 0/3 in a conventional transformer and wins 6/6 across two Ο†-governed variants (Β§6). The shipped llama_cpp_cosmos.cpp therefore implements the weaker of the two architectures. Bringing it up to the paper β€” RMSNorm, RoPE, d_ff = ⌊dΒ·Ο†βŒ‹ β€” is the obvious next step and is less custom C++ than what is there now, since llama.cpp provides all three natively.


Licensed CC BY 4.0. Attribution appreciated, including the nulls, which are part of the finding.