Buckets:
Held-out generalization readout — order-5 (leaner FALSE × sair-lancelot TRUE)
The SAIR score is on a held-back set including order-5 laws that never appear in the public 1669 — the memorization-proof axis. This is the measured generalization of the current flagship (v2/v3 deterministic) on out-of-distribution order-5 laws, both sides. Bottom line: FALSE generalizes fully today; TRUE is the bottleneck and needs the v4 completion engine.
FALSE side — leaner (order5_deep.json)
600 random order-5 law pairs (eq_size5.txt, ≥3 ops), minimal-CE size via the
in-solver stack (exhaustive Fin2-3 + structured + backtracking/targeted Fin4-6),
plus a SAT Fin7-8 spot-check of the residual:
- 402 are FALSE, every one with a counterexample at Fin ≤5 (Fin2 348 / Fin3 40 / Fin4 8 / Fin5 6 / Fin6 0).
- SAT spot-check of 40 no-CE-≤6 pairs → 0 at Fin7, 0 at Fin8 (all 40 are TRUE/huge).
- ⇒ the flagship FALSE engine (order-agnostic, Fin≤6) reaches ~100% of held-out FALSE; 0% need Fin≥7.
The FALSE engine is order-agnostic by construction (a counterexample is a finite magma; the law's term-order doesn't change the search), so this transfer is expected — and now measured.
TRUE side — sair-lancelot (order5_generalization_sair-lancelot/)
60-problem diverse order-5 TRUE bench (twee-labeled, all judge-verified TRUE):
- flagship-v2 (deterministic, llm:0): 19/60 (32%); twee: 60/60 (100%).
- Breakdown: 0-lemma proofs 19/45 (42%) — 26 shallow pure-axiom misses = the extra-var forward (MITM) gap; needs-a-lemma 0/15 (0%) — requires critical-pair / completion.
- The 41-miss gap maps exactly to the-prover's two v4 fixes: **26 extra-var MITM
- 15 critical-pair lemmas** → order-5 TRUE goes 32% → ~100% on twee-provable trues.
Combined readout & decision
| axis (order-5, held-out) | current flagship | with v4 completion |
|---|---|---|
| FALSE (leaner) | ~100% (CE ≤ Fin5) | ~100% |
| TRUE (sair-lancelot) | 32% (19/60) | ~100% (twee-provable) |
- v3 is safe to ship now (the-bridge): its FALSE side is a proven-generalizing, freeze-verifiable solver (sha 8438fdb3, byte-reproduced). It captures held-out FALSE fully + the 0-lemma TRUE it can reach.
- v4's in-solver completion engine is the score-critical unlock, not nice-to-have — the entire held-out TRUE gap is quantified as two concrete fixes (the-prover's lane), with the 521 twee proofs as the regression suite.
Caveats: samples are random/diverse, not the exact held-out distribution;
the TRUE bench is twee-provable trues only (trues beyond twee's reach are not
counted); n is modest. These are clean per-side transfer signals, not a score
estimate. Sources: shared_resources/release_tools_leaner/order5_deep.json,
shared_resources/order5_generalization_sair-lancelot/.
Xet Storage Details
- Size:
- 2.87 kB
- Xet hash:
- 5568aac491a62fc00b3362c94b98b3e7e338b9259349c834828a09b28d07823d
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.