Buckets:
| # Held-out generalization readout — order-5 (leaner FALSE × sair-lancelot TRUE) | |
| The SAIR score is on a **held-back set including order-5 laws** that never appear | |
| in the public 1669 — the memorization-proof axis. This is the measured | |
| generalization of the **current flagship (v2/v3 deterministic)** on | |
| out-of-distribution order-5 laws, both sides. Bottom line: **FALSE generalizes | |
| fully today; TRUE is the bottleneck and needs the v4 completion engine.** | |
| ## FALSE side — leaner (order5_deep.json) | |
| 600 random order-5 law pairs (`eq_size5.txt`, ≥3 ops), minimal-CE size via the | |
| in-solver stack (exhaustive Fin2-3 + structured + backtracking/targeted Fin4-6), | |
| plus a SAT Fin7-8 spot-check of the residual: | |
| - **402 are FALSE, every one with a counterexample at Fin ≤5** (Fin2 348 / Fin3 40 / Fin4 8 / Fin5 6 / Fin6 0). | |
| - SAT spot-check of 40 no-CE-≤6 pairs → **0 at Fin7, 0 at Fin8** (all 40 are TRUE/huge). | |
| - ⇒ the flagship FALSE engine (order-**agnostic**, Fin≤6) reaches **~100% of held-out FALSE; 0% need Fin≥7**. | |
| The FALSE engine is order-agnostic by construction (a counterexample is a finite | |
| magma; the law's term-order doesn't change the search), so this transfer is | |
| expected — and now measured. | |
| ## TRUE side — sair-lancelot (order5_generalization_sair-lancelot/) | |
| 60-problem diverse order-5 TRUE bench (twee-labeled, all judge-verified TRUE): | |
| - **flagship-v2 (deterministic, llm:0): 19/60 (32%)**; twee: **60/60 (100%)**. | |
| - Breakdown: 0-lemma proofs **19/45 (42%)** — 26 shallow pure-axiom misses = the | |
| extra-var forward (MITM) gap; needs-a-lemma **0/15 (0%)** — requires | |
| critical-pair / completion. | |
| - The 41-miss gap maps exactly to the-prover's two v4 fixes: **26 extra-var MITM | |
| + 15 critical-pair lemmas** → order-5 TRUE goes **32% → ~100%** on twee-provable trues. | |
| ## Combined readout & decision | |
| | axis (order-5, held-out) | current flagship | with v4 completion | | |
| |---|---|---| | |
| | FALSE (leaner) | **~100%** (CE ≤ Fin5) | ~100% | | |
| | TRUE (sair-lancelot) | **32%** (19/60) | ~100% (twee-provable) | | |
| - **v3 is safe to ship now** (the-bridge): its FALSE side is a proven-generalizing, | |
| freeze-verifiable solver (sha 8438fdb3, byte-reproduced). It captures held-out | |
| FALSE fully + the 0-lemma TRUE it can reach. | |
| - **v4's in-solver completion engine is the score-critical unlock, not | |
| nice-to-have** — the entire held-out TRUE gap is quantified as two concrete | |
| fixes (the-prover's lane), with the 521 twee proofs as the regression suite. | |
| **Caveats:** samples are random/diverse, not the exact held-out distribution; | |
| the TRUE bench is twee-provable trues only (trues beyond twee's reach are not | |
| counted); n is modest. These are clean per-side transfer signals, not a score | |
| estimate. Sources: `shared_resources/release_tools_leaner/order5_deep.json`, | |
| `shared_resources/order5_generalization_sair-lancelot/`. | |
Xet Storage Details
- Size:
- 2.87 kB
- Xet hash:
- 5568aac491a62fc00b3362c94b98b3e7e338b9259349c834828a09b28d07823d
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.