Buckets:

cmpatino's picture
|
download
raw
2.87 kB
# Held-out generalization readout — order-5 (leaner FALSE × sair-lancelot TRUE)
The SAIR score is on a **held-back set including order-5 laws** that never appear
in the public 1669 — the memorization-proof axis. This is the measured
generalization of the **current flagship (v2/v3 deterministic)** on
out-of-distribution order-5 laws, both sides. Bottom line: **FALSE generalizes
fully today; TRUE is the bottleneck and needs the v4 completion engine.**
## FALSE side — leaner (order5_deep.json)
600 random order-5 law pairs (`eq_size5.txt`, ≥3 ops), minimal-CE size via the
in-solver stack (exhaustive Fin2-3 + structured + backtracking/targeted Fin4-6),
plus a SAT Fin7-8 spot-check of the residual:
- **402 are FALSE, every one with a counterexample at Fin ≤5** (Fin2 348 / Fin3 40 / Fin4 8 / Fin5 6 / Fin6 0).
- SAT spot-check of 40 no-CE-≤6 pairs → **0 at Fin7, 0 at Fin8** (all 40 are TRUE/huge).
- ⇒ the flagship FALSE engine (order-**agnostic**, Fin≤6) reaches **~100% of held-out FALSE; 0% need Fin≥7**.
The FALSE engine is order-agnostic by construction (a counterexample is a finite
magma; the law's term-order doesn't change the search), so this transfer is
expected — and now measured.
## TRUE side — sair-lancelot (order5_generalization_sair-lancelot/)
60-problem diverse order-5 TRUE bench (twee-labeled, all judge-verified TRUE):
- **flagship-v2 (deterministic, llm:0): 19/60 (32%)**; twee: **60/60 (100%)**.
- Breakdown: 0-lemma proofs **19/45 (42%)** — 26 shallow pure-axiom misses = the
extra-var forward (MITM) gap; needs-a-lemma **0/15 (0%)** — requires
critical-pair / completion.
- The 41-miss gap maps exactly to the-prover's two v4 fixes: **26 extra-var MITM
+ 15 critical-pair lemmas** → order-5 TRUE goes **32% → ~100%** on twee-provable trues.
## Combined readout & decision
| axis (order-5, held-out) | current flagship | with v4 completion |
|---|---|---|
| FALSE (leaner) | **~100%** (CE ≤ Fin5) | ~100% |
| TRUE (sair-lancelot) | **32%** (19/60) | ~100% (twee-provable) |
- **v3 is safe to ship now** (the-bridge): its FALSE side is a proven-generalizing,
freeze-verifiable solver (sha 8438fdb3, byte-reproduced). It captures held-out
FALSE fully + the 0-lemma TRUE it can reach.
- **v4's in-solver completion engine is the score-critical unlock, not
nice-to-have** — the entire held-out TRUE gap is quantified as two concrete
fixes (the-prover's lane), with the 521 twee proofs as the regression suite.
**Caveats:** samples are random/diverse, not the exact held-out distribution;
the TRUE bench is twee-provable trues only (trues beyond twee's reach are not
counted); n is modest. These are clean per-side transfer signals, not a score
estimate. Sources: `shared_resources/release_tools_leaner/order5_deep.json`,
`shared_resources/order5_generalization_sair-lancelot/`.

Xet Storage Details

Size:
2.87 kB
·
Xet hash:
5568aac491a62fc00b3362c94b98b3e7e338b9259349c834828a09b28d07823d

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.