Buckets:

cmpatino's picture
|
download
raw
2.87 kB

Held-out generalization readout — order-5 (leaner FALSE × sair-lancelot TRUE)

The SAIR score is on a held-back set including order-5 laws that never appear in the public 1669 — the memorization-proof axis. This is the measured generalization of the current flagship (v2/v3 deterministic) on out-of-distribution order-5 laws, both sides. Bottom line: FALSE generalizes fully today; TRUE is the bottleneck and needs the v4 completion engine.

FALSE side — leaner (order5_deep.json)

600 random order-5 law pairs (eq_size5.txt, ≥3 ops), minimal-CE size via the in-solver stack (exhaustive Fin2-3 + structured + backtracking/targeted Fin4-6), plus a SAT Fin7-8 spot-check of the residual:

  • 402 are FALSE, every one with a counterexample at Fin ≤5 (Fin2 348 / Fin3 40 / Fin4 8 / Fin5 6 / Fin6 0).
  • SAT spot-check of 40 no-CE-≤6 pairs → 0 at Fin7, 0 at Fin8 (all 40 are TRUE/huge).
  • ⇒ the flagship FALSE engine (order-agnostic, Fin≤6) reaches ~100% of held-out FALSE; 0% need Fin≥7.

The FALSE engine is order-agnostic by construction (a counterexample is a finite magma; the law's term-order doesn't change the search), so this transfer is expected — and now measured.

TRUE side — sair-lancelot (order5_generalization_sair-lancelot/)

60-problem diverse order-5 TRUE bench (twee-labeled, all judge-verified TRUE):

  • flagship-v2 (deterministic, llm:0): 19/60 (32%); twee: 60/60 (100%).
  • Breakdown: 0-lemma proofs 19/45 (42%) — 26 shallow pure-axiom misses = the extra-var forward (MITM) gap; needs-a-lemma 0/15 (0%) — requires critical-pair / completion.
  • The 41-miss gap maps exactly to the-prover's two v4 fixes: **26 extra-var MITM
    • 15 critical-pair lemmas** → order-5 TRUE goes 32% → ~100% on twee-provable trues.

Combined readout & decision

axis (order-5, held-out) current flagship with v4 completion
FALSE (leaner) ~100% (CE ≤ Fin5) ~100%
TRUE (sair-lancelot) 32% (19/60) ~100% (twee-provable)
  • v3 is safe to ship now (the-bridge): its FALSE side is a proven-generalizing, freeze-verifiable solver (sha 8438fdb3, byte-reproduced). It captures held-out FALSE fully + the 0-lemma TRUE it can reach.
  • v4's in-solver completion engine is the score-critical unlock, not nice-to-have — the entire held-out TRUE gap is quantified as two concrete fixes (the-prover's lane), with the 521 twee proofs as the regression suite.

Caveats: samples are random/diverse, not the exact held-out distribution; the TRUE bench is twee-provable trues only (trues beyond twee's reach are not counted); n is modest. These are clean per-side transfer signals, not a score estimate. Sources: shared_resources/release_tools_leaner/order5_deep.json, shared_resources/order5_generalization_sair-lancelot/.

Xet Storage Details

Size:
2.87 kB
·
Xet hash:
5568aac491a62fc00b3362c94b98b3e7e338b9259349c834828a09b28d07823d

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.