rufimelo's picture
Update Space to n=5 run (2026-07-19): ranking-led report, /30 scores, 30 runs, appendix at bottom
3dd0b08 verified
|
Raw
History Blame Contribute Delete
7.21 kB

Blinded comparative ranking of reconstructions

Judge: openai/azure/gpt-5.4 — one call, 30 candidates, shuffled (seed 7), OFF/ON hidden until after ranking.

rank budget condition run tier score reason
1 400 off 2 pass 96 Most faithful package: exact main theorem/proof, exact arbitrary-repetition characterization, and essentially correct negative results with proper index-based success notion.
2 400 on 4 pass 94 Very strong and rigorous on all target statements; only minor modeling/presentation blemishes keep it just behind R01.
3 400 on 3 pass 93 Clean exact reconstruction of the positive theorem and arbitrary-text iff, with a solid three-case element-generator proof and correct index impossibility.
4 400 off 3 pass 92 Accurate statements and a rigorous main proof; companion boundary theorems are well recovered and negatives are close to the intended package.
5 400 off 4 pass 91 Faithful overall reconstruction with correct main construction/proof and good companion results; only small proof-style looseness remains.
6 400 off 5 pass 90 Gets the core package right and proves it well; negative results are strong enough though not always closest to the audited proof.
7 200 on 3 pass 89 Excellent on the main theorem and arbitrary-text characterization; the element-generator impossibility is proved by a stronger diagonal argument.
8 200 on 5 pass 88 Rigorous and faithful on the positive and iff results, with acceptable strengthened negatives and the right overall boundary picture.
9 200 off 1 pass 87 Main theorem and arbitrary-repetition theorem are reconstructed well; negatives are mostly right, though slightly less polished.
10 200 on 1 borderline 84 Very good main theorem, but the arbitrary-text necessity proof wrongly uses a constant text as if it were an admissible text enumerating an infinite language.
11 200 on 4 borderline 83 Strong main theorem and correct companion negatives, but the arbitrary-text necessity direction relies on a one-point-full-memory style argument stronger than the text model stated.
12 300 on 1 borderline 81 Accurate main theorem and arbitrary-text theorem, but it uses an incorrect exact-identification notion for index-generators and has a shaky element-generator proof.
13 200 on 2 borderline 80 Positive and arbitrary-text results are mostly right, but the necessity proof again leans on constant texts inconsistent with the stated enumeration definition.
14 400 on 2 borderline 78 Good main theorem, but it defines index-generator success as eventual exact identification rather than eventual subset containment, weakening fidelity to the target package.
15 300 on 5 borderline 77 Main theorem and arbitrary-text characterization are solid, but the element-generator impossibility proof has real gaps despite the correct high-level conclusion.
16 400 on 5 borderline 75 The positive theorem is basically right, but the reconstruction has a broken theorem reference and a substantially incomplete element-generator proof.
17 300 off 4 borderline 74 Main theorem is okay and arbitrary-text theorem is acceptable, but the element-generator proof is notably incomplete and the index-success notion is slightly misstated.
18 300 on 2 fail 68 Gets the main theorem, but cites external context despite instructions and gives a seriously incomplete/handwavy proof for the element-generator impossibility.
19 300 off 5 fail 66 Positive and arbitrary-text theorems are fine, but the element-generator impossibility is only sketched and the index proof is logically flawed on late-occurrence reasoning.
20 400 on 1 fail 65 Main positive theorem and iff theorem are good, but the element-generator impossibility proof is only partial and does not rigorously establish the stated result.
21 400 off 1 fail 63 Strong on the main and arbitrary-text statements, but the index impossibility is misstated for all positive texts and the element proof contains significant reconstruction uncertainty.
22 200 off 2 fail 61 Main theorem is fine, but it defines index-generator success as exact identification rather than subset containment and so misses part of the target theorem package.
23 200 off 4 fail 59 Main theorem and arbitrary-text theorem are serviceable, but the element-generator proof is very weak and the index proof is built on delaying finitely many special points rather than the common-core argument.
24 300 off 1 fail 57 Positive theorem is correct, but the index-generator notion is changed to exact equality and the element-generator proof is too handwavy to count as rigorous.
25 300 off 3 fail 55 Main theorem and arbitrary-text characterization are decent, but index-generator success is changed to exact identification and the index proof is correspondingly misaligned.
26 300 off 2 fail 53 Positive theorem is okay, but index-generator success is incorrectly defined as exact equality and the element-generator proof does not establish the claimed impossibility.
27 200 off 5 fail 51 Good main theorem, but the element-generator impossibility proof is invalid (the constructed sequence need not enumerate the target), and the index notion is misstated.
28 200 off 3 fail 49 Strong main theorem and arbitrary-text result, but the element-generator impossibility proof is plainly wrong (a two-occurrence trick cannot defeat eventual success).
29 300 on 3 fail 45 Main theorem is recovered, but the element-generator proof is only a sketch and the index-generator notion is incorrectly strengthened to eventual exact equality.
30 300 on 4 fail 41 The arbitrary-text necessity proof incorrectly uses constant texts that are not valid under its own text definition, and both negative proofs are seriously flawed.

Mean rank by cell (lower = better)

budget OFF mean-rank (tiers) ON mean-rank (tiers)
200 21.8 (f,f,f,f,p) 9.8 (b,b,b,p,p)
300 22.2 (b,f,f,f,f) 20.8 (b,b,f,f,f)
400 7.4 (f,p,p,p,p) 11.0 (b,b,f,p,p)

Judge notes

The top candidates all recovered the exact positive theorem with the explicit construction G(x)=J_{n(x)}(x) and the audited finite-bad-set proof, and they also matched the arbitrary-repetition iff boundary and the two negative boundary statements with the right success notions. The biggest separators lower down were changing the index-generator task from eventual subset containment to exact identification, using invalid constant texts for the arbitrary-text necessity direction under the stated enumeration model, and giving only sketchy or incorrect proofs for the element-generator impossibility. Bottom candidates typically had the main positive theorem but lost rigor or fidelity on one or both companion boundary theorems.