Context experiment — n=5 run (2026-07-19)
Target paper: On Language Generation in the Limit with Bounded Memory (arXiv 2605.30324)
Reference paper: Kleinberg & Mullainathan, Language Generation in the Limit (arXiv 2404.06757)
Model: gpt-5.4 all roles · Scale: {200, 300, 400 words} × {without PDF, with PDF} × 5 trials (30 reconstructions)
Frozen summaries measured: 200→c200-f, 300→c300-c, 400→c400-d (same seeds as the earlier n=3 run).
How we scored
Per-trial pass/fail is brittle: judging each reconstruction alone against a fixed bar, an LLM judge applies that bar inconsistently, so near-identical work flips pass↔fail on noise. So the authoritative result uses a single judge that ranked all 30 reconstructions together, blinded to which had the PDF. We report the rank as a quality score from 1 to 30, higher is better (best of the 30 scores 30).
Result — single-judge ranking (higher = better)
| Summary length | Without PDF | With PDF | Effect of the PDF |
|---|---|---|---|
| 200 words | 9.2 | 21.2 | helps a lot |
| 300 words | 8.8 | 10.2 | helps slightly |
| 400 words | 23.6 | 20.0 | hurts slightly |
Per-trial scores (of 30):
| Cell | trial scores |
|---|---|
| 400 without | 30, 27, 26, 25, 10 |
| 400 with | 29, 28, 17, 15, 11 |
| 200 with | 24, 23, 21, 20, 18 |
| 300 with | 19, 16, 13, 2, 1 |
| 200 without | 22, 9, 8, 4, 3 |
| 300 without | 14, 12, 7, 6, 5 |
What this run says
- 400 words is the strongest budget, with or without the PDF. The four best reconstructions of all 30 are all 400-word trials (both conditions). More words reliably produce better reconstructions — this is the run's most robust signal.
- The PDF helps at short budgets and is redundant at long ones. With PDF beats without at 200 (21.2 vs 9.2, a large gap) and 300 (10.2 vs 8.8), but slightly trails at 400 (20.0 vs 23.6). Coherent reading: the reference paper supplies scaffolding a thin summary lacks, but adds nothing once the summary is already rich.
- Pass/fail would have told the opposite story. The per-trial pass counts (in the appendix) put "with PDF" at or below "without PDF" in every cell — the reverse of the ranking. The two scorers disagree on the sign of the context effect; the joint ranking is the one to trust, and it favors the PDF at short budgets.
The catch — this contradicts the earlier n=3 run
In the earlier n=3 study, 200-with-PDF was the worst cell; here at n=5 it is one of the best (score 21.2). Same prompt, same condition, opposite result. The between-run swing is large enough that the short-budget context effect is not yet safe to bank: one run says the PDF hurts at 200, the other says it helps a lot there. What does reproduce across both runs is that 400 words is reliably the best budget.
Why the scores move so much (mechanism)
Every reconstruction recovers the main positive theorem and the arbitrary-text characterization. The judge's separators are all in the delicate boundary results — chiefly the element-generator impossibility proof (needs an exhaustive case split), plus using invalid constant texts for the arbitrary-text necessity direction, and stating the index-generator result as exact identification rather than eventual subset containment. Whether an attempt nails these is close to a coin-flip, which is why cells span a wide range and why run-to-run results swing.
Bottom line
At n=5, the authoritative ranking says: more words help (400 is best), and the reference paper helps when the summary is short but not when it is long. But this short-budget benefit reverses versus the earlier run, so it is a suggestive signal on one paper, not a settled effect. The only claim stable across both runs is that 400 words is the reliable budget.