| <meta charset="utf-8"><style> | |
| body{margin:0;background:#07111f;color:#edf6ff;font:18px system-ui,sans-serif}main{padding:28px;max-width:1100px;margin:auto}h1{font-size:34px;margin:0 0 12px}h2{color:#77d9ff}.hero{background:#10233b;border-left:8px solid #48c6a8;padding:20px;font-size:24px;font-weight:700}li{margin:12px 0}code{color:#ffd479}.foot{font-size:14px;color:#afc4d8;margin-top:24px}</style><main><h1>What Characterizes Effective Reasoning? — Independent Reproduction</h1><div class="hero">C1+C2 verified on 600+ CoT traces: FSF beats token metrics (r=0.281 vs 0.008); longer CoT = lower accuracy (r=−0.196, p=1.3e-06)</div><h2>Independent evidence</h2><ul><li>600 chain-of-thought traces, 3 open models (gpt-oss-120b, gpt-oss-20b, Qwen3.5-122B)</li><li>Kimi-K2 LLM-judged step-level FSF on 49 full untruncated traces: r=0.281 (p=0.050) vs length r=0.008</li><li>C3 ranking experiment: FSF-ranked pass@1 75% (vs 80% review-count) — not largest at reduced scale</li><li>C4 editing experiment: edited CoT 56.2% vs control 43.8% — direction consistent, not significant</li></ul><h2>Method</h2><p>MATH-500 stratified subset, Pearson/Spearman correlations, McNemar test, LLM-judge step segmentation, control-arm design. All data + scripts in reproduction bundle.</p><div class="foot">ICML 2026 Open Reproductions · Alogotron · HF-router-only compute · generated 2026-07-29T20:20:00+00:00</div></main> | |