What Characterizes Effective Reasoning? — Independent Reproduction

C1+C2 verified on 600+ CoT traces: FSF beats token metrics (r=0.281 vs 0.008); longer CoT = lower accuracy (r=−0.196, p=1.3e-06)

Independent evidence

Method

MATH-500 stratified subset, Pearson/Spearman correlations, McNemar test, LLM-judge step segmentation, control-arm design. All data + scripts in reproduction bundle.

ICML 2026 Open Reproductions · Alogotron · HF-router-only compute · generated 2026-07-29T20:20:00+00:00