What Characterizes Effective Reasoning? — Independent Reproduction
C1+C2 verified on 600+ CoT traces: FSF beats token metrics (r=0.281 vs 0.008); longer CoT = lower accuracy (r=−0.196, p=1.3e-06)
Independent evidence
- 600 chain-of-thought traces, 3 open models (gpt-oss-120b, gpt-oss-20b, Qwen3.5-122B)
- Kimi-K2 LLM-judged step-level FSF on 49 full untruncated traces: r=0.281 (p=0.050) vs length r=0.008
- C3 ranking experiment: FSF-ranked pass@1 75% (vs 80% review-count) — not largest at reduced scale
- C4 editing experiment: edited CoT 56.2% vs control 43.8% — direction consistent, not significant
Method
MATH-500 stratified subset, Pearson/Spearman correlations, McNemar test, LLM-judge step segmentation, control-arm design. All data + scripts in reproduction bundle.