| [ |
| "Capability boundaries are estimated via smoothed pinball-loss quantile regression at quantile τ=0.98 with smoothing parameter κ=50 and regularization λ=10⁻³ (Section 2.1).", |
| "Post-training capability boundaries are modeled as sigmoid functions of log10(FLOPs), q_τ^sig(z;θ)=y0+L·σ(a+βz), achieving in-distribution pinball loss of 4.08×10⁻³ versus 4.00×10⁻³ for a more flexible I-spline estimator, while generalizing better out-of-distribution (4.93×10⁻³ vs 4.92×10⁻³) (Section 3.1.1, Table 2).", |
| "At 10²⁴ FLOPs, the estimated 0.98-quantile boundary reaches 0.539 accuracy on MATH Level 5 versus 0.563 on MMLU-Pro, 0.700 on BBH, 0.828 on IFEval, 0.535 on MUSR, and 0.424 on GPQA (Table 1).", |
| "Unlike the stable boundaries for BBH, GPQA, MMLU-Pro, and MUSR, the MATH Level 5 and IFEval capability boundaries are non-stationary over time, with later-period models exceeding earlier predicted boundaries (Figure 2).", |
| "A balanced I-optimal sampling design recovers near-full-data capability frontiers using roughly 20% of the evaluation budget, with GPQA and MUSR requiring as little as 5% (Section 4, Figure 5).", |
| "A cross-benchmark shift test applied to AIME-2025 scores finds a positive but not statistically significant shift (p=0.15), giving no clear aggregate evidence of post-release contamination inflating math benchmark scores (Section 5.2.2, Equation 3)." |
| ] |
|
|