Buckets:
| { | |
| "openreview_id": "OT9cxeWbEO", | |
| "arxiv": "2602.11557", | |
| "claims": [ | |
| { | |
| "claim": 1, | |
| "source_status": "confirmed", | |
| "source_page": 7, | |
| "anchor": "Section 4.1, Theorem 4.1", | |
| "note": "The theorem states rho=gamma-4(n/b-1)R>0, equivalently b>4Rn/(gamma+4R), and convergence to margin rho. The discussion further observes this integer condition actually collapses to full batch b=n under the paper's assumptions." | |
| }, | |
| { | |
| "claim": 2, | |
| "source_status": "misstated", | |
| "source_page": 8, | |
| "anchor": "Section 4.2, Theorem 4.4 and Remark 4.5", | |
| "note": "The batch-momentum mechanism is supported, but the source labels it Theorem 4.4, not 4.3. It requires the positive effective-margin condition rho=gamma-2(1-beta_1)m(m^2-1)R>0; beta_1 approaching one permits small batches and closes the approximation gap but is not by itself the complete theorem condition." | |
| }, | |
| { | |
| "claim": 3, | |
| "source_status": "misstated", | |
| "source_page": 9, | |
| "anchor": "Section 4.2, Corollary 4.6", | |
| "note": "The paper attributes the explicit rates to Corollary 4.6, not 4.5. It states multiplicative slowdown factors scaling with m/(1-beta_1), but the full rates also contain learning-rate-regime-dependent threshold terms." | |
| }, | |
| { | |
| "claim": 4, | |
| "source_status": "misstated", | |
| "source_page": 10, | |
| "anchor": "Section 4.3, Theorem 4.7 and Corollary 4.8", | |
| "note": "Exact recovery of the full-batch max-margin bias for any batch size is supported, but the source numbers are 4.7 and 4.8, not 4.6 and 4.7. The theorem's m^(2+a) factor is for the no-momentum transient; with momentum it is divided by 1-beta_1, and Corollary 4.8 gives learning-rate-specific powers." | |
| }, | |
| { | |
| "claim": 5, | |
| "source_status": "misstated", | |
| "source_page": 11, | |
| "anchor": "Section 4.4, Definition 4.9 and Theorem 4.10", | |
| "note": "The orthogonal scale-skewed construction and distinct per-sample limiting bias are present, but the result is Theorem 4.10; 4.9 is the definition of the bias directions. The registered theorem citation is wrong." | |
| }, | |
| { | |
| "claim": 6, | |
| "source_status": "confirmed", | |
| "source_page": 12, | |
| "anchor": "Section 5, Figure 1", | |
| "note": "The source uses n=200 synthetic multiclass data and sweeps b in {20,200}, beta in {0,0.5,0.99}; Figure 1 reports normalized SGD/momentum and variance-reduced variants and says large momentum or variance reduction restores full-batch-like convergence for small batches." | |
| } | |
| ] | |
| } |
Xet Storage Details
- Size:
- 2.57 kB
- Xet hash:
- 13182feb296be4adb00510548ef869c7ba0a5dcd0518acef8ccb3f77ab488ce9
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.