SabaPivot's picture
download
raw
2.57 kB
{
"openreview_id": "OT9cxeWbEO",
"arxiv": "2602.11557",
"claims": [
{
"claim": 1,
"source_status": "confirmed",
"source_page": 7,
"anchor": "Section 4.1, Theorem 4.1",
"note": "The theorem states rho=gamma-4(n/b-1)R>0, equivalently b>4Rn/(gamma+4R), and convergence to margin rho. The discussion further observes this integer condition actually collapses to full batch b=n under the paper's assumptions."
},
{
"claim": 2,
"source_status": "misstated",
"source_page": 8,
"anchor": "Section 4.2, Theorem 4.4 and Remark 4.5",
"note": "The batch-momentum mechanism is supported, but the source labels it Theorem 4.4, not 4.3. It requires the positive effective-margin condition rho=gamma-2(1-beta_1)m(m^2-1)R>0; beta_1 approaching one permits small batches and closes the approximation gap but is not by itself the complete theorem condition."
},
{
"claim": 3,
"source_status": "misstated",
"source_page": 9,
"anchor": "Section 4.2, Corollary 4.6",
"note": "The paper attributes the explicit rates to Corollary 4.6, not 4.5. It states multiplicative slowdown factors scaling with m/(1-beta_1), but the full rates also contain learning-rate-regime-dependent threshold terms."
},
{
"claim": 4,
"source_status": "misstated",
"source_page": 10,
"anchor": "Section 4.3, Theorem 4.7 and Corollary 4.8",
"note": "Exact recovery of the full-batch max-margin bias for any batch size is supported, but the source numbers are 4.7 and 4.8, not 4.6 and 4.7. The theorem's m^(2+a) factor is for the no-momentum transient; with momentum it is divided by 1-beta_1, and Corollary 4.8 gives learning-rate-specific powers."
},
{
"claim": 5,
"source_status": "misstated",
"source_page": 11,
"anchor": "Section 4.4, Definition 4.9 and Theorem 4.10",
"note": "The orthogonal scale-skewed construction and distinct per-sample limiting bias are present, but the result is Theorem 4.10; 4.9 is the definition of the bias directions. The registered theorem citation is wrong."
},
{
"claim": 6,
"source_status": "confirmed",
"source_page": 12,
"anchor": "Section 5, Figure 1",
"note": "The source uses n=200 synthetic multiclass data and sweeps b in {20,200}, beta in {0,0.5,0.99}; Figure 1 reports normalized SGD/momentum and variance-reduced variants and says large momentum or variance reduction restores full-batch-like convergence for small batches."
}
]
}

Xet Storage Details

Size:
2.57 kB
·
Xet hash:
13182feb296be4adb00510548ef869c7ba0a5dcd0518acef8ccb3f77ab488ce9

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.