SabaPivot's picture
download
raw
3.6 kB
{
"openreview_id": "MrIDZjIsNF",
"arxiv": "2602.01381",
"claims": [
{
"claim_index": 1,
"registered_claim": "Under Assumption 3.2 (uniform Bellman error bound ε), if ε = O(1/T), SMC-based inference-time scaling attains a target TV error with particle/time complexity that is polynomial rather than exponential in the horizon T (Section 5, Theorem 5.1, Corollary 5.2).",
"source_status": "confirmed",
"source_page": "4; 6",
"anchor": "Assumption 3.2; Theorem 5.1; Corollary 5.2",
"note": "The uniform Bellman-error assumption and particle bound imply polynomial complexity when ε scales as O(1/T); Corollary 5.2 states the corresponding naive-proposal SMC runtime.",
"claim": 1
},
{
"claim_index": 2,
"registered_claim": "Without reward guidance, the number of samples needed to hit the target region grows as Ω(L^(2T/3)), an exponential lower bound in T (Section 4, Theorem 4.1).",
"source_status": "misstated",
"source_page": 5,
"anchor": "Theorem 4.1 (LB1)",
"note": "The Ω(L^(2T/3)) lower bound is correct, but the theorem's algorithm is given and may query V-hat satisfying the ratio-bound Assumption 3.1. Thus 'without reward guidance' is not the theorem's literal scope; the paper's contribution summary more narrowly calls it 'without intermediate guidance.'",
"claim": 2
},
{
"claim_index": 3,
"registered_claim": "Even with a Bellman-error-bounded reward model, sampling complexity is lower-bounded by Ω((1+ε)^(2T/3)), showing guidance alone cannot remove exponential dependence unless ε shrinks with T (Section 4, Corollary 4.2).",
"source_status": "confirmed",
"source_page": 5,
"anchor": "Corollary 4.2 (LB2)",
"note": "Under Assumptions 3.1 and 3.2, the corollary gives Ω((1+ε)^(2T/3)); a fixed positive ε therefore retains exponential dependence.",
"claim": 3
},
{
"claim_index": 4,
"registered_claim": "For single-particle guided SMC, the total-variation error is bounded by 2Tε, so guidance fails to control error once ε ≥ 1/(2T) (Section 4, Theorem 4.3).",
"source_status": "ambiguous",
"source_page": 5,
"anchor": "Theorem 4.3 (SP-gSMC TV error) and following discussion",
"note": "The theorem proves the upper bound ||π-tilde_t−π-hat_t||_TV ≤ 2tε. The text says this guarantee becomes non-informative around ε≥1/(2T), but an upper bound becoming vacuous does not prove that the algorithm's actual error is uncontrolled or that guidance necessarily fails.",
"claim": 4
},
{
"claim_index": 5,
"registered_claim": "Theorem 5.1 establishes a particle complexity bound N ≥ L^6 T(1+ε)^(6(T-1))/(2δ_TV) for SMC to achieve TV error δ_TV (Section 5, Theorem 5.1).",
"source_status": "confirmed",
"source_page": 6,
"anchor": "Theorem 5.1 (Particles Complexity)",
"note": "The theorem displays the registered sufficient particle threshold for target total-variation error.",
"claim": 5
},
{
"claim_index": 6,
"registered_claim": "A resampling-pool Metropolis-Hastings chain-based approach achieves the target accuracy with time complexity Õ(L T^3 log(1/δ) log(1/δ_TV)) (Section 6, Theorem 6.1).",
"source_status": "confirmed",
"source_page": 7,
"anchor": "Algorithm 2; Theorem 6.1",
"note": "The resampling-pool MH construction has the stated soft-O runtime on the theorem's good event, which holds with probability at least 1−δ.",
"claim": 6
}
]
}

Xet Storage Details

Size:
3.6 kB
·
Xet hash:
64b57924e198fb1095095db52663da4267b2ece6ae79a401914eb2550efbe0eb

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.