squaredcuber's picture
|
download
raw
2.83 kB

Official claim mapping

The scored claim set is pinned in configs/official_claims_lock.json. The live judge merges {**legacy, **anchored}, so the anchored six-claim record replaces the historical three-claim record for OpenReview K1EPPO9t2c. The C1, C2, and C3 labels inside the frozen 14,000-task plan are internal evidence-route labels only. They are not judge-facing claims and changing them would invalidate the active run identity.

Claim Official anchored claim Primary evidence Conservative verdict rule
A1 The ambiguity-set geometry is learned as a bilevel OT-DRO problem. Validated end-to-end portfolio and absolute-regression main suites, raw upper/lower-level lineage, solver feasibility, and per-iteration geometry. Verify only when both main suites are complete with no capped or worsening-stop tasks.
A2 Theorem 5.1 gives convergence to a critical point under its stated conditions. Enumerated source-anchored assumptions, a hash-bound proof artifact, and a separately bound independent-reviewer receipt. Verification is hard-disabled until that complete contract exists. Finite trajectories and self-declared receipts never verify the theorem.
A3 Algorithm 1 differentiates through nonsmooth conic OT-DRO programs. Source/formula audit plus nonsmooth and active-set boundary evidence for conservative implicit differentiation without continuous-differentiability assumptions. The current six smooth finite-difference checks can support at most partial verification. They cannot verify A3.
A4 Portfolio relative improvement increases as sample size decreases. Exact k=2, J=30, n_b=20, gamma=0.05, beta=0.1 binding and the paper's exact Figure 2 relative-improvement estimand. The frozen plan does not bind k, J, or the paper formula, so A4 remains inconclusive. Available slope analysis is extension-only sensitivity evidence.
A5 The shaped ambiguity set retains the specified high-probability coverage. portfolio_gaussian_main coverage at every sample size, aggregated by distribution before bootstrapping. Verify only when every lower 95% confidence bound reaches the predeclared 1 - beta = 0.9 target.
A6 Ten independent regression trials consistently reduce average loss versus baseline OT-DRO. regression_absolute_main out-of-sample relative improvement at every sample size and for each of the ten distributions. Verify only when every sample-size lower confidence bound is positive and every one of the ten distribution means is positive.

The high-dimensional, discrete, Gaussian-mixture, squared-loss, and coverage-penalty-ablation suites remain valuable robustness analyses. They are reported under extensions and cannot create additional scored claims.

Xet Storage Details

Size:
2.83 kB
·
Xet hash:
89e39978b978b339ffd51f6dce7723cbe44d92bb3b08c4b632b9666bc75e5181

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.