Buckets:
| {"charts":{"coverage":[{"distribution_count":50,"lower":0.71,"mean":0.752,"observation_count":500,"sample_size":10,"upper":0.792},{"distribution_count":50,"lower":0.762,"mean":0.7979999999999999,"observation_count":500,"sample_size":20,"upper":0.8320000000000001},{"distribution_count":50,"lower":0.8079999999999999,"mean":0.838,"observation_count":500,"sample_size":30,"upper":0.8660000000000001},{"distribution_count":50,"lower":0.784,"mean":0.8220000000000001,"observation_count":500,"sample_size":40,"upper":0.858},{"distribution_count":50,"lower":0.836,"mean":0.868,"observation_count":500,"sample_size":50,"upper":0.898},{"distribution_count":50,"lower":0.82,"mean":0.852,"observation_count":500,"sample_size":60,"upper":0.882},{"distribution_count":50,"lower":0.7979999999999999,"mean":0.828,"observation_count":500,"sample_size":70,"upper":0.858},{"distribution_count":50,"lower":0.812,"mean":0.8420000000000001,"observation_count":500,"sample_size":80,"upper":0.872},{"distribution_count":50,"lower":0.8079999999999999,"mean":0.8420000000000001,"observation_count":500,"sample_size":90,"upper":0.8759999999999999},{"distribution_count":50,"lower":0.7959999999999999,"mean":0.8340000000000001,"observation_count":500,"sample_size":100,"upper":0.87}],"regression":[{"distribution_count":10,"lower":0.023130405886121093,"mean":0.04031158725906639,"observation_count":100,"sample_size":10,"upper":0.0575285694400882},{"distribution_count":10,"lower":0.035821314908235084,"mean":0.060473756276566114,"observation_count":100,"sample_size":20,"upper":0.08572473298803536},{"distribution_count":10,"lower":0.04502722885844905,"mean":0.07319949266678219,"observation_count":100,"sample_size":30,"upper":0.10275283323003626},{"distribution_count":10,"lower":0.05563431089837605,"mean":0.0930437678913241,"observation_count":100,"sample_size":40,"upper":0.13041228398812152},{"distribution_count":10,"lower":0.06178815897357138,"mean":0.10159636947906532,"observation_count":100,"sample_size":50,"upper":0.14096126966354738}]},"claims":{"A1":{"answer":"Both main bilevel pipelines executed, but capped or worsening-stop tasks prevent a complete end-to-end verification.","downgrade":"Capped or worsening-stop trajectories do not establish that Algorithm 2 finished learning the reported ambiguity geometry.","facts":[["Validated tasks",14000],["Accepted solver fraction",1.0],["Required tasks capped",1954]],"official_text":"The paper formulates learning the ambiguity set for OT-DRO as a bilevel problem in which the upper level tunes the ambiguity-set geometry parameter theta and the lower level solves a standard OT-DRO problem (Section 4.2).","title":"Did the complete bilevel pipeline run?","verdict":"partially_verified"},"A2":{"answer":"Finite trajectories are supporting evidence only, and no valid theorem receipt passed.","downgrade":"Finite capped trajectories cannot verify Theorem 5.1.","facts":[["Audit assumptions",12],["Executed horizon","finite and capped"],["Theorem verdict","inconclusive"]],"official_text":"Theorem 5.1 establishes that the proposed hypergradient descent procedure converges to a critical point of the bilevel problem under mild conditions, including square-summable step sizes (Section 5.1, Theorem 5.1).","title":"Does Theorem 5.1 follow as stated?","verdict":"inconclusive"},"A3":{"answer":"The required independently bound finite-difference receipt is missing or failed.","downgrade":"The pinned component receipt checks implemented gradients, but it does not bind Algorithm 1 at nonsmooth active-set boundaries.","facts":[["Pinned component routes",5],["Receipt scope","bounded components"],["Paper-scale receipt","not bound"]],"official_text":"Algorithm 1 computes hypergradients through the nonsmooth conservative implicit function theorem, allowing differentiation through conic (OT-DRO) programs without assuming the solution map is continuously differentiable (Section 5, Algorithm 1).","title":"Did we validate the nonsmooth hypergradient route?","verdict":"inconclusive"},"A4":{"answer":"The anchored k=2 and J=30 setup and Figure 2 relative-improvement formula are not bound by the frozen plan or a hash-bound source audit. A4 therefore remains inconclusive. The available slope analysis is reported only as a sensitivity extension.","downgrade":"The available sample-size sweep is a sensitivity extension, not the anchored Figure 2 estimand.","facts":[["Matched fields",3],["Required fields",5],["Figure 2 formula","unbound"]],"official_text":"In portfolio optimization experiments with k=2, J=30, n_b=20, gamma=0.05, beta=0.1, the relative improvement of the learned ambiguity set over a fixed Wasserstein ball increases as sample size decreases (Section 6.1, Figure 2).","title":"Does the exact Figure 2 trend reproduce?","verdict":"inconclusive"},"A5":{"answer":"Coverage is evaluated at every portfolio sample size against the predeclared 0.9 target using distribution-first confidence intervals.","downgrade":"Every distribution-first upper confidence bound is below the predeclared 0.90 target, but capped tasks block a terminal falsified verdict.","facts":[["Target",0.9],["Sample sizes",10],["All lower bounds at target",false]],"official_text":"Figure 3 shows the learned ambiguity set still enforces the coverage constraint, containing the true data-generating distribution with the specified high probability despite being shaped to reduce loss (Section 6.1, Figure 3).","title":"Did the learned set preserve 90% coverage?","verdict":"inconclusive"},"A6":{"answer":"Verification requires positive out-of-sample improvement at every sample size and a positive distribution-level mean for each of the ten independent trials.","downgrade":"The distribution-level mean is positive at each sample size, but the claim says consistently across all ten trials and that stricter test fails.","facts":[["Independent trials",10],["Sample sizes",5],["Every trial mean positive",false]],"official_text":"Across 10 independent linear regression trials, the learned decision-focused ambiguity set consistently yields less conservative (lower average loss) decisions than baseline OT-DRO ambiguity sets (Section 6.2, Figure 5).","title":"Was regression loss lower in all ten trials?","verdict":"partially_verified"}},"derived_payload_sha256":"sha256:70cd4c79a2b19793d1cd6895e8fc96c3089616d6231096d2d1d2f48e7ec73a26","limits":[{"cap":5000,"censored_task_count":4579,"kind":"iteration_cap","statement":"A task at 5,000 iterations is right-censored, not converged. Its terminal decision and metric are a capped trajectory value."},{"kind":"literal_stopping_rule","statement":"The paper's literal signed denominator may stop when the total penalized objective worsens. Such tasks are not labeled convergence.","worsening_stop_count":0},{"kind":"theorem_scope","statement":"Finite capped trajectories cannot verify Theorem 5.1."}],"matrix":{"censored_at_5000":4579,"recovered_rows":5,"rejected_rows":0,"solver_acceptance_fraction":1.0,"validated_rows":14000},"paper":{"openreview_id":"K1EPPO9t2c","submission":"12512","title":"Loss-Aware Distributionally Robust Optimization via Trainable Optimal Transport Ambiguity Sets"},"schema_version":4,"seal":{"aggregate_identity":"sha256:4b4b1ef2383fecf4baa1565e7f9fbf2bc1b300b0b55dc67ba934f2bac5cf771a","aggregate_results_sha256":"025f51a63fd75de100390b13374942cc5b57400d900b36c544777f9e67e0bd78","analysis_payload_sha256":"sha256:1ffc132991f8d970011d9c0aa49da3e04cc78d2b0363bab634cedd32714280fd","hypergradient_receipt_sha256":"85ba8545c76a780fdd68934686db1ce616e2a2162aa8c59e4860c9c802a28a43","manifest_hash":"sha256:ac1d3c6cb4b87b3b9deb430352aadc7a0bc07acb421c95f6b6858bee4a67be8d","official_claims_lock_sha256":"8c852634725aa2d678d87ec3517eca882a095af8770fcebe971547910943a046","recovery_attestation_hash":"sha256:99c791f43a6423bf4e5f9c5710b92d1cfe85994e33758405ef2f0b15d22a9fbf","source_analysis_file_sha256":"2c872ba10ec1c5d166a15155eb0516e89abdbfea8aaae1a4126c6bf1895074bd","theorem_audit_sha256":"5555a358cbe1040cd39c6637e2b57d9012318a0a9ae5d2a4a696e216466ae011"}} | |
Xet Storage Details
- Size:
- 8.06 kB
- Xet hash:
- 7ca6c4d625284b4142afeb7af13557c3227629a1e539f2d15e9246117cb89254
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.