| { |
| "schema_version": 1, |
| "title": "Repro: Nonparametric LLM Evaluation from Preference Data", |
| "emoji": "microscope", |
| "space_id": "snaykey/repro-nonparametric-llm-eval", |
| "paper": {"openreview_id": "rHndxbqWyh"}, |
| "tags": ["icml2026-repro", "paper-rHndxbqWyh"], |
| "updated_at": "2026-07-27T00:00:00+00:00", |
| "root": { |
| "slug": "index", |
| "title": "Repro: Nonparametric LLM Evaluation from Preference Data", |
| "file": "pages/index.md", |
| "children": [ |
| {"slug":"executive-summary","title":"Executive summary","file":"pages/executive-summary/page.md","children":[]}, |
| {"slug":"claim-1","title":"DMLRank defines Generalized Average Ranking Scores (GARS) as θ=E[F(μ(X))], a unifying functional that recovers Borda scores, Bradley-Terry log-odds projections, and Rank Centrality stationary-distribution scores as special cases (Section 4).","file":"pages/claim-1/page.md","children":[]}, |
| {"slug":"claim-2","title":"Theorem 5.1 derives a debiased efficient-influence-function estimator for GARS that is asymptotically normal, √n(θ̂_EIF−θ)→N_d(0,Σ), and achieves the semiparametric efficiency bound under cross-fitted nuisance estimation (Theorem 5.1).","file":"pages/claim-2/page.md","children":[]}, |
| {"slug":"claim-3","title":"On synthetic data with n=1000, the debiased Borda estimator has error 0.15±0.03 with 0.94±0.06 coverage, versus the plugin estimator's 0.38±0.08 error and only 0.17±0.09 coverage, showing plugin confidence intervals are invalid while debiased ones are not (Table 1).","file":"pages/claim-3/page.md","children":[]}, |
| {"slug":"claim-4","title":"Theorem 6.2 characterizes the A-optimal preference-labeling policy under a budget constraint as π*_jk(x)=clip_[α,1]√(tr(J_jkV_jkJ_jkᵀ)/(λ_A c_jk)), balancing informativeness, influence on the target functional, and per-comparison cost (Theorem 6.2).","file":"pages/claim-4/page.md","children":[]}, |
| {"slug":"claim-5","title":"Under a fixed budget of β=2000 comparisons, the A-optimal labeling policy achieves lower MSE than random data collection for all three GARS types tested: Borda (0.108±0.041 vs 0.133±0.060), Bradley-Terry (2.565±1.046 vs 2.910±1.143), and Rank Centrality (0.010±0.008 vs 0.017±0.007) (Table 2).","file":"pages/claim-5/page.md","children":[]}, |
| {"slug":"claim-6","title":"The framework is applied to real Chatbot Arena preference data with n=32,980 prompts and K=20 models, using toxicity score and TF-IDF prompt embeddings as covariates, where plugin estimators yield near-zero-width (invalid) confidence intervals while debiased estimators retain interpretable uncertainty (Section 7, Figure 3).","file":"pages/claim-6/page.md","children":[]}, |
| {"slug":"conclusion","title":"Conclusion","file":"pages/conclusion/page.md","children":[]} |
| ] |
| } |
| } |
|
|