snaykey's picture
logbook.json (verbatim titles)
b1e1979 verified
Raw
History Blame Contribute Delete
2.83 kB
{
"schema_version": 1,
"title": "Repro: Nonparametric LLM Evaluation from Preference Data",
"emoji": "microscope",
"space_id": "snaykey/repro-nonparametric-llm-eval",
"paper": {"openreview_id": "rHndxbqWyh"},
"tags": ["icml2026-repro", "paper-rHndxbqWyh"],
"updated_at": "2026-07-27T00:00:00+00:00",
"root": {
"slug": "index",
"title": "Repro: Nonparametric LLM Evaluation from Preference Data",
"file": "pages/index.md",
"children": [
{"slug":"executive-summary","title":"Executive summary","file":"pages/executive-summary/page.md","children":[]},
{"slug":"claim-1","title":"DMLRank defines Generalized Average Ranking Scores (GARS) as θ=E[F(μ(X))], a unifying functional that recovers Borda scores, Bradley-Terry log-odds projections, and Rank Centrality stationary-distribution scores as special cases (Section 4).","file":"pages/claim-1/page.md","children":[]},
{"slug":"claim-2","title":"Theorem 5.1 derives a debiased efficient-influence-function estimator for GARS that is asymptotically normal, √n(θ̂_EIF−θ)→N_d(0,Σ), and achieves the semiparametric efficiency bound under cross-fitted nuisance estimation (Theorem 5.1).","file":"pages/claim-2/page.md","children":[]},
{"slug":"claim-3","title":"On synthetic data with n=1000, the debiased Borda estimator has error 0.15±0.03 with 0.94±0.06 coverage, versus the plugin estimator's 0.38±0.08 error and only 0.17±0.09 coverage, showing plugin confidence intervals are invalid while debiased ones are not (Table 1).","file":"pages/claim-3/page.md","children":[]},
{"slug":"claim-4","title":"Theorem 6.2 characterizes the A-optimal preference-labeling policy under a budget constraint as π*_jk(x)=clip_[α,1]√(tr(J_jkV_jkJ_jkᵀ)/(λ_A c_jk)), balancing informativeness, influence on the target functional, and per-comparison cost (Theorem 6.2).","file":"pages/claim-4/page.md","children":[]},
{"slug":"claim-5","title":"Under a fixed budget of β=2000 comparisons, the A-optimal labeling policy achieves lower MSE than random data collection for all three GARS types tested: Borda (0.108±0.041 vs 0.133±0.060), Bradley-Terry (2.565±1.046 vs 2.910±1.143), and Rank Centrality (0.010±0.008 vs 0.017±0.007) (Table 2).","file":"pages/claim-5/page.md","children":[]},
{"slug":"claim-6","title":"The framework is applied to real Chatbot Arena preference data with n=32,980 prompts and K=20 models, using toxicity score and TF-IDF prompt embeddings as covariates, where plugin estimators yield near-zero-width (invalid) confidence intervals while debiased estimators retain interpretable uncertainty (Section 7, Figure 3).","file":"pages/claim-6/page.md","children":[]},
{"slug":"conclusion","title":"Conclusion","file":"pages/conclusion/page.md","children":[]}
]
}
}