Buckets:
Source audit
- Paper: arXiv 2605.30447.
- OpenReview: STcIzNrUBB.
- Official implementation and result CSVs: Advueu963/Calibrated_Preference_Learning at
dc0b53324cfc98625517f6e7cdeb5c65a77b89a6. - RewardBench2 score files: allenai/reward-bench-2-results at
b19c7033e964187d12e74a43a07f2d727a3d37e5.
The current paper source labels the rankwise-but-neither counterexample as Table 2 and the top-1-not-full-rank counterexample as Table 3. The second registered claim is correct in substance but points to Table 2 for a result demonstrated by Table 3.
The empirical audit does not retrain ranking models. It independently recomputes RewardBench2 top-1 ECE and correlations from all available immutable score files, and separately aggregates the official five-fold label-ranking result CSVs.
Xet Storage Details
- Size:
- 1.1 kB
- Xet hash:
- 225298d36a1a988daa578fd8ddc8d526fba829b4ed244cf46f06c0553ae0ca8a
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.