Buckets:
| # Source audit | |
| - Paper: [arXiv 2605.30447](https://arxiv.org/abs/2605.30447). | |
| - OpenReview: [STcIzNrUBB](https://openreview.net/forum?id=STcIzNrUBB). | |
| - Official implementation and result CSVs: [Advueu963/Calibrated_Preference_Learning at `dc0b53324cfc98625517f6e7cdeb5c65a77b89a6`](https://github.com/Advueu963/Calibrated_Preference_Learning/tree/dc0b53324cfc98625517f6e7cdeb5c65a77b89a6). | |
| - RewardBench2 score files: [allenai/reward-bench-2-results at `b19c7033e964187d12e74a43a07f2d727a3d37e5`](https://huggingface.co/datasets/allenai/reward-bench-2-results/tree/b19c7033e964187d12e74a43a07f2d727a3d37e5). | |
| The current paper source labels the rankwise-but-neither counterexample as Table 2 and the top-1-not-full-rank counterexample as Table 3. The second registered claim is correct in substance but points to Table 2 for a result demonstrated by Table 3. | |
| The empirical audit does not retrain ranking models. It independently recomputes RewardBench2 top-1 ECE and correlations from all available immutable score files, and separately aggregates the official five-fold label-ranking result CSVs. | |
Xet Storage Details
- Size:
- 1.1 kB
- Xet hash:
- 225298d36a1a988daa578fd8ddc8d526fba829b4ed244cf46f06c0553ae0ca8a
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.