SabaPivot's picture
|
download
raw
1.1 kB
# Source audit
- Paper: [arXiv 2605.30447](https://arxiv.org/abs/2605.30447).
- OpenReview: [STcIzNrUBB](https://openreview.net/forum?id=STcIzNrUBB).
- Official implementation and result CSVs: [Advueu963/Calibrated_Preference_Learning at `dc0b53324cfc98625517f6e7cdeb5c65a77b89a6`](https://github.com/Advueu963/Calibrated_Preference_Learning/tree/dc0b53324cfc98625517f6e7cdeb5c65a77b89a6).
- RewardBench2 score files: [allenai/reward-bench-2-results at `b19c7033e964187d12e74a43a07f2d727a3d37e5`](https://huggingface.co/datasets/allenai/reward-bench-2-results/tree/b19c7033e964187d12e74a43a07f2d727a3d37e5).
The current paper source labels the rankwise-but-neither counterexample as Table 2 and the top-1-not-full-rank counterexample as Table 3. The second registered claim is correct in substance but points to Table 2 for a result demonstrated by Table 3.
The empirical audit does not retrain ranking models. It independently recomputes RewardBench2 top-1 ECE and correlations from all available immutable score files, and separately aggregates the official five-fold label-ranking result CSVs.

Xet Storage Details

Size:
1.1 kB
·
Xet hash:
225298d36a1a988daa578fd8ddc8d526fba829b4ed244cf46f06c0553ae0ca8a

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.