Spaces:
Running
Running
Ctrl+K
- 00-scored-evidence-summary
- claim-1-bnrm-replaces-deterministic-scalar-reward-outputs
- claim-2-bnrm-extends-the-bradley-terry-preference-model-wi
- claim-3-with-40k-training-examples-bt-bnrm-improves-over-t
- claim-4-bnrm-matches-the-performance-of-a-baseline-trained
- claim-5-bnrm-reduces-the-pearson-correlation-between-rewar
- claim-6-in-rlhf-policy-evaluation-policies-trained-with-bn
- conclusion
- reproduction-protocol-and-provenance
- 1.16 kB