Add RewardBench 2 result: 25.3 average (random floor) with full protocol - the OOD boundary measured on a third-party benchmark 17ee54a verified DantheMan124 commited on 14 days ago
Fix tokenizer_config compatibility: drop list-form extra_special_tokens (crashes transformers 4.x; tokens remain declared in their standard fields) 4f31c0c verified DantheMan124 commited on 14 days ago
Add reward_train_max_length: 512 so the calibration regime ships with the weights 00a4b81 verified DantheMan124 commited on Jul 28
Fix usage snippet: pairwise comparison, not absolute score; add chance floor + public baseline rows 6b00257 verified DantheMan124 commited on Jul 28