OPD Before RL: Warm-Starting Rubric-Based RL with On-Policy Distillation Paper • 2610.02781 • Published 7 days ago • 8
Temperature-scaling surprisal estimates improve fit to human reading times -- but does it do so for the "right reasons"? Paper • 2311.09325 • Published Nov 15, 2023
FocalPO: Enhancing Preference Optimizing by Focusing on Correct Preference Rankings Paper • 2501.06645 • Published Jan 11, 2025