Synthesizing Instruction-Tuning Datasets with Contrastive Decoding Paper • 2604.13538 • Published Apr 15
On the Optimal Reasoning Length for RL-Trained Language Models Paper • 2602.09591 • Published Feb 10 • 6