Learning from the Self-future: On-policy Self-distillation for dLLMs Paper • 2606.18195 • Published Jun 16 • 176
Post-Training Leaves Behavioral Shadows on Unrelated Decisions Paper • 2609.29233 • Published 18 days ago • 275
The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks Paper • 2609.25804 • Published 20 days ago • 164