Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 2 days ago • 65
StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability? Paper • 2609.00787 • Published 4 days ago • 1
Weak-to-Strong Generalization via Direct On-Policy Distillation Paper • 2607.05394 • Published Jul 8 • 150
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe Paper • 2604.13016 • Published Apr 14 • 116