When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation Paper • 2609.20511 • Published 5 days ago • 99
LARK: Learnability-Grounded Trajectory Selection for Efficient Reasoning Distillation Paper • 2605.30651 • Published May 28
From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning Paper • 2606.18089 • Published Jul 5
Efficient Agentic Reasoning Through Self-Regulated Simulative Planning Paper • 2605.22138 • Published May 21 • 10
IsoCompute Playbook: Optimally Scaling Sampling Compute for LLM RL Paper • 2603.12151 • Published Mar 12 • 2
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation Paper • 2609.20511 • Published 5 days ago • 99