Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening Paper • 2609.18708 • Published 13 days ago • 82
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 26 days ago • 103
SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning Paper • 2608.14277 • Published Aug 14 • 36
SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning Paper • 2608.14277 • Published Aug 14 • 36
Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning Paper • 2608.14290 • Published Aug 14 • 35
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Paper • 2608.08160 • Published Aug 8 • 30