arxiv:2609.20511
Shangzhe Li
DVA13304
AI & ML interests
Reinforcement Learning, Imitation Learning, Learning Theory
Recent Activity
authored a paper 1 day ago
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation upvoted a paper 2 days ago
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation upvoted a paper 7 months ago
Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement LearningOrganizations
None yet