arxiv:2608.14277
๐ In a Training Loop
haodi lei
bingyang-lei
AI & ML interests
None yet
Recent Activity
upvoted a paper 11 days ago
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening upvoted a paper 24 days ago
Rethinking On-Policy Distillation of Large Language Models II: One Training Example updated a collection 24 days ago
Draft-OPD