Han Cui
hancui
AI & ML interests
None yet
Recent Activity
upvoted a paper about 14 hours ago
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening upvoted a paper about 1 month ago
SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning upvoted a paper 4 months ago
Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified ScalingOrganizations
None yet