shuo shen
hyperion-shuo
ยท
AI & ML interests
reinforcement learning
Recent Activity
upvoted a paper about 11 hours ago
Bellman Policy Optimization liked a model 2 days ago
internlm/Intern-S2-397B upvoted a paper 14 days ago
OPV: Outcome-based Process Verifier for Efficient Long Chain-of-Thought Verification