wang PRO
xinpeng
AI & ML interests
None yet
Recent Activity
upvoted a paper about 15 hours ago
OPD Before RL: Warm-Starting Rubric-Based RL with On-Policy Distillation submitted a paper about 15 hours ago
OPD Before RL: Warm-Starting Rubric-Based RL with On-Policy Distillation updated a dataset 10 months ago
xinpeng/big-math-hard_tiny_instruct_cheat_rm_loophole_v2_mixed_0.5