arxiv:2410.01623
xichen
xichen-fy
ยท
AI & ML interests
LLMs
Recent Activity
upvoted a paper about 18 hours ago
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning upvoted a paper 2 months ago
Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation upvoted a paper over 1 year ago
Video-R1: Reinforcing Video Reasoning in MLLMsOrganizations
None yet