UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement Paper • 2609.38721 • Published 8 days ago • 293
Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs Paper • 2608.20492 • Published Aug 20 • 87
Alaya-EVOKE: From Linear-Scaling Supervision to Endless World Paper • 2608.13546 • Published Aug 13 • 143