How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining Paper • 2609.35457 • Published 2 days ago • 55
Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs Paper • 2608.20492 • Published Aug 20 • 87
Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs Paper • 2608.12781 • Published Aug 17 • 35