How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining Paper • 2609.35457 • Published 4 days ago • 65
Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning Paper • 2609.35767 • Published 4 days ago • 47
Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models Paper • 2609.34972 • Published 4 days ago • 29
Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It Paper • 2609.32444 • Published 6 days ago • 28
SenseNova-U1.5: Towards Native Unified Visual Intelligence Paper • 2609.11929 • Published 22 days ago • 277
SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem Paper • 2609.07064 • Published 25 days ago • 147