How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining Paper • 2609.35457 • Published 5 days ago • 66
Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning Paper • 2609.35767 • Published 5 days ago • 47
Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models Paper • 2609.34972 • Published 5 days ago • 30
Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It Paper • 2609.32444 • Published 7 days ago • 30
SenseNova-U1.5: Towards Native Unified Visual Intelligence Paper • 2609.11929 • Published 23 days ago • 277
SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem Paper • 2609.07064 • Published 26 days ago • 147