GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Paper • 2608.05747 • Published 1 day ago • 28
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Paper • 2608.05000 • Published 3 days ago • 50
Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts Paper • 2608.00574 • Published 7 days ago • 6