BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender Paper • 2609.15478 • Published 13 days ago • 34
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Paper • 2608.05139 • Published Aug 5 • 28
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Paper • 2608.05000 • Published Aug 6 • 64
Scaling Properties of Text Conditioning in Visual Generation Paper • 2607.29679 • Published Jul 31 • 42
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published Jul 10 • 83
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Paper • 2607.07508 • Published Jul 8 • 33
DataComp-VLM: Improved Open Datasets for Vision-Language Models Paper • 2606.28551 • Published Jun 26 • 52