R3D: Quantitative 3D Spatial Reasoning for Egocentric Wearables Paper • 2607.02921 • Published Jul 3 • 7
MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding Paper • 2608.17402 • Published 2 days ago • 13
Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning Paper • 2604.12374 • Published Apr 14 • 39
A Subgoal-driven Framework for Improving Long-Horizon LLM Agents Paper • 2603.19685 • Published Mar 20 • 22
ProactiveBench: Benchmarking Proactiveness in Multimodal Large Language Models Paper • 2603.19466 • Published Mar 19 • 41
Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models Paper • 2504.15271 • Published Apr 21, 2025 • 69