HydraHead: From Head-Level Functional Heterogeneity to Specialized Attention Hybridization Paper • 2606.20097 • Published Jun 18 • 18
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop Paper • 2605.18746 • Published May 18 • 6
Nemotron-Post-Training-v3 Collection Collection of datasets used in the post-training phase of Nemotron Nano, Super, and Ultra v3. • 50 items • Updated 17 days ago • 186
Innovator-VL Collection A Multimodal Large Language Model for Scientific Discovery • 9 items • Updated Mar 5 • 6
SSR Collection Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning • 5 items • Updated Mar 2 • 2
Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs Paper • 2512.16584 • Published Dec 18, 2025 • 2
R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning Paper • 2508.21113 • Published Aug 28, 2025 • 111
VisionThink Collection Efficient Reasoning Vision Language Model • 7 items • Updated Jul 18, 2025 • 7
view article Article Vision Language Models (Better, faster, stronger) +3 merve, sergiopaniego, ariG23498, pcuenq, andito • May 12, 2025 • 614
MiniCPM-o & MiniCPM-V Collection Multimodal models with leading performance. • 32 items • Updated May 24 • 85
We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning? Paper • 2407.01284 • Published Jul 1, 2024 • 81
Boosting Multimodal Reasoning with MCTS-Automated Structured Thinking Paper • 2502.02339 • Published Feb 4, 2025 • 23