OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Paper β’ 2607.28609 β’ Published 23 days ago β’ 73
Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone Paper β’ 2512.22615 β’ Published Dec 27, 2025 β’ 51
LongVideoAgent: Multi-Agent Reasoning with Long Videos Paper β’ 2512.20618 β’ Published Dec 23, 2025 β’ 56
ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration Paper β’ 2511.21689 β’ Published Nov 26, 2025 β’ 129
JanusCoder: Towards a Foundational Visual-Programmatic Interface for Code Intelligence Paper β’ 2510.23538 β’ Published Oct 27, 2025 β’ 99
CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training Paper β’ 2504.13161 β’ Published Apr 17, 2025 β’ 98
Jailbreaking as a Reward Misspecification Problem Paper β’ 2406.14393 β’ Published Jun 20, 2024 β’ 14
Running on CPU Upgrade 14.1k Open LLM Leaderboard π 14.1k Track, rank and evaluate open LLMs and chatbots