RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and Embodiments Paper • 2610.10409 • Published 1 day ago • 28
RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and Embodiments Paper • 2610.10409 • Published 1 day ago • 28
Collaborative Personalized Preference Alignment for LLMs under Data Deficiency Paper • 2610.05898 • Published 4 days ago • 7
Collaborative Personalized Preference Alignment for LLMs under Data Deficiency Paper • 2610.05898 • Published 4 days ago • 7
Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence Paper • 2608.31075 • Published Aug 31 • 30
Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence Paper • 2608.31075 • Published Aug 31 • 30
Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence Paper • 2608.31075 • Published Aug 31 • 30
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents Paper • 2607.23588 • Published Jul 26 • 129
Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment Paper • 2605.20834 • Published May 20 • 4
Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment Paper • 2605.20834 • Published May 20 • 4
Shaping Schema via Language Representation as the Next Frontier for LLM Intelligence Expanding Paper • 2605.09271 • Published May 10 • 7
Shaping Schema via Language Representation as the Next Frontier for LLM Intelligence Expanding Paper • 2605.09271 • Published May 10 • 7
Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows Paper • 2604.28139 • Published Apr 30 • 42