HuatuoGPT-3: RL-Only Domain Adaptation from Base Models Paper • 2610.05966 • Published 3 days ago • 21
Spatial Memory Intelligence: Endowing World Models with Understanding-Driven Long-Term Memory Paper • 2610.02521 • Published 7 days ago • 53
COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization Paper • 2609.11682 • Published 28 days ago • 46
SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models Paper • 2609.02886 • Published Sep 2 • 118
SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation Paper • 2608.04419 • Published Aug 5 • 30
Debating Truth: Debate-driven Claim Verification with Multiple Large Language Model Agents Paper • 2507.19090 • Published Jul 25, 2025
Novel Claim or Déjà Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking Paper • 2607.23514 • Published Jul 26 • 14
Novel Claim or Déjà Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking Paper • 2607.23514 • Published Jul 26 • 14
Provably Efficient Exploration in Inverse Constrained Reinforcement Learning Paper • 2409.15963 • Published May 16, 2025
Real-Time Verification of Embodied Reasoning for Generative Skill Acquisition Paper • 2505.11175 • Published May 19, 2025
Toward Humanoid Brain-Body Co-design: Joint Optimization of Control and Morphology for Fall Recovery Paper • 2510.22336 • Published Nov 5, 2025
Interpretation-Oriented Cloud Removal via Observation-Anchored Residual Flow with Geo-Contextual Alignment Paper • 2607.02471 • Published Jul 2 • 3
Interpretation-Oriented Cloud Removal via Observation-Anchored Residual Flow with Geo-Contextual Alignment Paper • 2607.02471 • Published Jul 2 • 3
GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine? Paper • 2606.17861 • Published Jun 16 • 60
TRL-Bench: Standardizing Cross-Paradigm Representation-Level Evaluation of Tabular Encoders Paper • 2606.09323 • Published Jun 8 • 54
Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning Paper • 2604.16029 • Published Apr 17 • 22
BrandFusion: A Multi-Agent Framework for Seamless Brand Integration in Text-to-Video Generation Paper • 2603.02816 • Published Mar 3 • 2
Janus: Disaggregating Attention and Experts for Scalable MoE Inference Paper • 2512.13525 • Published Dec 15, 2025 • 6
MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols Paper • 2508.18240 • Published Aug 22, 2025 • 1
EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs Paper • 2509.09174 • Published Sep 11, 2025 • 62