RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 5 days ago • 261
Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability Paper • 2610.03509 • Published 4 days ago • 13
WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents Paper • 2609.36887 • Published 7 days ago • 18
Beyond Future Prediction: Denoising as Generative Adaptation for Robot Control Paper • 2609.28339 • Published 13 days ago • 24
4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes Paper • 2610.03715 • Published 4 days ago • 21
Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies Paper • 2609.38155 • Published 7 days ago • 114
Think Before You Score: Thinking Reward Model for Visual Generation Paper • 2609.37372 • Published 7 days ago • 101
Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 10 days ago • 323
LEGO-Anything: Coding Agents for 3D Scene Reconstruction Paper • 2609.36380 • Published 8 days ago • 138