NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video Paper • 2608.13210 • Published 9 days ago • 8
Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning Paper • 2608.18746 • Published 3 days ago • 15
H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models Paper • 2608.13049 • Published 9 days ago • 17
SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure Paper • 2608.11079 • Published 11 days ago • 17
HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Paper • 2608.06301 • Published 16 days ago • 35
GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience Paper • 2608.02392 • Published 19 days ago • 15