Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Paper • 2608.08160 • Published 12 days ago • 29
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design Paper • 2608.13560 • Published 7 days ago • 51
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers Paper • 2608.06867 • Published 13 days ago • 106
MatrAIx: Simulating the World with 8.3 Billion Persona Agents Paper • 2608.04205 • Published 16 days ago • 47
LLM-as-a-Verifier: A General-Purpose Verification Framework Paper • 2607.05391 • Published Jul 6 • 18
ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Paper • 2608.03972 • Published 16 days ago • 5
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations Paper • 2607.28956 • Published 20 days ago • 96
To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing Paper • 2607.28887 • Published 21 days ago • 20
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Paper • 2607.08317 • Published Jul 9 • 37
EgoCS-400K: An Egocentric Gameplay Dataset for World Models Paper • 2606.18180 • Published Jun 16 • 16
InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Paper • 2606.12195 • Published Jun 10 • 23
DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning Paper • 2606.08035 • Published Jun 6 • 16
Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception Paper • 2602.11858 • Published Feb 12 • 61