RRSI: Regularized Recursive Self-Improvement of Agent Harnesses Paper • 2609.24972 • Published 3 days ago • 177
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills Paper • 2609.02749 • Published 22 days ago • 402
UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City Paper • 2608.27456 • Published 28 days ago • 87
Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence Paper • 2608.21156 • Published Aug 21 • 63
SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science? Paper • 2608.19799 • Published Aug 20 • 67
ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents Paper • 2608.11878 • Published Aug 12 • 12
SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs Paper • 2608.03573 • Published Aug 6 • 61
SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Paper • 2608.09802 • Published Aug 10 • 84
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published Aug 14 • 285
HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published Aug 17 • 122
EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning Paper • 2608.06197 • Published Aug 6 • 47
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents Paper • 2607.23588 • Published Jul 26 • 127
Qwen-AgentWorld: Language World Models for General Agents Paper • 2606.24597 • Published Jun 23 • 165
Toward Generalist Autonomous Research via Hypothesis-Tree Refinement Paper • 2606.11926 • Published Jun 10 • 80