TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces Paper • 2609.33295 • Published 5 days ago • 68
RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments Paper • 2609.15364 • Published 18 days ago • 84
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses Paper • 2608.12307 • Published Aug 12 • 118
When Users Change Their Mind: Evaluating Interruptible Agents in Long-Horizon Web Navigation Paper • 2604.00892 • Published Apr 1 • 5
EpochX: Building the Infrastructure for an Emergent Agent Civilization Paper • 2603.27304 • Published Mar 28 • 21
Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models Paper • 2512.24618 • Published Dec 31, 2025 • 156
TestNUC: Enhancing Test-Time Computing Approaches through Neighboring Unlabeled Data Consistency Paper • 2502.19163 • Published Feb 26, 2025 • 1
From Web Search towards Agentic Deep Research: Incentivizing Search with Reasoning Agents Paper • 2506.18959 • Published Jun 23, 2025 • 5
Towards Agentic RAG with Deep Reasoning: A Survey of RAG-Reasoning Systems in LLMs Paper • 2507.09477 • Published Jul 13, 2025 • 89
A Survey on Large Language Model based Human-Agent Systems Paper • 2505.00753 • Published May 1, 2025 • 1