Running Reproduction: AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use 🎯 Collaborate with an AI agent to record benchmark findings
Running Reproduction: PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks 🎯 Collaborate with your coding agent to log PowerPoint task results