ClawBench & Agentic Web Benchmarks
ClawBench and related papers and open datasets for real-world web and computer-use agent evaluation.
Paper • 2604.08523 • Published • 265Note Paper: https://arxiv.org/abs/2604.08523 · code: https://github.com/TIGER-AI-Lab/ClawBench · project: https://claw-bench.com/
TIGER-Lab/ClawBench
Viewer • Updated • 283 • 526 • 1Note V1/V2 tasks and rubrics · repo: https://github.com/TIGER-AI-Lab/ClawBench
NAIL-Group/ClawBenchV1Trace
Updated • 173 • 1Note Full V1 execution traces for regrading and analysis.
TIGER-Lab/ClawBenchV2Trace
Updated • 344 • 1Note Full V2 execution traces for regrading and analysis.
GAIA: a benchmark for General AI Assistants
Paper • 2311.12983 • Published • 249Note GAIA: general AI assistant benchmark with real-world questions.
TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
Paper • 2412.14161 • Published • 51Note TheAgentCompany: consequential real-world tasks for LLM agents.
WebArena: A Realistic Web Environment for Building Autonomous Agents
Paper • 2307.13854 • Published • 27Note WebArena: realistic websites for autonomous web agents.
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Paper • 2404.07972 • Published • 52Note OSWorld: open-ended tasks in real computer environments.
The BrowserGym Ecosystem for Web Agent Research
Paper • 2412.05467 • Published • 24Note The BrowserGym Ecosystem for Web Agent Research.
ClawBench Leaderboard
🦀Can AI agents complete everyday online tasks?
Note Interactive leaderboard: https://huggingface.co/spaces/TIGER-Lab/ClawBench · project: https://claw-bench.com/
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
Paper • 2403.07718 • Published • 2Note WorkArena: enterprise knowledge-work tasks for web agents.
WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work Tasks
Paper • 2407.05291 • Published • 2Note WorkArena++: compositional planning and reasoning for web-agent tasks.
AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?
Paper • 2407.15711 • Published • 9Note AssistantBench: realistic, time-consuming web-agent tasks.
Mind2Web: Towards a Generalist Agent for the Web
Paper • 2306.06070 • Published • 21Note Mind2Web: generalist web-agent benchmark and dataset.
Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge
Paper • 2506.21506 • Published • 52Note Mind2Web 2: long-horizon agentic-search benchmark.
WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
Paper • 2401.13919 • Published • 33Note WebVoyager: end-to-end multimodal web-agent benchmark.
WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Paper • 2207.01206 • Published • 3Note WebShop: grounded language agents on real-world shopping tasks.
OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
Paper • 2510.24563 • Published • 23Note OSWorld-MCP: benchmark for MCP tool invocation in computer-use agents.
BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Paper • 2504.12516 • Published • 2Note BrowseComp: persistent browsing questions requiring hard-to-find web information.
BrowseComp-V^3: A Visual, Vertical, and Verifiable Benchmark for Multimodal Browsing Agents
Paper • 2602.12876 • Published • 14Note BrowseComp-V3: visual, vertical, and verifiable multimodal browsing benchmark.
K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts
Paper • 2606.02404 • Published • 59Note K-BrowseComp: web-browsing benchmark grounded in Korean contexts.
AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
Paper • 2405.14573 • PublishedNote AndroidWorld: dynamic benchmark for multimodal mobile agents.
AgentWebBench: Benchmarking Multi-Agent Coordination in Agentic Web
Paper • 2604.10938 • PublishedNote AgentWebBench: multi-agent coordination in agentic web settings.
BEARCUBS: A benchmark for computer-using web agents
Paper • 2503.07919 • PublishedNote BEARCUBS: compact benchmark for computer-using web agents.
WebCanvas: Benchmarking Web Agents in Online Environments
Paper • 2406.12373 • PublishedNote WebCanvas: web agents in online environments with live evaluation.
An Illusion of Progress? Assessing the Current State of Web Agents
Paper • 2504.01382 • Published • 4Note Online-Mind2Web: live evaluation across 300 realistic tasks and 136 websites.
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
Paper • 2402.05930 • Published • 39Note WebLINX: multi-turn conversational website navigation benchmark.
ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents
Paper • 2410.06703 • Published • 3Note ST-WebAgentBench: safety and trustworthiness benchmark for web agents.
WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation
Paper • 2607.06118 • PublishedNote WebRetriever: comprehensive benchmark for efficient web-agent evaluation.
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
Paper • 2407.00993 • Published • 1Note Mobile-Bench: evaluation benchmark for LLM-based mobile agents.
SPA-Bench: A Comprehensive Benchmark for SmartPhone Agent Evaluation
Paper • 2410.15164 • Published • 1Note SPA-Bench: smartphone-agent benchmark in interactive environments.
MobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive, and MCP-Augmented Environments
Paper • 2512.19432 • Published • 13Note MobileWorld: agent-user interactive and MCP-augmented mobile benchmark.
MobileWorldBench: Towards Semantic World Modeling For Mobile Agents
Paper • 2512.14014 • Published • 3Note MobileWorldBench: semantic world-model benchmark for mobile agents.
PhoneWorld: Scaling Phone-Use Agent Environments
Paper • 2605.29486 • Published • 12Note PhoneWorld: scalable phone-use agent environments.
AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents
Paper • 2602.06855 • Published • 83Note AIRS-Bench: end-to-end AI research science benchmark for LLM agents.
PaperBench: Evaluating AI's Ability to Replicate AI Research
Paper • 2504.01848 • Published • 37Note PaperBench: evaluating agents on reproducing AI research.
AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts
Paper • 2601.11044 • Published • 35Note AgencyBench: broad real-world autonomous-agent benchmark.
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Paper • 2410.07095 • Published • 8Note MLE-bench: machine-learning engineering benchmark for agents.
MLGym: A New Framework and Benchmark for Advancing AI Research Agents
Paper • 2502.14499 • Published • 196Note MLGym: framework and benchmark for AI research agents.
AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents
Paper • 2407.18901 • Published • 36Note AppWorld: controllable multi-app environment and benchmark for interactive coding agents. Project: https://appworld.dev/ · code: https://github.com/StonyBrookNLP/appworld
τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Paper • 2406.12045 • Published • 10Note Tool-agent-user benchmark with state-based evaluation and pass^k reliability. Paper: https://arxiv.org/abs/2406.12045 · code: https://github.com/sierra-research/tau-bench
TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents
Paper • 2606.28480 • Published • 48Note TUA-Bench evaluates general-purpose terminal-use agents across real-world, scientific, and engineering workflows. Paper: https://arxiv.org/abs/2606.28480
AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions
Paper • 2605.25707 • Published • 6Note Computer-use agent robustness benchmark under realistic environment corruptions. Paper: https://arxiv.org/abs/2605.25707 · project: https://AgentHijack.github.io
OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
Paper • 2606.29537 • Published • 22Note Long-horizon real-world computer-use workflows with authentic artifacts and stateful tasks. Paper: https://arxiv.org/abs/2606.29537
WorldGUI: Dynamic Testing for Comprehensive Desktop GUI Automation
Paper • 2502.08047 • Published • 28Note Dynamic desktop GUI benchmark with varied initial states across 10 applications. Paper: https://arxiv.org/abs/2502.08047 · code: https://github.com/showlab/WorldGUI
OSUniverse: Benchmark for Multimodal GUI-navigation AI Agents
Paper • 2505.03570 • Published • 8Note Multimodal desktop GUI-navigation benchmark with multi-application tasks and automated validation. Paper: https://arxiv.org/abs/2505.03570 · code: https://github.com/agentsea/osuniverse
Tur[k]ingBench: A Challenge Benchmark for Web Agents
Paper • 2403.11905 • PublishedNote Web-agent benchmark built from natural crowdsourcing HTML pages, with 32.2K instances across 158 tasks. Paper: https://arxiv.org/abs/2403.11905
WebSuite: Systematically Evaluating Why Web Agents Fail
Paper • 2406.01623 • PublishedNote Diagnostic web-agent benchmark organized around an action taxonomy to analyze failure modes. Paper: https://arxiv.org/abs/2406.01623
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Paper • 2508.20453 • Published • 63Note MCP-Bench evaluates multi-step real-world tasks across 28 MCP servers and 250 tools. Paper: https://arxiv.org/abs/2508.20453 · code: https://github.com/Accenture/mcp-bench
ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Paper • 2408.04682 • Published • 18Note Stateful, conversational tool-use benchmark with implicit dependencies and milestone-based evaluation. Paper: https://arxiv.org/abs/2408.04682 · code: https://github.com/apple/ToolSandbox
MCPToolBench++: A Large Scale AI Agent Model Context Protocol MCP Tool Use Benchmark
Paper • 2508.07575 • Published • 1Note Large-scale MCP tool-use benchmark spanning thousands of servers and multi-step calls. Paper: https://arxiv.org/abs/2508.07575
AgentBench: Evaluating LLMs as Agents
Paper • 2308.03688 • Published • 26Note Multi-environment agent benchmark spanning web browsing, web shopping, operating systems, databases, and other interactive settings. Paper: https://arxiv.org/abs/2308.03688 · code: https://github.com/THUDM/AgentBench
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Paper • 2304.08244 • Published • 1Note Tool-augmented LLM benchmark covering API invocation, tool selection, and execution. Paper: https://arxiv.org/abs/2304.08244
SWE-bench Goes Live!
Paper • 2505.23419 • Published • 21Note Live, contamination-resistant benchmark of issue resolution on recently created GitHub repositories. Paper: https://arxiv.org/abs/2505.23419 · project: https://www.swebench.com/
GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
Paper • 2508.18993 • Published • 5Note Code-agent benchmark for solving real-world tasks through repository-level leverage. Paper: https://arxiv.org/abs/2508.18993 · code: https://github.com/QuantaAlpha/GitTaskBench
τ-Knowledge: Evaluating Conversational Agents over Unstructured Knowledge
Paper • 2603.04370 • Published • 3Note Updated τ-bench implementation for tool-agent-user interaction and reliability evaluation. Paper: https://arxiv.org/abs/2603.04370 · code: https://github.com/sierra-research/tau2-bench
ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use
Paper • 2504.07981 • Published • 7Note Professional high-resolution GUI grounding benchmark covering 23 applications, five industries, and three operating systems. Paper: https://arxiv.org/abs/2504.07981 · leaderboard: https://gui-agent.github.io/grounding-leaderboard
ProBench: Benchmarking GUI Agents with Accurate Process Information
Paper • 2511.09157 • PublishedNote Mobile GUI benchmark that evaluates both final state and intermediate process information across 200+ tasks. Paper: https://arxiv.org/abs/2511.09157
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
Paper • 2401.13649 • Published • 1Note Multimodal web-agent benchmark extending WebArena with visual observations and tasks. Paper: https://arxiv.org/abs/2401.13649 · code: https://github.com/web-arena-x/visualwebarena
Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration
Paper • 1802.08802 • Published • 2Note Interactive web-interface benchmark with 100+ browser environments and programmatic rewards. Paper: https://arxiv.org/abs/1802.08802 · code: https://github.com/Farama-Foundation/miniwob-plusplus
AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents
Paper • 2406.13352 • PublishedNote Dynamic benchmark for prompt-injection attacks and defenses in tool-using agents. Paper: https://arxiv.org/abs/2406.13352 · code: https://github.com/ethz-spylab/agentdojo
Agent-SafetyBench: Evaluating the Safety of LLM Agents
Paper • 2412.14470 • Published • 12Note Comprehensive safety benchmark for LLM agents across interactive environments and risk categories. Paper: https://arxiv.org/abs/2412.14470 · code: https://github.com/thu-coai/Agent-SafetyBench
Identifying the Risks of LM Agents with an LM-Emulated Sandbox
Paper • 2309.15817 • PublishedNote LM-emulated sandbox benchmark for identifying risks in tool-using agents, covering 36 toolkits and 144 tests. Paper: https://arxiv.org/abs/2309.15817 · code: https://github.com/ryoungj/ToolEmu
Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents
Paper • 2602.13379 • Published • 3Note Multi-turn tool-using agent safety benchmark with taxonomy of risks and defenses. Paper: https://arxiv.org/abs/2602.13379 · code: https://github.com/CHATS-lab/ToolShield
MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Paper • 2310.03302 • PublishedNote Machine-learning experimentation benchmark for evaluating language agents on realistic ML workflows. Paper: https://arxiv.org/abs/2310.03302
MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents
Paper • 2503.01935 • Published • 30Note Multi-agent benchmark evaluating collaboration and competition dynamics across interactive tasks. Paper: https://arxiv.org/abs/2503.01935
OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
Paper • 2402.17553 • Published • 26Note Multimodal desktop-and-web benchmark measuring executable program generation for computer tasks. Paper: https://arxiv.org/abs/2402.17553
Android in the Wild: A Large-Scale Dataset for Android Device Control
Paper • 2307.10088 • Published • 12Note Large-scale Android device-control dataset with natural-language instructions, screenshots, and action traces. Paper: https://arxiv.org/abs/2307.10088 · code: https://github.com/google-research/google-research/tree/master/android_in_the_wild
AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents
Paper • 2410.24024 • Published • 49Note Reproducible Android-agent benchmark with 138 tasks across nine apps and predefined virtual devices. Paper: https://arxiv.org/abs/2410.24024 · code: https://github.com/THUDM/Android-Lab
AmineHA/WebArena-Verified
Viewer • Updated • 1.07k • 80 • 1Note HF dataset release of WebArena-Verified, a curated and reproducibly evaluated WebArena task set. Dataset: https://huggingface.co/datasets/AmineHA/WebArena-Verified · code: https://github.com/ServiceNow/webarena-verified
WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks
Paper • 2506.01952 • Published • 10Note Reproducible web benchmark with 532 tedious, memory-heavy, calculation, and long-term memory tasks. Paper: https://arxiv.org/abs/2506.01952 · project: https://webchorearena.github.io/
WebGames: Challenging General-Purpose Web-Browsing AI Agents
Paper • 2502.18356 • Published • 14Note Hermetic benchmark suite of 50+ interactive browser challenges with verifiable ground-truth solutions. Paper: https://arxiv.org/abs/2502.18356 · project: https://webgames.convergence.ai
ScienceWorld: Is your Agent Smarter than a 5th Grader?
Paper • 2203.07540 • PublishedNote Interactive text environment benchmark for grounded scientific reasoning and multi-step experimentation. Paper: https://arxiv.org/abs/2203.07540 · code: https://github.com/allenai/ScienceWorld
InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback
Paper • 2306.14898 • PublishedNote Interactive coding benchmark with execution feedback across Bash, SQL, Python, and other environments. Paper: https://arxiv.org/abs/2306.14898 · project: https://intercode-benchmark.github.io
Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking
Paper • 2606.15673 • Published • 13Note Process-level web-agent benchmark with 1,800 instances and automatic semantic state tracking. Paper: https://arxiv.org/abs/2606.15673 · project: https://jiwanchung.github.io/webstep/
BLADE: Benchmarking Language Model Agents for Data-Driven Science
Paper • 2408.09667 • PublishedNote Benchmark for data-driven scientific agents with expert analyses across 12 datasets and research questions. Paper: https://arxiv.org/abs/2408.09667
ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
Paper • 2410.05080 • Published • 21Note Science-agent benchmark with 102 expert-validated tasks drawn from 44 peer-reviewed publications. Paper: https://arxiv.org/abs/2410.05080
DSBench: How Far Are Data Science Agents to Becoming Data Science Experts?
Paper • 2409.07703 • Published • 66Note Realistic data-science-agent benchmark covering data analysis tasks and relative performance gaps. Paper: https://arxiv.org/abs/2409.07703
AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
Paper • 2510.21652 • Published • 4Note Scientific research-agent benchmark covering literature review, experiment reproduction, data analysis, and research planning. Paper: https://arxiv.org/abs/2510.21652
AgenticDataBench: A Comprehensive Benchmark for Data Agents
Paper • 2607.01647 • Published • 37Note Comprehensive benchmark for data agents with realistic tasks, fine-grained ground truth, and skill coverage across data workflows. Paper: https://arxiv.org/abs/2607.01647
SafeArena: Evaluating the Safety of Autonomous Web Agents
Paper • 2503.04957 • Published • 21Note Web-agent safety benchmark with 250 safe and 250 harmful tasks across four websites. Paper: https://arxiv.org/abs/2503.04957 · project: https://safearena.github.io
EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions
Paper • 2606.23654 • Published • 80Note Enterprise benchmark distilled from real workplace agent sessions into 852 reproducible tasks with artifacts and rubrics. Paper: https://arxiv.org/abs/2606.23654 · code: https://github.com/FrontisAI/EnterpriseClawBench
GTA: A Benchmark for General Tool Agents
Paper • 2407.08713 • Published • 17Note General Tool Agent benchmark with 229 human-written multimodal tasks and executable tool chains. Paper: https://arxiv.org/abs/2407.08713 · code: https://github.com/open-compass/GTA
GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows
Paper • 2604.15715 • Published • 4Note Hierarchical tool-agent benchmark spanning atomic tool use and long-horizon deliverable workflows. Paper: https://arxiv.org/abs/2604.15715 · code: https://github.com/open-compass/GTA
Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows
Paper • 2605.27922 • PublishedNote Benchmark isolating model-versus-harness effects across realistic agent workflows and shared task environments. Paper: https://arxiv.org/abs/2605.27922 · project: https://www.harness-bench.ai/
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
Paper • 2605.10912 • Published • 46Note Long-horizon benchmark running CLI agent harnesses with real tools in reproducible Docker tasks. Paper: https://arxiv.org/abs/2605.10912
REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites
Paper • 2504.11543 • Published • 2Note Benchmark and flexible harness for multi-turn agents on deterministic simulations of real-world websites. Paper: https://arxiv.org/abs/2504.11543
AgenticVBench: Can AI Agents Complete Real-World Post-Production Tasks?
Paper • 2605.27705 • Published • 1Note Professional post-production workflow benchmark with 100 agentic tasks from 20 industry experts; analyzes harness effects and deliverable quality. Paper: https://arxiv.org/abs/2605.27705
MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research
Paper • 2505.19955 • Published • 14Note Open-ended machine-learning research-agent benchmark with 201 tasks, rubric-based judging, and modular agent scaffold. Paper: https://arxiv.org/abs/2505.19955
Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale
Paper • 2409.08264 • Published • 48Note Scalable Windows desktop-agent benchmark with 150+ diverse OS tasks and reproducible parallel evaluation. Paper: https://arxiv.org/abs/2409.08264 · code: https://github.com/microsoft/WindowsAgentArena
CRAB: Cross-environment Agent Benchmark for Multimodal Language Model Agents
Paper • 2407.01511 • Published • 2Note Cross-environment benchmark for multimodal agents across websites, desktop computers, and mobile phones. Paper: https://arxiv.org/abs/2407.01511
PaperArena: An Evaluation Benchmark for Tool-Augmented Agentic Reasoning on Scientific Literature
Paper • 2510.10909 • PublishedNote Scientific-literature benchmark requiring cross-paper synthesis and multi-tool orchestration. Paper: https://arxiv.org/abs/2510.10909 · project: https://paperarena-ai.github.io/
MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
Paper • 2512.24565 • PublishedNote Real-world MCP tool-use benchmark with simulated tools, distractor selection, multi-step tasks, and execution-efficiency metrics. Paper: https://arxiv.org/abs/2512.24565
TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
Paper • 2510.04550 • Published • 2Note Trajectory-aware tool-use benchmark reporting intermediate tool selection and execution metrics beyond final accuracy. Paper: https://arxiv.org/abs/2510.04550 · dataset: https://huggingface.co/datasets/bigboss24/TRAJECT-Bench
MMAU: A Holistic Benchmark of Agent Capabilities Across Diverse Domains
Paper • 2407.18961 • Published • 40Note Holistic offline benchmark decomposing agent capabilities into understanding, reasoning, planning, problem-solving, and self-correction. Paper: https://arxiv.org/abs/2407.18961 · project: https://machinelearning.apple.com/research/mmau
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
Paper • 2510.11977 • PublishedNote Holistic agent evaluation leaderboard with large-scale logs, cost-aware comparisons, and behavior analysis. Paper: https://arxiv.org/abs/2510.11977 · project: https://hal.cs.princeton.edu/
Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents
Paper • 2606.04874 • PublishedNote Diagnostic planning benchmark with multimodal cases, feedback-conditioned planning, and robustness to broken or extraneous tools. Paper: https://arxiv.org/abs/2606.04874
MobiAgent: A Systematic Framework for Customizable Mobile Agents
Paper • 2509.00531 • Published • 8Note Mobile-agent framework releasing the MobiFlow benchmarking suite for real-world mobile scenarios. Paper: https://arxiv.org/abs/2509.00531
MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research
Paper • 2605.26114 • Published • 66Note Verifiable, highly parallel mobile GUI-agent simulation platform with deterministic state-diff evaluation. Paper: https://arxiv.org/abs/2605.26114 · code: https://github.com/Purewhiter/mobilegym
OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
Paper • 2605.18758 • Published • 16Note Omni-modal smartphone GUI benchmark combining audio, video, and image inputs. Paper: https://arxiv.org/abs/2605.18758 · project: https://omni-gui.github.io
MemGUI-Bench: Benchmarking Memory of Mobile GUI Agents in Dynamic Environments
Paper • 2602.06075 • Published • 14Note Memory-centric mobile GUI benchmark with 128 tasks across 26 apps and cross-session evaluation. Paper: https://arxiv.org/abs/2602.06075 · project: https://lgy0404.github.io/MemGUI-Bench/
MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
Paper • 2507.19478 • Published • 33Note Hierarchical multi-platform GUI-agent benchmark across Windows, macOS, Linux, iOS, Android, and Web. Paper: https://arxiv.org/abs/2507.19478
MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents
Paper • 2605.18652 • Published • 8Note Paper: https://arxiv.org/abs/2605.18652 · introduces MementoGUI-Bench for long-horizon GUI-agent decision-making · project: https://zzzmyyzeng.github.io/MementoGUI
Benchmarking and Improving GUI Agents in High-Dynamic Environments
Paper • 2604.25380 • PublishedNote Paper: https://arxiv.org/abs/2604.25380 · introduces DynamicGUIBench for high-dynamic GUI environments
MacArena: Benchmarking Computer Use Agents on an Online macOS Environment
Paper • 2606.06560 • Published • 4Note Paper: https://arxiv.org/abs/2606.06560 · MacArena: 421 manually verified macOS computer-use tasks across 50 applications.
WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces
Paper • 2606.09426 • Published • 107Note Paper: https://arxiv.org/abs/2606.09426 · project: https://weavebench.github.io/ · code: https://github.com/weavebench/WeaveBench · long-horizon GUI+CLI benchmark.
TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks
Paper • 2605.22535 • Published • 11Note Paper: https://arxiv.org/abs/2605.22535 · project: https://terminalworld.ai/ · code: https://github.com/EuniAI/TerminalWorld · real-world terminal-agent benchmark.
RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World Users
Paper • 2504.10445 • PublishedNote Paper: https://arxiv.org/abs/2504.10445 · project: https://scai.cs.jhu.edu/projects/RealWebAssist/ · long-horizon web assistance benchmark.
GUIDE: A Benchmark for Understanding and Assisting Users in Open-Ended GUI Tasks
Paper • 2603.25864 • PublishedNote Paper: https://arxiv.org/abs/2603.25864 · project: https://guide-bench.github.io/ · GUIDE evaluates GUI behavior understanding, intent prediction, and assistance.
PIRA-Bench: A Transition from Reactive GUI Agents to GUI-based Proactive Intent Recommendation Agents
Paper • 2603.08013 • Published • 15Note Paper: https://arxiv.org/abs/2603.08013 · PIRA-Bench evaluates proactive intent recommendation from continuous GUI observations.
MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents
Paper • 2606.16748 • Published • 7Note Paper: https://arxiv.org/abs/2606.16748 · project: https://mypcbench.com/ · code: https://github.com/ljang0/MyPCBench · benchmark for personalized computer-use agents.
G-FOCUS: Towards a Robust Method for Assessing UI Design Persuasiveness
Paper • 2505.05026 • Published • 18Note Paper: https://arxiv.org/abs/2505.05026 · G-FOCUS studies robust assessment of UI design persuasiveness and the WiserUI-Bench evaluation setting.
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
Paper • 2503.15661 • Published • 3Note Paper: https://arxiv.org/abs/2503.15661 · UI-Vision is a desktop-centric GUI benchmark for contextual visual interaction understanding.
Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction
Paper • 2606.29445 • Published • 28Note Paper: https://arxiv.org/abs/2606.29445 · project: https://vg-gui-tasker.github.io/ · code: https://github.com/VG-GUI-TASKER/VG-GUI-TASKER · VG-GUIBench evaluates video-guided long-horizon GUI tasks.
LongCLI-Bench: A Preliminary Benchmark and Study for Long-horizon Agentic Programming in Command-Line Interfaces
Paper • 2602.14337 • Published • 15Note Paper: https://arxiv.org/abs/2602.14337 · LongCLI-Bench evaluates long-horizon agentic programming workflows in command-line environments.
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
Paper • 2607.08964 • Published • 76Note Paper: https://arxiv.org/abs/2607.08964 · Long-Horizon-Terminal-Bench evaluates 46 long-horizon terminal tasks across nine categories.
Covering Human Action Space for Computer Use: Data Synthesis and Benchmark
Paper • 2605.12501 • Published • 16Note Paper: https://arxiv.org/abs/2605.12501 · Microsoft-led benchmark and data-synthesis pipeline covering human action space for computer use.
AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
Paper • 2604.02947 • Published • 19Note Paper: https://arxiv.org/abs/2604.02947 · AgentHazard benchmarks harmful behavior and robustness of computer-use agents.
LPS-Bench: Benchmarking Safety Awareness of Computer-Use Agents in Long-Horizon Planning under Benign and Adversarial Scenarios
Paper • 2602.03255 • PublishedNote Paper: https://arxiv.org/abs/2602.03255 · code: https://github.com/tychenn/LPS-Bench · LPS-Bench evaluates long-horizon planning-time safety awareness.
Code Agent can be an End-to-end System Hacker: Benchmarking Real-world Threats of Computer-use Agent
Paper • 2510.06607 • Published • 4Note Paper: https://arxiv.org/abs/2510.06607 · AdvCUA benchmarks real-world security threats of computer-use agents in multi-host environments.
Human-Guided Harm Recovery for Computer Use Agents
Paper • 2604.18847 • PublishedNote Paper: https://arxiv.org/abs/2604.18847 · BackBench evaluates recovery from harmful states in computer-use tasks.
GUI vs. CLI: Execution Bottlenecks in Screen-Only and Skill-Mediated Computer-Use Agents
Paper • 2606.24551 • Published • 28Note Paper: https://arxiv.org/abs/2606.24551 · matched 440-task benchmark comparing GUI and CLI execution layers for computer-use agents.
BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks
Paper • 2510.02418 • Published • 2Note Paper: https://arxiv.org/abs/2510.02418 · BrowserArena evaluates web agents on live real-world navigation tasks with human step-level feedback.
Embodied Web Agents: Bridging Physical-Digital Realms for Integrated Agent Intelligence
Paper • 2506.15677 • Published • 23Note Paper: https://arxiv.org/abs/2506.15677 · project: https://embodied-web-agent.github.io/ · Embodied Web Agents Benchmark spans coordinated physical and digital web tasks.
Web-Shepherd: Advancing PRMs for Reinforcing Web Agents
Paper • 2505.15277 • Published • 105Note Paper: https://arxiv.org/abs/2505.15277 · WebRewardBench evaluates process reward models for web-agent trajectory assessment.
MAS-Bench: A Unified Benchmark for Shortcut-Augmented Hybrid Mobile GUI Agents
Paper • 2509.06477 • Published • 3Note Paper: https://arxiv.org/abs/2509.06477 · project: https://pengxiang-zhao.github.io/MAS-Bench · MAS-Bench evaluates shortcut-augmented hybrid mobile GUI agents.
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
Paper • 2504.18575 • PublishedNote Paper: https://arxiv.org/abs/2504.18575 · code: https://github.com/facebookresearch/wasp · WASP benchmarks end-to-end web-agent security against prompt injection.
VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability Diagnostics
Paper • 2604.06182 • Published • 4Note Paper: https://arxiv.org/abs/2604.06182 · code: https://github.com/inclusionAI/UI-Venus/tree/VenusBench-Mobile · VenusBench-Mobile evaluates user-centric mobile GUI agents under realistic environment variation.
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
Paper • 2510.01354 • Published • 3Note Paper: https://arxiv.org/abs/2510.01354 · code: https://github.com/Norrrrrrr-lyn/WAInjectBench · WAInjectBench evaluates prompt-injection detection for web agents.
BrowseSafe: Understanding and Preventing Prompt Injection Within AI Browser Agents
Paper • 2511.20597 • PublishedNote Paper: https://arxiv.org/abs/2511.20597 · BrowseSafe benchmarks prompt-injection risks and defenses for AI browser agents.
SimuWoB: Simulating Real-World Mobile Apps for Fast and Faithful GUI Agent Benchmarking
Paper • 2605.25160 • Published • 9Note Paper: https://arxiv.org/abs/2605.25160 · SimuWoB benchmarks mobile GUI agents in high-fidelity synthetic app environments with automatic rewards.
GUI-CEval: A Hierarchical and Comprehensive Chinese Benchmark for Mobile GUI Agents
Paper • 2603.15039 • PublishedNote Paper: https://arxiv.org/abs/2603.15039 · GUI-CEval evaluates mobile GUI agents across perception, planning, reflection, execution, and evaluation.
Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization
Paper • 2604.09574 • Published • 30Note Paper: https://arxiv.org/abs/2604.09574 · Turing Test on Screen introduces the Agent Humanization Benchmark for mobile GUI interaction behavior.
WebGuard: Building a Generalizable Guardrail for Web Agents
Paper • 2507.14293 • Published • 1Note Paper: https://arxiv.org/abs/2507.14293 · WebGuard provides a dataset and evaluation framework for web-agent action-risk guardrails.
TimeWarp: Evaluating Web Agents by Revisiting the Past
Paper • 2603.04949 • PublishedNote Paper: https://arxiv.org/abs/2603.04949 · TimeWarp evaluates web-agent robustness across historical UI and layout versions.
A3: Android Agent Arena for Mobile GUI Agents
Paper • 2501.01149 • Published • 22Note Paper: https://arxiv.org/abs/2501.01149 · project: https://yuxiangchai.github.io/Android-Agent-Arena/ · A3 evaluates mobile GUI agents on 21 apps and real-world tasks.
StressWeb: A Diagnostic Benchmark for Web Agent Robustness under Realistic Interaction Variability
Paper • 2604.16385 • Published • 1Note Paper: https://arxiv.org/abs/2604.16385 · StressWeb diagnoses web-agent robustness under realistic interaction variability.
GhostEI-Bench: Do Mobile Agents Resilience to Environmental Injection in Dynamic On-Device Environments?
Paper • 2510.20333 • PublishedNote Paper: https://arxiv.org/abs/2510.20333 · GhostEI-Bench evaluates mobile-agent resilience to adversarial environmental UI injections.
DigiData: Training and Evaluating General-Purpose Mobile Control Agents
Paper • 2511.07413 • Published • 7Note Paper: https://arxiv.org/abs/2511.07413 · project: https://facebookresearch.github.io/DigiData/ · DigiData-Bench evaluates general-purpose mobile control agents on complex real-world tasks.
AutoGUI-v2: A Comprehensive Multi-Modal GUI Functionality Understanding Benchmark
Paper • 2604.24441 • Published • 4Note Paper: https://arxiv.org/abs/2604.24441 · AutoGUI-v2 evaluates GUI functionality understanding and interaction outcome prediction across six operating systems.
WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction Traces
Paper • 2603.05295 • PublishedNote Paper: https://arxiv.org/abs/2603.05295 · dataset: https://huggingface.co/datasets/webagentlab/WebChain · WebChain and WebChainBench provide large-scale human-annotated real-world web interaction trajectories.
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
Paper • 2603.22529 • Published • 7Note Paper: https://arxiv.org/abs/2603.22529 · project: https://ego2web.github.io/ · Ego2Web benchmarks web agents grounded in egocentric video and online task execution.
WebTestBench: Evaluating Computer-Use Agents towards End-to-End Automated Web Testing
Paper • 2603.25226 • PublishedNote Paper: https://arxiv.org/abs/2603.25226 · WebTestBench evaluates computer-use agents on end-to-end automated web testing workflows.
Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents
Paper • 2605.29447 • Published • 21Note Paper: https://arxiv.org/abs/2605.29447 · GUI-RobustEval benchmarks error awareness and recovery across policy-induced GUI failures.
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields
Paper • 2606.11042 • Published • 22Note Paper: https://arxiv.org/abs/2606.11042 · Workflow-GYM evaluates long-horizon computer-use workflows in real-world professional software environments.
WindowsWorld: A Process-Centric Benchmark of Autonomous GUI Agents in Professional Cross-Application Environments
Paper • 2604.27776 • Published • 15Note Paper: https://arxiv.org/abs/2604.27776 · code: https://github.com/HITsz-TMG/WindowsWorld · WindowsWorld benchmarks process-aware cross-application professional GUI workflows.
Gym-Anything: Turn any Software into an Agent Environment
Paper • 2604.06126 • Published • 1Note Paper: https://arxiv.org/abs/2604.06126 · project: https://cmu-l3.github.io/gym-anything/ · Gym-Anything produces CUA-World, 10K+ tasks across 200+ software environments.
HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks
Paper • 2604.09937 • PublishedNote Paper: https://arxiv.org/abs/2604.09937 · HealthAdminBench evaluates computer-use agents on healthcare administration workflows.
HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents
Paper • 2606.31179 • Published • 7Note Paper: https://arxiv.org/abs/2606.31179 · HealthAgentBench evaluates agentic healthcare workflows across seven environment categories.
CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents
Paper • 2605.25624 • Published • 35Note CUA-Gym: 32,112 verified computer-use training tuples across 110 environments; code/data/project links: https://arxiv.org/abs/2605.25624
TRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenarios
Paper • 2602.01675 • Published • 10Note TRIP-Bench evaluates long-horizon interactive agents in realistic travel-planning scenarios with multi-turn constraints and tool use; paper: https://arxiv.org/abs/2602.01675
OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
Paper • 2506.14866 • Published • 5Note OS-Harm: 150 computer-use safety tasks; project/repo: https://github.com/tml-epfl/os-harm · paper: https://arxiv.org/abs/2506.14866
DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards
Paper • 2508.17398 • Published • 2Note DashboardQA benchmarks GUI agents on interactive dashboards; repo: https://github.com/vis-nlp/DashboardQA · paper: https://arxiv.org/abs/2508.17398
Benchmark Test-Time Scaling of General LLM Agents
Paper • 2602.18998 • Published • 10Note General AgentBench is a unified benchmark spanning search, coding, reasoning, and tool use; repo: https://github.com/cxcscmu/General-AgentBench · paper: https://arxiv.org/abs/2602.18998
Can AI Agents Answer Your Data Questions? A Benchmark for Data Agents
Paper • 2603.20576 • Published • 4Note Data Agent Benchmark (DAB) evaluates agents over multi-database enterprise data workflows; repo: https://github.com/ucbepic/DataAgentBench · paper: https://arxiv.org/abs/2603.20576
UI-CUBE: Enterprise-Grade Computer Use Agent Benchmarking Beyond Task Accuracy to Operational Reliability
Paper • 2511.17131 • PublishedNote UI-CUBE evaluates enterprise computer-use reliability across 226 UI tasks and workflow scenarios; paper: https://arxiv.org/abs/2511.17131
thomas-kuntz/os-harm
Viewer • Updated • 149 • 272Note OS-Harm dataset: 150 computer-use safety tasks; source project: https://github.com/tml-epfl/os-harm
ahmed-masry/DashboardQA
Viewer • Updated • 292 • 95Note DashboardQA dataset for interactive dashboard question answering and GUI-agent evaluation; source repo: https://github.com/vis-nlp/DashboardQA
agentjudge-anon/GeneralAgentBench
Viewer • Updated • 1.45k • 54Note General AgentBench dataset spanning search, coding, reasoning, and tool-use tasks; source repo: https://github.com/cxcscmu/General-AgentBench
ruiyingm/DataAgentBench-data
Updated • 945Note DataAgentBench data for multi-database enterprise data-agent workflows; source repo: https://github.com/ucbepic/DataAgentBench
General Agent Evaluation
Paper • 2602.22953 • Published • 12Note General Agent Evaluation proposes a unified protocol and open leaderboard across six environments; HF paper page: https://huggingface.co/papers/2602.22953
MedBrowseComp: Benchmarking Medical Deep Research and Computer Use
Paper • 2505.14963 • Published • 2Note MedBrowseComp benchmarks medical deep research and computer use over live domain-specific knowledge bases; project: https://moreirap12.github.io/mbc-browse-app/ · paper: https://arxiv.org/abs/2505.14963
Computer Use at the Edge of the Statistical Precipice
Paper • 2605.08261 • PublishedNote Computer Use at the Edge of the Statistical Precipice analyzes evaluation pitfalls and statistical reliability in computer-use benchmarks; paper: https://arxiv.org/abs/2605.08261
MacAgentBench: Benchmarking AI Agents on Real-World macOS Desktop
Paper • 2606.22557 • PublishedNote MacAgentBench evaluates 676 macOS tasks across 25 applications with deterministic and fine-grained scoring; code: https://github.com/JetAstra/MacAgentBench · paper: https://arxiv.org/abs/2606.22557
Hardening Agent Benchmarks with Adversarial Hacker-Fixer Loops
Paper • 2606.08960 • Published • 1Note Hardening Agent Benchmarks with Adversarial Hacker-Fixer Loops studies exploitable verifiers across terminal-agent benchmarks; paper: https://arxiv.org/abs/2606.08960
WebForge: Breaking the Realism-Reproducibility-Scalability Trilemma in Browser Agent Benchmark
Paper • 2604.10988 • Published • 3Note WebForge-Bench contains 934 reproducible browser-agent tasks with controlled multidimensional difficulty; code: https://github.com/yuandaxia2001/WebForge · paper: https://arxiv.org/abs/2604.10988
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
Paper • 2504.08942 • Published • 29Note AgentRewardBench evaluates automatic judges for web-agent trajectories using 1,302 expert-reviewed trajectories; project: https://agent-reward-bench.github.io/ · paper: https://arxiv.org/abs/2504.08942
CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark
Paper • 2409.11363 • Published • 2Note CORE-Bench evaluates computational reproducibility agents on 270 tasks based on 90 scientific papers; code: https://github.com/siegelz/core-bench · paper: https://arxiv.org/abs/2409.11363
siegelz/core-bench
Preview • Updated • 421 • 9Note CORE-Bench dataset for computational reproducibility agents; source repo: https://github.com/siegelz/core-bench
McGill-NLP/safearena
Updated • 116 • 5Note SafeArena dataset for safety evaluation of autonomous web agents; source project: https://safearena.github.io
AssistantBench/AssistantBench
Viewer • Updated • 214 • 2.14k • 23Note AssistantBench dataset for realistic, time-consuming web-agent tasks; paper: https://arxiv.org/abs/2407.15711
osunlp/Mind2Web
Viewer • Updated • 253 • 4.3k • 128Note Mind2Web dataset for generalist web-agent tasks and trajectories; benchmark paper: https://arxiv.org/abs/2306.06070
xlangai/osworld_v2_tasks
Viewer • Updated • 1 • 1.68k • 18Note OSWorld 2.0 long-horizon desktop task set; benchmark paper: https://arxiv.org/abs/2606.29537
McGill-NLP/weblinx-browsergym
Updated • 2.21k • 4Note WebLINX/BrowserGym resources for browser-agent interaction; project: https://mcgill-nlp.github.io/weblinx/
Establishing Best Practices for Building Rigorous Agentic Benchmarks
Paper • 2507.02825 • Published • 1Note Agentic Benchmark Checklist (ABC) provides rigorous guidelines for task setup and reward design in agentic benchmarks; paper: https://arxiv.org/abs/2507.02825
AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios
Paper • 2505.16944 • Published • 8Note AgentIF benchmarks instruction following in realistic agentic scenarios with 707 human-annotated instructions; dataset: https://huggingface.co/datasets/THU-KEG/AgentIF · paper: https://arxiv.org/abs/2505.16944
PACE: A Proxy for Agentic Capability Evaluation
Paper • 2607.02032 • Published • 20Note PACE predicts expensive agentic benchmark performance from a small atomic subset; paper: https://arxiv.org/abs/2607.02032
THU-KEG/AgentIF
Viewer • Updated • 707 • 482 • 8Note AgentIF dataset: 707 human-annotated instructions for realistic agentic scenarios; paper: https://arxiv.org/abs/2505.16944
OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth?
Paper • 2507.19132 • PublishedNote OS-MAP evaluates 416 computer-use tasks across 15 applications by automation level and generalization scope; code: https://github.com/OS-Copilot/OS-Map · paper: https://arxiv.org/abs/2507.19132
Measuring Harmfulness of Computer-Using Agents
Paper • 2508.00935 • PublishedNote CUAHarm benchmarks misuse risks of computer-using agents on 104 expert-written scenarios; code: https://github.com/db-ol/CUAHarm · paper: https://arxiv.org/abs/2508.00935
CUAHarm/CUAHarm
Viewer • Updated • 104 • 50 • 1Note CUAHarm dataset: expert-written computer-use misuse scenarios; paper: https://arxiv.org/abs/2508.00935
EconWebArena: Benchmarking Autonomous Agents on Economic Tasks in Realistic Web Environments
Paper • 2506.08136 • PublishedNote EconWebArena benchmarks 360 multimodal economic web tasks across 82 authoritative websites; dataset: https://huggingface.co/datasets/EconWebArena/EconWebArena · paper: https://arxiv.org/abs/2506.08136
WebArXiv: Evaluating Multimodal Agents on Time-Invariant arXiv Tasks
Paper • 2507.00938 • PublishedNote WebArXiv provides 275 static, time-invariant web-agent tasks with deterministic ground truths; paper: https://arxiv.org/abs/2507.00938
EconWebArena/EconWebArena
Viewer • Updated • 360 • 22.2k • 2Note EconWebArena dataset: 360 curated economic web-agent tasks; paper: https://arxiv.org/abs/2506.08136
WARC-Bench: Web Archive Based Benchmark for GUI Subtask Executions
Paper • 2510.09872 • PublishedNote WARC-Bench provides 438 sandboxed web-archive GUI subtasks with deterministic rewards; ICLR 2026 paper: https://arxiv.org/abs/2510.09872
Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies
Paper • 2605.03596 • Published • 12Note Workspace-Bench evaluates 388 tasks over large-scale file dependencies, with 20,476 files and 7,399 rubrics; paper: https://arxiv.org/abs/2605.03596
DeskCraft: Benchmarking Desktop Agents on Professional Workflows and Human-in-the-Loop Collaboration
Paper • 2606.03103 • Published • 3Note DeskCraft benchmarks 538 professional desktop tasks with proactive human-agent interaction; code: https://github.com/mrwwk/DeskCraft · paper: https://arxiv.org/abs/2606.03103
ProSoftArena: Benchmarking Hierarchical Capabilities of Multimodal Agents in Professional Software Environments
Paper • 2601.02399 • PublishedNote ProSoftArena evaluates 436 professional-software GUI tasks across 13 applications; project: https://prosoftarena.github.io · paper: https://arxiv.org/abs/2601.02399
Workspace-Bench/Workspace-Bench
Viewer • Updated • 388 • 4.46k • 4Note Workspace-Bench full task dataset; paper: https://arxiv.org/abs/2605.03596
Workspace-Bench/Workspace-Bench-Lite
Viewer • Updated • 100 • 2.4k • 2Note Workspace-Bench-Lite 100-task subset for lower-cost workspace-agent evaluation; paper: https://arxiv.org/abs/2605.03596
Workspace-Bench/Workspace-Bench-Workspaces
Updated • 495Note Workspace-Bench workspace fixtures and file-dependency environments; paper: https://arxiv.org/abs/2605.03596
ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
Paper • 2604.05172 • Published • 25Note ClawsBench (distinct from ClawBench) evaluates productivity-agent capability and safety in simulated multi-service workspaces; paper: https://arxiv.org/abs/2604.05172
benchflow/ClawsBench
Viewer • Updated • 7.83k • 261 • 4Note ClawsBench dataset with agent traces for simulated productivity workflows; paper: https://arxiv.org/abs/2604.05172
Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows
Paper • 2604.28139 • Published • 43Note Claw-Eval-Live is a live workflow-agent benchmark with refreshable demand signals and verifiable execution evidence; project: https://claw-eval-live.github.io · paper: https://arxiv.org/abs/2604.28139
claw-eval-live/claw-eval-live
Viewer • Updated • 105 • 32 • 1Note Claw-Eval-Live release snapshot with 105 workflow tasks, mock services, fixtures, and graders; paper: https://arxiv.org/abs/2604.28139
Can LLMs Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise Environments
Paper • 2510.27287 • PublishedNote EnterpriseBench evaluates 500 enterprise tasks spanning software engineering, HR, finance, and administration; paper: https://arxiv.org/abs/2510.27287
EnterpriseBench Corecraft: Training Generalizable Agents on High-Fidelity RL Environments
Paper • 2602.16179 • PublishedNote EnterpriseBench CoreCraft provides a high-fidelity customer-support simulation with 2,500 entities and 23 tools for agent training/evaluation; paper: https://arxiv.org/abs/2602.16179
EXP-Bench: Can AI Conduct AI Research Experiments?
Paper • 2505.24785 • Published • 24Note EXP-Bench evaluates 461 end-to-end AI research experiment tasks sourced from 51 top-tier papers; code: https://github.com/Just-Curieous/Curie/tree/main/benchmark/exp_bench
FDABench: A Benchmark for Data Agents on Analytical Queries over Heterogeneous Data
Paper • 2509.02473 • PublishedNote FDABench evaluates 2,007 data-agent analytical tasks over heterogeneous sources; paper: https://arxiv.org/abs/2509.02473
GAIR/AgencyBench
Viewer • Updated • 42 • 694 • 3Note AgencyBench dataset for long-context, long-horizon autonomous-agent evaluation; paper: https://arxiv.org/abs/2601.11044
FDAbench2026/FDAbench-Full
Viewer • Updated • 2.01k • 655 • 1Note FDABench full dataset for heterogeneous data-agent analytical tasks; paper: https://arxiv.org/abs/2509.02473
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
Paper • 2506.11763 • Published • 74Note DeepResearch Bench evaluates 100 PhD-level research tasks across 22 fields with retrieval and report-quality metrics; code: https://github.com/Ayanami0730/deep_research_bench
EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer
Paper • 2607.05202 • Published • 1Note EvoAgentBench evaluates procedural self-evolution and ability transfer across web research, reasoning, software engineering, and knowledge-work domains; dataset: https://huggingface.co/datasets/EverMind-AI/EvoAgentBench
AgentSearchBench: A Benchmark for AI Agent Search in the Wild
Paper • 2604.22436 • Published • 14Note AgentSearchBench evaluates execution-grounded retrieval and ranking of nearly 10,000 real-world AI agents; code: https://github.com/Bingo-W/AgentSearchBench
EverMind-AI/EvoAgentBench
Viewer • Updated • 578 • 402 • 14Note EvoAgentBench dataset for ability-guided agent self-evolution evaluation; paper: https://arxiv.org/abs/2607.05202
AgentSearch/AgentSearchBench-Tasks
Viewer • Updated • 4.01k • 178 • 3Note AgentSearchBench task split for execution-grounded agent discovery; paper: https://arxiv.org/abs/2604.22436
DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis
Paper • 2508.20033 • Published • 10Note DeepScholar-Bench evaluates live research synthesis, retrieval quality, citation verifiability, and knowledge synthesis; code: https://github.com/guestrin-lab/deepscholar-bench
DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert Report
Paper • 2601.08536 • Published • 3Note DeepResearch Bench II evaluates 132 cross-domain research tasks with 9,430 expert-derived fine-grained rubrics; paper: https://arxiv.org/abs/2601.08536
Enterprise Deep Research: Steerable Multi-Agent Deep Research for Enterprise Analytics
Paper • 2510.17797 • Published • 11Note Enterprise Deep Research (Salesforce AI Research) provides a steerable multi-agent research system with public code and EDR-200 data; code: https://github.com/SalesforceAIResearch/enterprise-deep-research
DRBench: A Realistic Benchmark for Enterprise Deep Research
Paper • 2510.00172 • Published • 3Note DRBench evaluates 15 realistic enterprise deep-research tasks spanning public web and private company knowledge; code: https://github.com/ServiceNow/drbench
deepscholar-bench/DeepScholarBench
Viewer • Updated • 2.81k • 67 • 2Note DeepScholar-Bench dataset for related-work synthesis and citation analysis; paper: https://arxiv.org/abs/2508.20033
Salesforce/EDR-200
Viewer • Updated • 201 • 91 • 15Note Salesforce Enterprise Deep Research trajectory dataset; paper: https://arxiv.org/abs/2510.17797
ServiceNow/drbench
Viewer • Updated • 215 • 4.01k • 5Note ServiceNow DRBench dataset for enterprise deep-research evaluation; paper: https://arxiv.org/abs/2510.00172
LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild
Paper • 2510.14240 • Published • 13Note Salesforce AI Research; live web deep-research benchmark with 100 expert-curated tasks and citation-grounded evaluation. Paper: https://arxiv.org/abs/2510.14240 · code: https://github.com/SalesforceAIResearch/LiveResearchBench · dataset: https://huggingface.co/datasets/Salesforce/LiveResearchBench · project: https://livedeepresearch.github.io/
MMDeepResearch-Bench: A Benchmark for Multimodal Deep Research Agents
Paper • 2601.12346 • Published • 52Note Multimodal deep-research benchmark with visual evidence, citation alignment, and 140 expert-crafted tasks. Paper: https://arxiv.org/abs/2601.12346 · code: https://github.com/AIoT-MLSys-Lab/MMDeepResearch-Bench · dataset: https://huggingface.co/datasets/MMDR-2025/MMdeepresearch · project: https://mmdeepresearch-bench.github.io/
How Far Are We from Genuinely Useful Deep Research Agents?
Paper • 2512.01948 • Published • 58Note FINDER deep-research benchmark with 100 human-curated tasks, 419 checklist items, and a failure taxonomy for evidence integration. Paper: https://arxiv.org/abs/2512.01948
How Far Are We From True Auto-Research?
Paper • 2605.19156 • Published • 2Note ResearchArena evaluates end-to-end agentic research workflows with artifact-aware and manuscript-level review. Paper: https://arxiv.org/abs/2605.19156
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction
Paper • 2607.20911 • Published • 25Note Tencent WorkBuddy Bench: multi-domain coding-agent benchmark spanning code, web, office, and security tasks with contamination-resistant construction. Paper: https://arxiv.org/abs/2607.20911 · dataset: https://huggingface.co/datasets/tencent/workbuddy-bench
Agentic-MME: What Agentic Capability Really Brings to Multimodal Intelligence?
Paper • 2604.03016 • Published • 37Note Agentic-MME evaluates multimodal agent capability through tool usage, process efficiency, and outcome verification. Paper: https://arxiv.org/abs/2604.03016
A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks
Paper • 2605.28556 • Published • 75Note TASTE improves agent-benchmark coverage and difficulty through adaptive tool-sequence generation. Paper: https://arxiv.org/abs/2605.28556
DREAM: Deep Research Evaluation with Agentic Metrics
Paper • 2602.18940 • Published • 14Note DREAM evaluates deep-research agents with agentic, temporally aware, citation-grounded metrics. Paper: https://arxiv.org/abs/2602.18940
SGR-Bench: Benchmarking Search Agents on State-Gated Retrieval
Paper • 2605.22219 • PublishedNote SGR-Bench evaluates web search agents on state-gated retrieval across 100 expert-curated tasks and 12 public data ecosystems. Paper: https://arxiv.org/abs/2605.22219 · dataset: https://huggingface.co/datasets/PKUAIWeb/SGR-BENCH
PKUAIWeb/SGR-BENCH
Viewer • Updated • 100 • 228Note Official SGR-Bench dataset for state-gated retrieval web-agent evaluation. Paper: https://arxiv.org/abs/2605.22219
deepweb-bench-anon/deepweb-bench
Updated • 79 • 3Note DeepWeb-Bench public dataset for long-horizon, cross-source deep research evaluation. Paper: https://arxiv.org/abs/2605.21482 · project: https://sixiongxie1001-dot.github.io/deep-research-benchmark2.0/
CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing
Paper • 2605.19484 • Published • 21Note CutVerse evaluates GUI agents on 186 long-horizon professional media post-production tasks across seven applications. Paper: https://arxiv.org/abs/2605.19484
π-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows
Paper • 2605.14678 • Published • 108Note π-Bench evaluates proactive personal assistant agents on 100 multi-turn tasks with hidden intents, inter-task dependencies, and cross-session continuity. Paper: https://arxiv.org/abs/2605.14678
CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agents
Paper • 2606.22883 • Published • 37Note CLI-Universe synthesizes verifiable terminal-agent tasks with Dockerized environments and strict execution checks. Paper: https://arxiv.org/abs/2606.22883
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
Paper • 2510.24358 • PublishedNote PRDBench evaluates project-level coding agents on 50 real-world Python projects with PRD-based and agent-as-judge criteria. Paper: https://arxiv.org/abs/2510.24358
PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems
Paper • 2606.22388 • Published • 95Note PlanBench-XL evaluates long-horizon planning over 327 retail tasks and 1,665 tools, including missing, failing, and distracting tool conditions. Paper: https://arxiv.org/abs/2606.22388 · dataset: https://huggingface.co/datasets/JiayuJeff/PlanBench-XL
ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research
Paper • 2606.07591 • Published • 102Note ResearchClawBench evaluates end-to-end autonomous scientific research across 40 tasks and 10 scientific domains with multimodal expert rubrics. Paper: https://arxiv.org/abs/2606.07591 · dataset: https://huggingface.co/datasets/InternScience/ResearchClawBench
InternScience/ResearchClawBench
Benchmark • Updated • 57 • 32.7k • 14Note Official ResearchClawBench benchmark dataset for autonomous scientific research evaluation. Paper: https://arxiv.org/abs/2606.07591
EvoBrowseComp: Benchmarking Search Agents on Evolving Knowledge
Paper • 2606.13120 • Published • 4Note EvoBrowseComp is an evolving web-search benchmark with 800 contamination-free questions synthesized from live-web traversal. Paper: https://arxiv.org/abs/2606.13120 · dataset: https://huggingface.co/datasets/Krystalan/EvoBrowseComp
Krystalan/EvoBrowseComp
Viewer • Updated • 800 • 106 • 1Note Official EvoBrowseComp dataset for contamination-resistant, evolving web-search-agent evaluation. Paper: https://arxiv.org/abs/2606.13120
DataClawBench: An Agent Benchmark for Exploratory Real-World Financial Data Analysis
Paper • 2605.02503 • Published • 1Note DataClawBench evaluates exploratory real-world financial data analysis with 492 cross-domain tasks and native data noise. Paper: https://arxiv.org/abs/2605.02503 · dataset: https://huggingface.co/datasets/GTML-LAB/DataClaw
GTML-LAB/DataClaw
Viewer • Updated • 492 • 687 • 1Note Official DataClaw dataset for exploratory agentic data-analysis evaluation. Paper: https://arxiv.org/abs/2605.02503
StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents
Paper • 2607.22798 • Published • 55Note StateAct, from Salesforce AI Research, evaluates a program-state-first harness for long-horizon computer-use agents and reports results on OSWorld 2.0 and related suites. Paper: https://arxiv.org/abs/2607.22798
Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks
Paper • 2606.12344 • Published • 71Note Claw-SWE-Bench evaluates OpenClaw-style agent harnesses on 350 multilingual software-engineering tasks with standardized adapters and cost accounting. Paper: https://arxiv.org/abs/2606.12344 · code: https://github.com/opensquilla/claw-swe-bench · dataset: https://huggingface.co/datasets/TokenRhythm/Claw-SWE-Bench
TokenRhythm/Claw-SWE-Bench
Viewer • Updated • 430 • 260 • 6Note Official Claw-SWE-Bench dataset for reproducible OpenClaw-style coding-agent harness evaluation. Paper: https://arxiv.org/abs/2606.12344
A History-Aware Visually Grounded Critic for Computer Use Agents
Paper • 2606.11078 • Published • 2Note HiViG studies history-aware, visually grounded critics for long-horizon GUI-agent evaluation across web, mobile, and desktop benchmarks. Paper: https://arxiv.org/abs/2606.11078
When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents
Paper • 2602.08235 • Published • 1Note AutoElicit systematically evaluates unsafe unintended behaviors of computer-use agents under benign inputs. Paper: https://arxiv.org/abs/2602.08235 · code: https://github.com/OSU-NLP-Group/AutoElicit
osunlp/AutoElicit-Bench
Viewer • Updated • 117 • 64 • 1Note AutoElicit benchmark for realistic unintended-behavior testing of computer-use agents. Paper: https://arxiv.org/abs/2602.08235
osunlp/AutoElicit-Exec
Viewer • Updated • 132 • 247 • 1Note Execution traces for AutoElicit computer-use safety evaluation. Paper: https://arxiv.org/abs/2602.08235
CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents
Paper • 2603.24440 • Published • 99Note CUA-Suite from ServiceNow provides continuous expert demonstrations, UI-Vision benchmark resources, and GroundCUA annotations for desktop computer-use agents. Paper: https://arxiv.org/abs/2603.24440 · dataset: https://huggingface.co/datasets/ServiceNow/VideoCUA · project: https://cua-suite.github.io
ServiceNow/VideoCUA
Updated • 1.72k • 33Note Official VideoCUA continuous human-demonstration dataset from CUA-Suite. Paper: https://arxiv.org/abs/2603.24440
MedCUA-Bench: A Screenshot-Only Benchmark for Clinical Computer-Use Agents
Paper • 2606.03203 • PublishedNote MedCUA-Bench evaluates clinical computer-use agents across 18 scenarios, 10 medical domains, deterministic completion checks, and five clinical-safety dimensions. Paper: https://arxiv.org/abs/2606.03203
RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS Environments
Paper • 2505.21936 • Published • 1Note RedTeamCUA and RTC-Bench evaluate indirect prompt injection in hybrid web-OS computer-use environments. Paper: https://arxiv.org/abs/2505.21936
WebPII: Benchmarking Visual PII Detection for Computer-Use Agents
Paper • 2603.17357 • PublishedNote WebPII benchmarks visual PII detection for computer-use agents with 44,865 annotated e-commerce UI images. Paper: https://arxiv.org/abs/2603.17357
A11y-CUA Dataset: Characterizing the Accessibility Gap in Computer Use Agents
Paper • 2602.09310 • PublishedNote A11y-CUA characterizes the accessibility gap between blind/low-vision users, sighted users, and computer-use agents across 60 everyday desktop tasks. Paper: https://arxiv.org/abs/2602.09310
berkeley-hci/A11y-CUA
Updated • 521 • 3Note Full multimodal A11y-CUA dataset with synchronized interaction traces, accessibility trees, screen video, and audio. Paper: https://arxiv.org/abs/2602.09310
berkeley-hci/Reduced-A11y-CUA
Updated • 164Note Structured-data-only A11y-CUA subset for accessibility-focused computer-use research. Paper: https://arxiv.org/abs/2602.09310
Capable but Careless: Do Computer-Use Agents Follow Contextual Integrity?
Paper • 2606.23189 • Published • 2Note AgentCIBench evaluates contextual-integrity and cross-application privacy failures in computer-use agents with executable, deterministic scenarios. Paper: https://arxiv.org/abs/2606.23189 · dataset: https://huggingface.co/datasets/UKPLab/agentcibench
UKPLab/agentcibench
Viewer • Updated • 203 • 66 • 2Note Official AgentCIBench dataset for contextual-integrity evaluation of computer-use agents. Paper: https://arxiv.org/abs/2606.23189
Yunhao-Feng/AgentHazard
Viewer • Updated • 2.65k • 128Note Official AgentHazard benchmark dataset with 2,653 execution-level safety instances. Paper: https://arxiv.org/abs/2604.02947
Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness
Paper • 2510.01670 • Published • 8Note BLIND-ACT evaluates blind goal-directedness in computer-use agents with 90 OSWorld-based tasks covering feasibility, safety, ambiguity, and context. Paper: https://arxiv.org/abs/2510.01670 · Microsoft Research publication: https://www.microsoft.com/en-us/research/publication/just-do-it-computer-use-agents-exhibit-blind-goal-directedness/
Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps
Paper • 2605.17554 • PublishedNote Paper: https://arxiv.org/abs/2605.17554 · public paper page: https://huggingface.co/papers/2605.17554
AgentProcessBench: Diagnosing Step-Level Process Quality in Tool-Using Agents
Paper • 2603.14465 • Published • 23Note Paper: https://arxiv.org/abs/2603.14465 · code: https://github.com/RUCBM/AgentProcessBench · project: https://rucbm.github.io/AgentProcessBench-Homepage/
LulaCola/AgentProcessBench
Viewer • Updated • 1k • 475 • 16Note Dataset: https://huggingface.co/datasets/LulaCola/AgentProcessBench · paper: https://arxiv.org/abs/2603.14465
AIM-Harvard/MedBrowseComp
Viewer • Updated • 1.14k • 154 • 8Note Dataset: https://huggingface.co/datasets/AIM-Harvard/MedBrowseComp · paper: https://arxiv.org/abs/2505.14963
Can Deep Research Agents Find and Organize? Evaluating the Synthesis Gap with Expert Taxonomies
Paper • 2601.12369 • Published • 4Note Paper: https://arxiv.org/abs/2601.12369 · code: https://github.com/KongLongGeFDU/TaxoBench
DeepResearch-9K: A Challenging Benchmark Dataset of Deep-Research Agent
Paper • 2603.01152 • PublishedNote Paper: https://arxiv.org/abs/2603.01152 · dataset: https://huggingface.co/datasets/artillerywu/DeepResearch-9K · code: https://github.com/Applied-Machine-Learning-Lab/DeepResearch-R1
artillerywu/DeepResearch-9K
Viewer • Updated • 13k • 400 • 18Note Dataset: https://huggingface.co/datasets/artillerywu/DeepResearch-9K · paper: https://arxiv.org/abs/2603.01152
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios
Paper • 2607.14989 • PublishedNote Paper: https://arxiv.org/abs/2607.14989 · OmniaBench evaluates general agents across diverse application scenarios.
EgoBench: An Interactive Egocentric Multimodal Benchmark for Tool-Using Agents
Paper • 2605.27820 • PublishedNote Paper: https://arxiv.org/abs/2605.27820 · EgoBench evaluates interactive egocentric multimodal tool-using agents.
Halluminate/WebBench
Viewer • Updated • 2.65k • 120 • 15Note Dataset: https://huggingface.co/datasets/Halluminate/WebBench · browser-agent benchmark with realistic web workflows.
wanlilll/WeaveBench
Updated • 900 • 7Note Dataset: https://huggingface.co/datasets/wanlilll/WeaveBench · paper: https://arxiv.org/abs/2606.09426 · project: https://weavebench.github.io/
Marti844/SaaS-Bench-docker
Updated • 523Note Dataset/container resources associated with SaaS-Bench; paper: https://arxiv.org/abs/2605.15777 · code: https://github.com/UniPat-AI/SaaS-Bench
Lakera/b3-agent-security-benchmark-weak
Viewer • Updated • 630 • 1.3k • 5Note Public dataset for agent security benchmark research: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak
obaydata/mcp-agent-trajectory-benchmark
Viewer • Updated • 38 • 762 • 3Note Public MCP agent trajectory benchmark dataset: https://huggingface.co/datasets/obaydata/mcp-agent-trajectory-benchmark
lyyang2766/Passing-the-Turing-Test-on-Screen-Agent-Humanization-Benchmark
Updated • 74Note Dataset for screen-agent humanization evaluation; related paper: https://arxiv.org/abs/2604.09574
mercor/apex-agents
Benchmark • Updated • 480 • 47.8k • 151Note APEX-Agents benchmark dataset for deep-research agents: https://huggingface.co/datasets/mercor/apex-agents
harborframework/terminal-bench-2.0
Benchmark • Updated • 11.4k • 47Note Terminal-Bench 2.0 benchmark dataset for terminal agents: https://huggingface.co/datasets/harborframework/terminal-bench-2.0
likaixin/ScreenSpot-Pro
Benchmark • Updated • 8.22k • 67Note ScreenSpot-Pro benchmark dataset for GUI grounding and computer-use agents: https://huggingface.co/datasets/likaixin/ScreenSpot-Pro
gaia-benchmark/GAIA
Viewer • Updated • 932 • 16.9k • 735Note GAIA benchmark dataset for general AI assistants: https://huggingface.co/datasets/gaia-benchmark/GAIA · benchmark paper: https://arxiv.org/abs/2311.12983
openai/BrowseCompLongContext
Viewer • Updated • 295 • 1.98k • 54Note OpenAI BrowseComp long-context dataset: https://huggingface.co/datasets/openai/BrowseCompLongContext · benchmark context: https://openai.com/index/browsecomp/
smolagents/browse_comp
Viewer • Updated • 1.27k • 3.1k • 6Note BrowseComp dataset mirror for browsing-agent evaluation: https://huggingface.co/datasets/smolagents/browse_comp · benchmark context: https://openai.com/index/browsecomp/
mercor/APEX-v1-extended
Benchmark • Updated • 100 • 8.32k • 16Note APEX-v1 extended benchmark dataset for deep-research agents: https://huggingface.co/datasets/mercor/APEX-v1-extended
ServiceNow/WorkArena-Instances
Updated • 16.6k • 5Note ServiceNow WorkArena benchmark instances for web agents: https://huggingface.co/datasets/ServiceNow/WorkArena-Instances · paper: https://arxiv.org/abs/2403.07718
McGill-NLP/WebLINX-full
Updated • 27.9k • 8Note WebLINX real-world conversational web navigation benchmark: https://huggingface.co/datasets/McGill-NLP/WebLINX-full · paper: https://arxiv.org/abs/2402.05930 · code: https://github.com/mcgill-nlp/weblinx
osunlp/Multimodal-Mind2Web
Viewer • Updated • 14.2k • 8.76k • 96Note Multimodal Mind2Web dataset for web-agent evaluation: https://huggingface.co/datasets/osunlp/Multimodal-Mind2Web · Mind2Web paper: https://arxiv.org/abs/2306.06070
futurehouse/lab-bench
Viewer • Updated • 1.97k • 9.85k • 50Note LAB-Bench scientific research agent evaluation dataset: https://huggingface.co/datasets/futurehouse/lab-bench · benchmark paper: https://arxiv.org/abs/2407.10362
EdisonScientific/labbench2
Viewer • Updated • 3.82k • 5.26k • 55Note LABBench2 dataset for biology research agent evaluation: https://huggingface.co/datasets/EdisonScientific/labbench2 · paper: https://arxiv.org/abs/2604.09554
facebook/airs-bench
Viewer • Updated • 20 • 56 • 6Note Meta FAIR AIRS-Bench dataset for AI research science agents: https://huggingface.co/datasets/facebook/airs-bench · paper: https://arxiv.org/abs/2602.06855 · code: https://github.com/facebookresearch/airs-bench
ai-safety-institute/AgentHarm
Viewer • Updated • 468 • 3.86k • 58Note AgentHarm harmfulness benchmark for LLM agents: https://huggingface.co/datasets/ai-safety-institute/AgentHarm · paper: https://arxiv.org/abs/2410.09024
eth-sri/agentbench
Viewer • Updated • 138 • 316 • 1Note ETH SRI AgentBench dataset for agentic software-engineering workflows: https://huggingface.co/datasets/eth-sri/agentbench · code: https://github.com/eth-sri/agentbench
Mininglamp-2718/WebRetriever
Preview • Updated • 80 • 3Note WebRetriever large-scale web-agent evaluation dataset: https://huggingface.co/datasets/Mininglamp-2718/WebRetriever · paper: https://arxiv.org/abs/2607.06118
microsoft/WebTailBench
Preview • Updated • 330 • 17Note Microsoft WebTailBench computer-use benchmark: https://huggingface.co/datasets/microsoft/WebTailBench · paper: https://arxiv.org/abs/2511.19663 · code: https://github.com/microsoft/fara
xlangai/computer-agent-arena
Viewer • Updated • 114k • 269 • 1Note Computer Agent Arena interaction-trajectory dataset: https://huggingface.co/datasets/xlangai/computer-agent-arena · project: https://arena.xlang.ai/
HuggingSelf/StressWeb
Updated • 91 • 2Note StressWeb web-agent robustness benchmark: https://huggingface.co/datasets/HuggingSelf/StressWeb · paper: https://arxiv.org/abs/2604.16385
xlangai/AgentNet
Preview • Updated • 2.12k • 89Note Large-scale desktop computer-use agent trajectory dataset: https://huggingface.co/datasets/xlangai/AgentNet
ST-WebAgentBench/st-webagentbench
Viewer • Updated • 3.06k • 308 • 5Note Safety and trustworthiness benchmark for web agents: https://huggingface.co/datasets/ST-WebAgentBench/st-webagentbench · paper: https://arxiv.org/abs/2410.06703
DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation
Paper • 2512.17776 • Published • 1Note DEER benchmark for deep-research expert reports: https://arxiv.org/abs/2512.17776 · LG AI Research
LG-AI-Research/DEER-Deep-Research-Benchmark
Updated • 93Note DEER deep-research benchmark dataset: https://huggingface.co/datasets/LG-AI-Research/DEER-Deep-Research-Benchmark · paper: https://arxiv.org/abs/2512.17776
local-deep-research/ldr-benchmarks
Viewer • Updated • 24 • 167 • 11Note Community benchmark results for local deep-research agents: https://huggingface.co/datasets/local-deep-research/ldr-benchmarks · code: https://github.com/LearningCircuit/ldr-benchmarks
muset-ai/DeepResearch-Bench-Dataset
Preview • Updated • 516 • 10Note Deep research agent benchmark dataset: https://huggingface.co/datasets/muset-ai/DeepResearch-Bench-Dataset · paper: https://arxiv.org/abs/2506.11763
muset-ai/DeepResearch-Bench-II-Dataset
Viewer • Updated • 135 • 5.56k • 2Note Deep research agent benchmark dataset: https://huggingface.co/datasets/muset-ai/DeepResearch-Bench-II-Dataset · paper: https://arxiv.org/abs/2601.08536
ParseBench: A Document Parsing Benchmark for AI Agents
Paper • 2604.08538 • Published • 10Note ParseBench document-parsing benchmark for AI agents: https://arxiv.org/abs/2604.08538 · code: https://github.com/run-llama/ParseBench
llamaindex/ParseBench
Benchmark • Updated • 169k • 11.1k • 112Note ParseBench dataset for agent-facing document parsing: https://huggingface.co/datasets/llamaindex/ParseBench · paper: https://arxiv.org/abs/2604.08538 · code: https://github.com/run-llama/ParseBench
xlangai/ubuntu_osworld_verified_trajs
Updated • 3.62k • 19Note OSWorld-Verified computer-use trajectories: https://huggingface.co/datasets/xlangai/ubuntu_osworld_verified_trajs · project: https://xlang.ai/blog/osworld-verified
cua-lite/Lite.CUAWorld
Viewer • Updated • 3.63k • 231 • 1Note Lite CUA-World computer-use benchmark assets: https://huggingface.co/datasets/cua-lite/Lite.CUAWorld
hud-evals/OSWorld-Verified
Viewer • Updated • 369 • 287Note OSWorld-Verified evaluation dataset for computer-use agents: https://huggingface.co/datasets/hud-evals/OSWorld-Verified
hasura/agentic-data-access-benchmark
Viewer • Updated • 144 • 22Note Hasura Agentic Data Access Benchmark for closed-domain, multi-step data retrieval: https://huggingface.co/datasets/hasura/agentic-data-access-benchmark · code: https://github.com/hasura/agentic-data-access-benchmark
google/deepsearchqa
Viewer • Updated • 900 • 40.5k • 127Note Google DeepSearchQA benchmark for difficult multi-step information seeking: https://huggingface.co/datasets/google/deepsearchqa
Salesforce/LiveResearchBench
Viewer • Updated • 623 • 916 • 7Note Salesforce LiveResearchBench for user-centric deep research in the wild: https://huggingface.co/datasets/Salesforce/LiveResearchBench · paper: https://arxiv.org/abs/2510.14240 · project: https://livedeepresearch.github.io/
Lk123/AutoResearchBench
Updated • 165 • 5Note AutoResearchBench scientific literature discovery benchmark: https://huggingface.co/datasets/Lk123/AutoResearchBench · code: https://github.com/CherYou/AutoResearchBench
cx-cmu/AgentWebBench-corpus
Viewer • Updated • 1 • 238Note AgentWebBench corpus for multi-agent coordination in agentic web: https://huggingface.co/datasets/cx-cmu/AgentWebBench-corpus · paper: https://arxiv.org/abs/2604.10938
yonsei-dli/Persona2Web
Viewer • Updated • 650 • 195Note Persona2Web personalized web-agent benchmark dataset: https://huggingface.co/datasets/yonsei-dli/Persona2Web · paper: https://arxiv.org/abs/2602.17003
sunblaze-ucb/AgentSynth
Viewer • Updated • 1.21k • 147 • 6Note AgentSynth scalable synthesis pipeline for generalist computer-use agent tasks and trajectories: https://huggingface.co/datasets/sunblaze-ucb/AgentSynth
markov-ai/computer-use
Viewer • Updated • 313 • 456 • 69Note Successful OSWorld computer-use agent trajectories: https://huggingface.co/datasets/markov-ai/computer-use
agents-last-exam/agents-last-exam
Viewer • Updated • 153 • 932 • 206Note Agents Last Exam task-card metadata: https://huggingface.co/datasets/agents-last-exam/agents-last-exam · source: https://github.com/rdi-berkeley/agents-last-exam · site: https://agents-last-exam.org
agents-last-exam/agents-last-exam-data
Updated • 11.8k • 5Note Agents Last Exam open task-input data: https://huggingface.co/datasets/agents-last-exam/agents-last-exam-data · source: https://github.com/rdi-berkeley/agents-last-exam
AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation
Paper • 2604.18240 • Published • 17Note AJ-Bench evaluates agent-as-a-judge for environment-aware evaluation: https://arxiv.org/abs/2604.18240
Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History
Paper • 2602.17003 • PublishedNote Persona2Web personalized open-web agent benchmark: https://arxiv.org/abs/2602.17003
microsoft/synthetic-computers-at-scale
Viewer • Updated • 98 • 290 • 19Note Microsoft synthetic computer-use trajectories dataset: https://huggingface.co/datasets/microsoft/synthetic-computers-at-scale
AI45Research/SATraj-OS
Viewer • Updated • 6.03k • 1.02k • 3Note SATraj-OS computer-use trajectory dataset: https://huggingface.co/datasets/AI45Research/SATraj-OS
anon-aiba-2026/aiba-benchmark
Viewer • Updated • 132 • 8Note AIBA benchmark for computer-use/browser-agent atomic-ability and anti-cheat evaluation: https://huggingface.co/datasets/anon-aiba-2026/aiba-benchmark
goldenash/MAD-Bench
Viewer • Updated • 360 • 16Note MAD-Bench for deceptive behavior in multimodal computer-use agents: https://huggingface.co/datasets/goldenash/MAD-Bench
chakra-labs/dojo-bench-mini
Viewer • Updated • 219 • 160Note Dojo-Bench-Mini public computer-use benchmark: https://huggingface.co/datasets/chakra-labs/dojo-bench-mini
claw-eval/Claw-Eval
Benchmark • Updated • 3.38k • 31Note Claw-Eval benchmark dataset for evaluating agentic systems: https://huggingface.co/datasets/claw-eval/Claw-Eval
internlm/WildClawBench
Benchmark • Updated • 33.7k • 63Note WildClawBench real-world long-horizon agent evaluation benchmark: https://huggingface.co/datasets/internlm/WildClawBench
mercor/ACE
Benchmark • Updated • 592 • 624 • 5Note Mercor ACE benchmark for long-horizon consumer application agent tasks: https://huggingface.co/datasets/mercor/ACE
cua-lite/Lite.OSWorld
Viewer • Updated • 4.82k • 1.6k • 1Note Lite OSWorld computer-use benchmark tasks: https://huggingface.co/datasets/cua-lite/Lite.OSWorld
xlangai/windows_osworld
Updated • 316 • 4Note Windows OSWorld computer-use task dataset: https://huggingface.co/datasets/xlangai/windows_osworld
xlangai/macos_osworld
Updated • 33 • 1Note macOS OSWorld computer-use task dataset: https://huggingface.co/datasets/xlangai/macos_osworld
xlangai/ubuntu_osworld
Updated • 6.17k • 5Note Ubuntu OSWorld computer-use task dataset: https://huggingface.co/datasets/xlangai/ubuntu_osworld
MMInstruction/OSWorld-G
Viewer • Updated • 510 • 537 • 6Note OSWorld-G GUI grounding dataset associated with computer-use evaluation: https://huggingface.co/datasets/MMInstruction/OSWorld-G
TESS-Computer/tess-agentnet
Viewer • Updated • 329k • 41Note TESS AgentNet computer-use trajectories for VLA/GUI-agent evaluation: https://huggingface.co/datasets/TESS-Computer/tess-agentnet · source AgentNet: https://huggingface.co/datasets/xlangai/AgentNet
VPI-Bench/vpi-bench
Viewer • Updated • 306 • 286 • 1Note VPI-Bench visual prompt-injection benchmark for computer-use agents: https://huggingface.co/datasets/VPI-Bench/vpi-bench · paper: https://arxiv.org/abs/2506.02456
Hcompany/DragOn
Updated • 203 • 5Note DragOn benchmark and dataset for drag-based GUI interactions: https://huggingface.co/datasets/Hcompany/DragOn
ai-multiple/aim-ui-grounding
Viewer • Updated • 10 • 40 • 3Note UI grounding benchmark preview for computer-use agents: https://huggingface.co/datasets/ai-multiple/aim-ui-grounding · methodology: https://research.aimultiple.com/computer-use-agents/
OpenGVLab/MMBench-GUI
Preview • Updated • 213 • 38Note MMBench-GUI dataset: https://huggingface.co/datasets/OpenGVLab/MMBench-GUI · paper: https://arxiv.org/abs/2507.19478 · code: https://github.com/open-compass/MMBench-GUI
Aoraku/VG-GUI-Bench
Updated • 1.45k • 2Note Video-Guided GUI Agent benchmark dataset: https://huggingface.co/datasets/Aoraku/VG-GUI-Bench · code: https://github.com/VG-GUI-TASKER/VG-GUI-TASKER
vyokky/GUI-360
Preview • Updated • 12.1k • 18Note GUI-360 comprehensive computer-use agent dataset and benchmark: https://huggingface.co/datasets/vyokky/GUI-360 · paper: https://arxiv.org/abs/2511.04307
iMeanAI/GAE-Bench-lite
Viewer • Updated • 1.81M • 436Note GUI Agents Embedding Benchmark Lite: https://huggingface.co/datasets/iMeanAI/GAE-Bench-lite
lgy0404/LearnGUI
Updated • 438 • 8Note LearnGUI mobile GUI-agent demonstration benchmark: https://huggingface.co/datasets/lgy0404/LearnGUI · paper: https://arxiv.org/abs/2504.13805
TIGER-Lab/BrowserAgent-Data
Viewer • Updated • 44.4k • 260 • 5Note Official TIGER-Lab browser-agent trajectories and data, directly relevant to real-world web-agent evaluation.
TIGER-Lab/BrowserAgent-SeedData
Viewer • Updated • 252k • 130Note Official TIGER-Lab seed tasks/data for browser-agent evaluation and training.
Salesforce/CRMArena
Viewer • Updated • 1.19k • 770 • 8Note Salesforce benchmark data for evaluating agents on CRM workflows.
OpenHandsCommunity/eval-output-webarena
Updated • 72Note OpenHands community WebArena evaluation outputs for reproducible web-agent comparisons.
DataCanvasAILab/Titan-CV-Agent-Benchmark
Viewer • Updated • 93 • 226 • 2Note Computer-vision agent benchmark resources for evaluating visual interaction capabilities.
webarena-x/webarena-infinity-trajectories
Viewer • Updated • 23.6k • 831 • 2Note WebArena-X trajectories for long-horizon browser-agent evaluation.
parcs-benchmark/parcs-agent-benchmark
Viewer • Updated • 70k • 36Note PARCS agent benchmark resources for evaluating agent task completion.
MemGym/memgym-rm-scenario-ood-webarena
Viewer • Updated • 426 • 38Note Out-of-distribution WebArena scenarios for memory-aware web-agent evaluation.
Intelligent-Internet/ii-agent_gaia-benchmark_validation
Viewer • Updated • 165 • 855 • 8Note GAIA benchmark validation resources for general assistant evaluation.
DeepNLP/ai-agent-benchmark
Viewer • Updated • 33 • 16 • 1Note Agent benchmark dataset covering task-oriented AI-agent evaluation.
LangAGI-Lab/mini_rm_benchmark_for_web_agent
Viewer • Updated • 128 • 32Note Web-agent reward-model benchmark for evaluating preference and outcome scoring.
design-agent/in2n-benchmark-bundle
Viewer • Updated • 25.9k • 131Note Benchmark bundle for evaluating interactive design agents.
ScareRezume/agent-sandbox-negotiation-benchmark
Viewer • Updated • 24.1k • 14 • 1Note Sandbox negotiation benchmark for tool-using agent evaluation.
clicksprotocol/agent-treasury-benchmark
Preview • Updated • 26Note Agent treasury benchmark for evaluating tool use and financial workflows.
Agnuxo/benchclaw-ai-agent-benchmarking
Updated • 15Note OpenClaw-oriented agent benchmarking resources for related autonomous-agent evaluation.
webarenapro/task-intents-and-rubrics
Viewer • Updated • 300 • 14Note WebArena task intents and rubrics for fine-grained browser-agent evaluation.
webarenapro/agent_trajectory_reviews
Viewer • Updated • 1 • 35Note Reviewed WebArena agent trajectories for analysis of browser-agent behavior.
AutoSurfer/WebArenaSFT_V4_Refined
Viewer • Updated • 12.7k • 21Note Refined WebArena task/trajectory data for web-agent evaluation and training.
WPRM/annotated_webarena_checklist
Viewer • Updated • 812 • 11Note Annotated WebArena checklist resources for reproducible browser-agent evaluation.
DLIlab/webarena_two_shot_instructions
Viewer • Updated • 130 • 32Note WebArena instruction variants for controlled web-agent evaluation.
sudac/browser-world-models-webarena-reddit
Viewer • Updated • 726 • 42Note WebArena Reddit scenarios used in browser-world-model and agent evaluation.
PRHW/loom-benchmark-webarena
Updated • 1.67kNote WebArena-derived benchmark resources for long-horizon web agents.
vals-ai/finance_agent_benchmark
Viewer • Updated • 50 • 948 • 8Note Finance-agent benchmark dataset for evaluating agents in structured real-world workflows.
rogue-security/coding-agent-security-benchmark
Viewer • Updated • 332 • 2Note Security benchmark resources for evaluating coding-agent behavior.
Agent-Threat-Rule/atr-skill-benchmark
Viewer • Updated • 498 • 54Note Agent skill and threat benchmark for evaluating tool-using agents.
MSakae/industrial-agent-benchmark
Viewer • Updated • 180 • 36 • 1Note Industrial-agent benchmark resources for task-oriented agent evaluation.