Agent Harness Benchmarks & Evals
Curated benchmarks and papers for evaluating agent harnesses in realistic web, computer, and tool-using environments.
Paper • 2604.08523 • Published • 265Note ClawBench: real-world browser-agent tasks with safe submission interception and detailed execution traces. Paper: https://arxiv.org/abs/2604.08523 · code: https://github.com/TIGER-AI-Lab/ClawBench · project: https://claw-bench.com/
StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs
Paper • 2505.20139 • Published • 19Note StructEval: structured-output generation benchmark with executable/rendered artifact evaluation. Paper: https://arxiv.org/abs/2505.20139 · code: https://github.com/TIGER-AI-Lab/StructEval
Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows
Paper • 2605.27922 • PublishedNote Harness-Bench: diagnostic evaluation of model-harness pairings across realistic workflows.
AgentBench: Evaluating LLMs as Agents
Paper • 2308.03688 • Published • 26Note AgentBench: multi-environment evaluation of LLMs as agents.
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Paper • 2404.07972 • Published • 52Note OSWorld: open-ended computer-use tasks in real computer environments.
WebArena: A Realistic Web Environment for Building Autonomous Agents
Paper • 2307.13854 • Published • 27Note WebArena: realistic web tasks for browser agents.
ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
Paper • 2604.05172 • Published • 25Note ClawsBench: productivity-agent evaluation in simulated workspaces with completion and safety metrics.
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
Paper • 2605.10912 • Published • 46Note WildClawBench: long-horizon real-world agent evaluation benchmark.
harborframework/terminal-bench-2.0
Benchmark • Updated • 10.7k • 47Note Terminal-Bench 2.0 task data for reproducible coding-agent evaluation in containerized environments.
claw-eval/Claw-Eval
Benchmark • Updated • 3.31k • 31Note Claw-Eval task suite for evaluating agent completion, safety, and robustness.