Buckets:
| title: Agentic Benchmarks (Execution-Graded Environments for RL'd Agents) | |
| maturity: developing | |
| sources: | |
| - arxiv:2310.06770 | |
| - arxiv:2307.13854 | |
| - arxiv:2406.12045 | |
| - arxiv:2308.03688 | |
| - arxiv:2210.03629 | |
| - arxiv:2201.11903 | |
| open_questions: | |
| - "Agentic benchmarks double as verifiable-reward *training environments* — so how much of the rapid score climb (e.g. SWE-bench 1.96% → the 60–70% frontier reported since) is genuine capability vs optimization *to the benchmark's checkable signal*, including test/state-match gaming? The grading is verifiable but `r=1` is necessary-not-sufficient (τ-bench), so the eval and the training target share a reward-hacking surface." | |
| - "pass^1 (average success) hides that agents are wildly inconsistent — τ-bench's pass^8 falls below 25% where pass^1 is ~61%. Should RL-for-agents optimize *reliability* (pass^k) rather than mean reward, and does mean-reward RL leave the consistency gap untouched?" | |
| - "Scores are heavily scaffold- and version-dependent (function-calling vs ReAct vs Act; retrieval vs oracle context; the LM user-simulator's own quality). A bare 'benchmark number' is close to meaningless — how should the corpus report agentic results so they stay comparable as scaffolds churn?" | |
| - "How well does the self-hosted / continually-updatable design (WebArena Docker, SWE-bench post-cutoff issues) actually defeat contamination and live-site drift versus static benchmarks — and does it hold as these benchmarks themselves become saturated training targets?" | |
| # Agentic Benchmarks (Execution-Graded Environments for RL'd Agents) | |
| Most evaluation in this wiki grades a *single response* — a preference win-rate | |
| (`evaluation/alignment-and-winrate-evals`) or a static answer key | |
| (`evaluation/capability-and-safety-benchmarks`). **Agentic benchmarks** grade something | |
| harder and, for reinforcement learning (RL), more consequential: an autonomous **agent** | |
| that takes **many actions over a long horizon** inside an interactive environment — a | |
| code repository, a live-like website, a database behind tool APIs, a simulated user — and | |
| is scored by whether its actions *achieved the goal*, checked **programmatically by | |
| executing them**, not by matching a reference string. This article is the deep-dive child | |
| of `evaluation/capability-and-safety-benchmarks`; its thesis is that these benchmarks are | |
| **not merely evals but verifiable-reward *environments*** — the same execution-based, | |
| ground-truth signal that Reinforcement Learning from Verifiable Rewards (RLVR, | |
| `verifiable-rewards-and-reasoning/rlvr-overview`) optimizes — so the frontier eval and the | |
| frontier training target have become **the same object**, which is both why they matter | |
| and why they must be read carefully. | |
| ## 1. What makes a benchmark "agentic" | |
| Four properties recur across the canonical suites, and together they separate agentic | |
| benchmarks from the static kind: | |
| - **Multi-turn, long-horizon interaction.** The agent observes, acts, observes the | |
| environment's response, and repeats — for a handful to dozens of rounds. AgentBench and | |
| τ-bench both formalize the setting as a **partially-observable Markov decision process | |
| (POMDP)** `⟨S, A, T, R, U, O⟩` [source:arxiv:2308.03688][source:arxiv:2406.12045]; | |
| round counts range from ~5 (database queries) to ~35 (embodied house-holding) in | |
| AgentBench alone. | |
| - **A real (or realistic) environment with tools.** A Docker Ubuntu shell and a live MySQL | |
| database (AgentBench's Operating System and Database tasks [source:arxiv:2308.03688]); | |
| four self-hosted websites plus a map/calculator/scratchpad and a Wikipedia (WebArena | |
| [source:arxiv:2307.13854]); a full Python repository at a real pre-fix commit | |
| (SWE-bench [source:arxiv:2310.06770]); JSON databases mutated only through Python API | |
| tools under a written policy (τ-bench [source:arxiv:2406.12045]). | |
| - **Execution-based, functional-correctness grading** (§2) — the defining feature. | |
| - **A prompting/scaffold assumption.** These evaluate an LLM *as* an agent, usually | |
| zero-shot via **chain-of-thought (CoT)** [source:arxiv:2201.11903] in a **Thought + | |
| Action** loop adapted from **ReAct** [source:arxiv:2210.03629]. AgentBench deliberately | |
| uses "the easiest, cheapest, most common" single-trial CoT — no self-consistency, no | |
| tree search — to reflect how people actually deploy models [source:arxiv:2308.03688]; | |
| τ-bench instead compares native **function calling (FC)** against text-ReAct and an | |
| Act-only ablation, finding FC consistently strongest [source:arxiv:2406.12045]. This | |
| makes every score **scaffold-dependent** (§4). | |
| ## 2. The grading mechanism *is* a verifiable reward (the RL-central point) | |
| The load-bearing commonality: success is a **programmatic function of the world's end | |
| state**, computed by running the agent's actions — not a similarity to a gold trajectory. | |
| This is exactly a **verifiable reward** (`reward-modeling/verifiable-rewards`), which is | |
| why each benchmark doubles as an RL environment: | |
| - **SWE-bench — hidden test suite.** The agent emits a patch; the harness applies it with | |
| unix `patch` and runs the repo's tests. Resolution requires the patch to apply *and* | |
| both the **FAIL_TO_PASS** tests (which verify the fix) and a median ~51 **PASS_TO_PASS** | |
| tests (which verify nothing else broke) to pass [source:arxiv:2310.06770]. The model may | |
| solve the issue *differently* from the reference PR — grading is execution, not | |
| text-match. | |
| - **WebArena — functional correctness.** A reward function `r(a, s)` over the action/state | |
| sequence checks the achieved end-state, in two families (its Table 1): `r_info` | |
| (`exact_match` / `must_include` / an LLM `fuzzy_match` for information-seeking answers) | |
| and `r_prog` (per-task **locators** — a database query, a site API call, or a JavaScript | |
| DOM selection — assert e.g. that an order was really placed) [source:arxiv:2307.13854]. | |
| The paper's central argument is that end-state grading is **more reliable than comparing | |
| action sequences** because it admits *multiple valid paths* to the goal. | |
| - **τ-bench — state match × required output.** Binary reward `r = r_action × r_output`: | |
| `r_action = 1` iff the **final database state exactly matches the unique annotated goal**, | |
| and `r_output = 1` iff the replies contained all required information (e.g. a quoted | |
| refund amount) [source:arxiv:2406.12045]. Tasks are annotated (and validated with >40 | |
| GPT-4-turbo trials each) so exactly one database outcome is correct, letting the noisy | |
| conversation vary while grading stays objective. | |
| - **AgentBench — task-specific success.** Per-environment metrics — success rate, | |
| answer F1, win rate, game progress, step success rate — combined into a weighted | |
| **Overall Score** whose per-task weight is the *reciprocal of the average score across | |
| all tested models*, so easy high-scoring tasks don't dominate (its Table 2) | |
| [source:arxiv:2308.03688]. | |
| Because each signal is ground-truth and automatically checkable, it is **directly usable | |
| as an RL reward** — no learned reward model, no human label — placing agentic benchmarks | |
| in the same family as SWE-bench's and τ-bench's own framing as verifiable-reward targets | |
| that "RL-for-agents optimizes toward" [source:arxiv:2310.06770][source:arxiv:2406.12045]. | |
| SWE-bench even ships a **training split** (SWE-bench-train, ~19k issue-PR pairs from 37 | |
| *disjoint* repos) and fine-tuned **SWE-Llama** models, making the eval/train duality | |
| explicit [source:arxiv:2310.06770]. | |
| ## 3. The four canonical suites | |
| ### 3.1 SWE-bench — repository-scale code (2,294 tasks) | |
| Real GitHub issue→pull-request tasks mined by a **three-stage pipeline** (its Figure 2) | |
| over ~90,000 PRs from 12 popular Python repos: scrape PRs → keep merged PRs that resolve a | |
| linked issue *and* touch test files → an **execution filter** that keeps an instance only | |
| if ≥1 test flips fail→pass after the non-test changes [source:arxiv:2310.06770]. Each task | |
| hands the model an issue (avg 195 words) and the **entire repo** at the base commit | |
| (mean ~438K lines across ~3,000 files), so the model must **localize a few lines in a sea | |
| of context**, edit across files (gold patches touch avg 1.7 files / 3.0 functions / 32.8 | |
| lines), and respect existing style. Context is supplied by realistic **BM25 sparse | |
| retrieval** or an **"oracle"** upper-bound that reveals the gold-edited files. It was | |
| brutal at release — the best system (Claude 2 + BM25) resolved **1.96%** — and the | |
| pipeline is **continually updatable** on post-cutoff issues, a deliberate contamination | |
| defense (§4). | |
| ### 3.2 WebArena — realistic web/tool use (812 tasks) | |
| A **self-hostable, reproducible** environment of four fully-functional site categories — | |
| e-commerce, a Reddit-like forum, GitLab, and a content-management system — populated with | |
| data sampled from their real counterparts, plus utility tools and knowledge resources, | |
| delivered as **Docker containers with `gym`-style APIs and deterministic resets** | |
| [source:arxiv:2307.13854]. Running offline sidesteps CAPTCHAs and live-site drift that | |
| make cross-system comparison unfair over time. The observation space renders as raw | |
| **HTML/DOM**, a **screenshot**, or the compact **accessibility tree**; it is the first web | |
| environment to support **multi-tab** tasks. The 812 tasks (from 241 templates) span | |
| information-seeking, navigation, and content/configuration, and some are deliberately | |
| **unachievable** ("N/A") to test whether an agent *refuses* rather than hallucinates. | |
| Headline gap: the best agent (GPT-4 + CoT) reaches **14.41%** end-to-end success versus | |
| **78.24%** for humans [source:arxiv:2307.13854]. | |
| ### 3.3 τ-bench — tool + policy + simulated user (165 tasks, 2 domains) | |
| The most "customer-service-realistic" setup: the agent must **call domain API tools, obey | |
| a written Markdown policy document, and converse with an LM-simulated user** (GPT-4-0613 | |
| holding a hidden task instruction) over up to 30 actions, across τ-retail (115 tasks) and | |
| τ-airline (50 tasks) [source:arxiv:2406.12045]. Crucially, **many policy restrictions are | |
| not API-enforced** — the agent must follow them on its own — and the goal annotation is | |
| hidden from the agent. Its signature contributions are two: (1) the **pass^k** reliability | |
| metric — the chance that *all* k i.i.d. trials of the same task succeed — which exposes | |
| that agents are wildly inconsistent (GPT-4o's success falls from pass^1 ≈ 61% to **pass^8 | |
| < 25%** on retail, its Figure 4); and (2) a concrete demonstration that state-match reward | |
| is **reward-hackable** — `r = 1` is "necessary but not sufficient," since an agent can | |
| reach the goal database state while *violating policy* (e.g. issuing a return without the | |
| required confirmation) [source:arxiv:2406.12045]. A policy-ablation is telling: removing | |
| the policy barely hurt GPT-4o on simple retail (−4.4%) but dropped it **22.4%** on | |
| airline, i.e. much retail "success" was commonsense tool use, not rule-following. | |
| ### 3.4 AgentBench — breadth across 8 environments | |
| Evaluates the LLM-as-agent across **8 interactive environments** in three groups — | |
| code-grounded (Operating System, Database, Knowledge Graph over Freebase), game-grounded | |
| (a card game, lateral-thinking puzzles, ALFWorld house-holding), and web-grounded | |
| (WebShop, Mind2Web browsing) — zero-shot with single-trial CoT [source:arxiv:2308.03688]. | |
| Across 29 models it found a **large closed-vs-open gap**: GPT-4 scored **4.01** overall vs | |
| the best ≤70B open model (CodeLlama-34B) at **0.96**, attributing open-model failure to | |
| weak long-horizon reasoning, decision-making, and instruction-following, with a | |
| 5-way failure taxonomy (context-limit / invalid-format / invalid-action / | |
| task-limit-exceeded / complete). It is the **breadth** complement to the three | |
| depth-in-one-domain benchmarks above. | |
| ## 4. Cross-cutting themes (why these evals behave differently) | |
| - **Huge human–agent gaps = RL headroom.** WebArena's 14% vs 78% human | |
| [source:arxiv:2307.13854], SWE-bench's 1.96% at release [source:arxiv:2310.06770], and | |
| AgentBench's open-model 0.96 [source:arxiv:2308.03688] are exactly the kind of large, | |
| checkable gaps that make an area attractive for RL — a dense-enough verifiable signal | |
| with a long way to climb. | |
| - **Reliability (pass^k) is a distinct objective from mean reward.** τ-bench's pass^1→pass^8 | |
| collapse names a target that average-reward training can leave untouched: an agent that | |
| succeeds *on average* may still fail *most repetitions* of the same task | |
| [source:arxiv:2406.12045]. This reframes the RL goal as **consistency/robustness**, not | |
| just expected return (`objectives-and-regularization/entropy-and-exploration` for the | |
| exploration/consistency tension). | |
| - **The reward is verifiable but *gameable* — the eval inherits reward hacking.** Because | |
| the training target and the eval are the same execution signal, the reward-hacking | |
| surface is shared: SWE-bench patches can pass the visible tests without being a correct | |
| fix; WebArena admits multiple paths (a feature, but also un-checked side effects); and | |
| τ-bench explicitly shows goal-state-reached-while-policy-violated | |
| [source:arxiv:2406.12045]. This is the **outcome-vs-process** tension | |
| (`reward-modeling/process-vs-outcome-rewards`) and a live reward-hacking caveat | |
| (`reward-modeling/reward-hacking`) baked into agentic evaluation. | |
| - **Contamination resistance by construction.** SWE-bench can be re-mined on issues created | |
| *after* a model's cutoff, and WebArena is self-hosted and deterministic — both directly | |
| target the contamination/saturation problem that dogs static benchmarks | |
| (`evaluation/capability-and-safety-benchmarks` §3). Whether this holds once the benchmark | |
| becomes a saturated training target is open (frontmatter). | |
| - **Scaffold- and version-dependence.** τ-bench's function-calling-vs-ReAct gap | |
| [source:arxiv:2406.12045][source:arxiv:2210.03629], SWE-bench's retrieval-vs-oracle gap | |
| [source:arxiv:2310.06770], and the LM user-simulator's own quality mean a bare score is | |
| uninterpretable without specifying model version + scaffold + context construction. | |
| Report agentic results *with their harness*, not as a single number. | |
| ## 5. Relationship to the rest of the wiki | |
| - **`verifiable-rewards-and-reasoning/rlvr-overview`** — agentic benchmarks are the | |
| multi-turn, tool-using extension of RLVR's checkable-reward idea beyond single-answer | |
| math/code. | |
| - **`reward-modeling/verifiable-rewards`** — the reward-family these environments belong | |
| to; **`reward-modeling/process-vs-outcome-rewards`** and **`reward-modeling/reward-hacking`** | |
| — the outcome-grading caveat they concretely exhibit. | |
| - **`evaluation/capability-and-safety-benchmarks`** — the hub: static capability/safety | |
| suites vs these interactive, execution-graded ones (this is the deep child). | |
| - **`evaluation/judging-bias-and-contamination`** — WebArena's `fuzzy_match` and τ-bench's | |
| user-sim are LLM-in-the-loop graders, inheriting judge-reliability concerns. | |
| - **`evaluation/alignment-and-winrate-evals`** — the preference-grading contrast: | |
| functional correctness replaces (gameable) human/LLM preference with (differently-gameable) | |
| execution. | |
| ## 6. Current status and trajectory | |
| *(Hedged; grounded in the processed corpus, which captures the four release papers, not | |
| the fast-moving leaderboard since.)* | |
| On the corpus evidence, agentic benchmarks have become **the frontier evaluation for | |
| capable models** precisely because static benchmarks saturate and these do not — they are | |
| long-horizon, execution-graded, and (SWE-bench, WebArena) contamination-resistant by | |
| design. Their defining move is grading by **verifiable, programmatic end-state**, which is | |
| why they are simultaneously the **training environments** RL-for-agents and RLVR optimize — | |
| the eval and the target have merged. Two cautions are load-bearing and likely durable: the | |
| grading is **reward-hackable** (`r=1` necessary-not-sufficient), so a rising score is not | |
| self-evidently rising capability; and results are **reliability-sensitive** (pass^k) and | |
| **scaffold-dependent**, so single numbers mislead. All specific figures here (SWE-bench | |
| 1.96%, WebArena 14.41% vs 78.24%, τ-bench pass^1 ≈ 61%/35%, AgentBench 4.01 vs 0.96) are | |
| **release-era snapshots of specific model versions** and have moved substantially since; | |
| cite them with date and scaffold, and treat the *methodology* (execution grading, pass^k, | |
| self-hosting, continual updating) as the durable contribution. This node covers the core | |
| four; the space is broad and growing (verified/curated SWE-bench variants, OS/desktop and | |
| longer-horizon environments) — `not-reported ≠ not-exist`, so this is an expandable hub, | |
| not a closed list. | |
| ## 7. References | |
| - **SWE-bench: Can Language Models Resolve Real-World GitHub Issues?** — Jimenez et al., | |
| Princeton, ICLR 2024 [source:arxiv:2310.06770]: the 2,294-task execution-based coding | |
| benchmark (3-stage construction, FAIL_TO_PASS/PASS_TO_PASS, BM25 vs oracle context, 1.96% | |
| at release), SWE-bench-train/Lite, and the continually-updatable contamination defense. | |
| - **WebArena: A Realistic Web Environment for Building Autonomous Agents** — Zhou et al., | |
| CMU, ICLR 2024 [source:arxiv:2307.13854]: the self-hosted 4-site environment, the | |
| `r_info`/`r_prog` functional-correctness reward (Table 1), multi-tab + accessibility-tree | |
| observations, unachievable-task refusal test, and the 14.41% vs 78.24% human gap. | |
| - **τ-bench: Tool-Agent-User Interaction in Real-World Domains** — Yao et al., Sierra, 2024 | |
| [source:arxiv:2406.12045]: the tool+policy+user-sim POMDP, state-match×output reward, the | |
| **pass^k** reliability metric (pass^1→pass^8 collapse), and the goal-reached-while-policy- | |
| violated reward-hacking caution. | |
| - **AgentBench: Evaluating LLMs as Agents** — Liu et al., ICLR 2024 [source:arxiv:2308.03688]: | |
| the 8-environment multi-turn suite, reciprocal-difficulty-weighted Overall Score, and the | |
| closed-vs-open capability gap (GPT-4 4.01 vs CodeLlama-34B 0.96). | |
| - **ReAct: Synergizing Reasoning and Acting in Language Models** — Yao et al. 2022/2023 | |
| [source:arxiv:2210.03629]: the Thought+Action agent scaffold these benchmarks evaluate. | |
| - **Chain-of-Thought Prompting** — Wei et al. 2022 [source:arxiv:2201.11903]: the reasoning | |
| prompt underlying the single-trial CoT agent baseline. | |
| - Forward links: `evaluation/capability-and-safety-benchmarks` (hub), | |
| `verifiable-rewards-and-reasoning/rlvr-overview`, `reward-modeling/verifiable-rewards`, | |
| `reward-modeling/process-vs-outcome-rewards`, `reward-modeling/reward-hacking`, | |
| `evaluation/judging-bias-and-contamination`, `evaluation/alignment-and-winrate-evals`. | |
Xet Storage Details
- Size:
- 18.8 kB
- Xet hash:
- c6a3d515779c4e85374787b34771c88f399246f3c6456ebd348612e4c654a5ac
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.