--- license: mit tags: - code - evaluation - benchmark - coding-agents - reliability - security pipeline_tag: text-generation --- # CodeBench **CodeBench** is an evaluation framework for measuring the *reliable* code generation ability of AI coding agents, beyond inflated pass@k scores. ## Key Findings | Finding | Result | |---|---| | **H1 — Category error** | Standard pass@k treats test-case counts as independent sample counts, producing different scores for agents with identical true reliability | | **H2 — Score inflation** | Current pass@5 ≈ 0.96–0.97 collapses to reliability@5 ≈ 0.00–0.12 when the category error is corrected | | **H3 — Proxy validity** | Single-rollout proxy score has low Spearman correlation with reliability@k; ≥5 rollouts needed for reliable ranking | | **H4 — Security gap** | Security-adjusted reliability@5 is meaningfully lower than reliability@5; leaderboard rank flips when security is accounted for | ## Metrics - **`reliability@k`** — correct operationalization of Chen et al. (2021) pass@k, using per-(task, agent) rollout counts and binary execution success - **`security_adjusted_reliability@k`** — reliability@k counting only rollouts that are both correct *and* produce code with no insecure patterns (eval, exec, os.system, yaml.load without Loader, pickle.loads) ## Agents Evaluated | Agent | Provider | Model | |---|---|---| | `anote-code` | Anthropic | claude-sonnet-4-6 (Anote system prompt) | | `claude-code` | Anthropic | claude-sonnet-4-6 | | `codex` | OpenAI | gpt-4o | ## Figures ### Figure 1 — Baseline: pass@1 vs current pass@5 ![fig1](figures/fig1_baseline.png) ### Figure 2 — H1: Category-error proof ![fig2](figures/fig2_h1_proof.png) ### Figure 3 — H2: Score inflation magnitude ![fig3](figures/fig3_h2_comparison.png) ### Figure 4 — H3: Proxy vs reliability@k correlation ![fig4](figures/fig4_h3_correlation.png) ### Figure 5 — H4: Security-adjusted reliability leaderboard ![fig5](figures/fig5_h4_security_leaderboard.png) ## Datasets - [anote-ai/codebench-tasks](https://huggingface.co/datasets/anote-ai/codebench-tasks) — benchmark task definitions - [anote-ai/codebench-results](https://huggingface.co/datasets/anote-ai/codebench-results) — experiment results (h4_security + swebench_smoke) ## Citation ``` @misc{codebench2026, title = {CodeBench: Measuring Reliable Code Generation}, author = {Anote AI}, year = {2026} } ```