Buckets:
| # SWE-bench Pro: 10 agent harnesses x 2 models | |
| Every rollout behind the harness-vs-model comparison: 10 coding agent harnesses run | |
| against both models on the same 250 held-out SWE-bench Pro tasks, one rollout per task, | |
| 5,000 rollouts in total. | |
| | model | precision | pass@1 range across harnesses | | |
| |--------------------|-----------|-------------------------------| | |
| | GLM-5.2 744B-A40B | FP8 | 23.2% (goose) - 52.4% (mini_swe_agent) | | |
| | Gemma 4 26B-A4B | BF16 | 15.2% (goose) - 36.0% (crush) | | |
| The harness rankings do not transfer between the two models: the rank correlation between | |
| them is -0.05. Vendor-shipped harnesses fall on the small model (codex 2nd of 10 to 9th, | |
| claude_code 3rd to 7th) while model-agnostic ones climb (crush 7th to 1st, opencode 8th | |
| to 2nd). | |
| ## Layout | |
| ``` | |
| swebench_pro/test/<model>/<harness>/<task_id>/t0/ one rollout | |
| swebench_pro/test/<harness>/<model>.summary.json pass@1, Wilson 95% CI, coverage | |
| ``` | |
| Each rollout directory holds: | |
| | file | what it is | | |
| |-------------------------|-------------------------------------------------------------------| | |
| | `scores.json` | verdict, fail-to-pass counts, input/output tokens, wall clock, harness version | | |
| | `hf_router_usage.jsonl` | one line per model call with its prompt and completion sizes | | |
| | `agent_cmd.sh` | the exact command the harness was launched with | | |
| | `agent_stdout.log` | everything the harness printed | | |
| | `grade/` | the graded patch (`patch.diff`), test `output.json`, run scripts, stdout/stderr | | |
| | `candidate.patch` | the patch the agent produced, where the harness writes one separately | | |
| | `mini_traj.json` | mini_swe_agent's structured trajectory | | |
| | `codex_home/sessions/` | codex's session transcripts | | |
| ## What is not here | |
| Harness home directories (caches, sqlite journals, downloaded skills, shell snapshots) | |
| are excluded. They are an order of magnitude larger than the record itself -- one codex | |
| rollout is 68MB of which under 1MB describes the run -- and they are where a harness | |
| writes a credential if it writes one at all. Only the two directories above, which hold | |
| real transcripts, survive that cut. | |
| Token-shaped strings matching `hf_[A-Za-z0-9]{20,}` are replaced with | |
| `<REDACTED_HF_TOKEN>`. Private-key headers and `sk-` prefixes that remain are content | |
| from the benchmark's own task repositories (`gravitational/teleport`, `flipt-io/flipt`) | |
| and from branch names like `task-Fix-Timezone`, not credentials. | |
| ## Caveats | |
| 250 tasks and one rollout put a Wilson 95% interval of roughly +/-6pp around any single | |
| pass@1, so neighbouring harnesses are not separated. The reshuffling of the ranking | |
| between the two models is a joint pattern across ten harnesses and is the robust result. | |
| Costs quoted elsewhere use list API rates as of 2026-08-06: GLM-5.2 at Z.ai's own card | |
| ($1.40/$4.40 per M tokens), Gemma 4 at DeepInfra, the cheapest of 13 providers | |
| ($0.07/$0.34). | |
Xet Storage Details
- Size:
- 3.18 kB
- Xet hash:
- 0ac823d6c492bd7b36050ceef34d84ff0afa84ab20e650179404763f6ab6566c
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.