Update README.md
Browse files
README.md
CHANGED
|
@@ -28,7 +28,7 @@ written solution. The recommended checkpoint is **`opd-32b-bf16-step-225`** ("st
|
|
| 28 |
|
| 29 |
| Benchmark | Result | Grader | Setting | Solutions |
|
| 30 |
|---|---|---|---|---|
|
| 31 |
-
| **IMO 2026** | **21 / 42 — Bronze medal** | Human expert (ex-IMO) | step-225, `high` budget | [
|
| 32 |
| **IMO 2025** | **30 / 42 — Silver medal** | GPT-5.6-sol | step-225, `high` budget | [solutions.csv](https://github.com/fieldsmodelorg/AIMO-Proof-Pilot/blob/main/benchmarks/imo-2025/solutions.csv) |
|
| 33 |
| **IMO-ProofBench-V2** | **66.4%** (mean **4.65 / 7**) | GPT-5.6-sol | step-225, `medium` budget | [solutions.csv](https://github.com/fieldsmodelorg/AIMO-Proof-Pilot/blob/main/benchmarks/imo-proofbench-v2/solutions.csv) |
|
| 34 |
|
|
@@ -131,7 +131,7 @@ live in [a separate HuggingFace repository](https://huggingface.co/fieldsmodelor
|
|
| 131 |
step-225 at the `high` budget, graded per-problem by human experts (former IMO
|
| 132 |
competitors) against broad IMO standards (no markscheme was available). The proofs are
|
| 133 |
the tournament
|
| 134 |
-
[`
|
| 135 |
per-problem
|
| 136 |
[`scores.csv`](https://github.com/fieldsmodelorg/AIMO-Proof-Pilot/blob/main/imo_2026_eval/scores.csv)
|
| 137 |
and full commentary are in
|
|
|
|
| 28 |
|
| 29 |
| Benchmark | Result | Grader | Setting | Solutions |
|
| 30 |
|---|---|---|---|---|
|
| 31 |
+
| **IMO 2026** | **21 / 42 — Bronze medal** | [Human expert (ex-IMO)](https://github.com/fieldsmodelorg/AIMO-Proof-Pilot/tree/main/imo_2026_eval) | step-225, `high` budget | [solutions.csv](https://github.com/fieldsmodelorg/AIMO-Proof-Pilot/blob/main/imo_2026_eval/raw/imo2026-step225-budget-high-tournament_submission.csv) |
|
| 32 |
| **IMO 2025** | **30 / 42 — Silver medal** | GPT-5.6-sol | step-225, `high` budget | [solutions.csv](https://github.com/fieldsmodelorg/AIMO-Proof-Pilot/blob/main/benchmarks/imo-2025/solutions.csv) |
|
| 33 |
| **IMO-ProofBench-V2** | **66.4%** (mean **4.65 / 7**) | GPT-5.6-sol | step-225, `medium` budget | [solutions.csv](https://github.com/fieldsmodelorg/AIMO-Proof-Pilot/blob/main/benchmarks/imo-proofbench-v2/solutions.csv) |
|
| 34 |
|
|
|
|
| 131 |
step-225 at the `high` budget, graded per-problem by human experts (former IMO
|
| 132 |
competitors) against broad IMO standards (no markscheme was available). The proofs are
|
| 133 |
the tournament
|
| 134 |
+
[`solutions.csv`](https://github.com/fieldsmodelorg/AIMO-Proof-Pilot/blob/main/imo_2026_eval/raw/imo2026-step225-budget-high-tournament_submission.csv);
|
| 135 |
per-problem
|
| 136 |
[`scores.csv`](https://github.com/fieldsmodelorg/AIMO-Proof-Pilot/blob/main/imo_2026_eval/scores.csv)
|
| 137 |
and full commentary are in
|