chankhavu commited on
Commit
dba668c
·
verified ·
1 Parent(s): 7862528

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +2 -2
README.md CHANGED
@@ -28,7 +28,7 @@ written solution. The recommended checkpoint is **`opd-32b-bf16-step-225`** ("st
28
 
29
  | Benchmark | Result | Grader | Setting | Solutions |
30
  |---|---|---|---|---|
31
- | **IMO 2026** | **21 / 42 — Bronze medal** | Human expert (ex-IMO) | step-225, `high` budget | [submission.csv](https://huggingface.co/datasets/imo2026-challenge/chankhavu-imo-reasoning-traces/blob/main/imo2026-step225-budget-high-tournament/submission.csv) |
32
  | **IMO 2025** | **30 / 42 — Silver medal** | GPT-5.6-sol | step-225, `high` budget | [solutions.csv](https://github.com/fieldsmodelorg/AIMO-Proof-Pilot/blob/main/benchmarks/imo-2025/solutions.csv) |
33
  | **IMO-ProofBench-V2** | **66.4%** (mean **4.65 / 7**) | GPT-5.6-sol | step-225, `medium` budget | [solutions.csv](https://github.com/fieldsmodelorg/AIMO-Proof-Pilot/blob/main/benchmarks/imo-proofbench-v2/solutions.csv) |
34
 
@@ -131,7 +131,7 @@ live in [a separate HuggingFace repository](https://huggingface.co/fieldsmodelor
131
  step-225 at the `high` budget, graded per-problem by human experts (former IMO
132
  competitors) against broad IMO standards (no markscheme was available). The proofs are
133
  the tournament
134
- [`submission.csv`](https://huggingface.co/datasets/imo2026-challenge/chankhavu-imo-reasoning-traces/blob/main/imo2026-step225-budget-high-tournament/submission.csv);
135
  per-problem
136
  [`scores.csv`](https://github.com/fieldsmodelorg/AIMO-Proof-Pilot/blob/main/imo_2026_eval/scores.csv)
137
  and full commentary are in
 
28
 
29
  | Benchmark | Result | Grader | Setting | Solutions |
30
  |---|---|---|---|---|
31
+ | **IMO 2026** | **21 / 42 — Bronze medal** | [Human expert (ex-IMO)](https://github.com/fieldsmodelorg/AIMO-Proof-Pilot/tree/main/imo_2026_eval) | step-225, `high` budget | [solutions.csv](https://github.com/fieldsmodelorg/AIMO-Proof-Pilot/blob/main/imo_2026_eval/raw/imo2026-step225-budget-high-tournament_submission.csv) |
32
  | **IMO 2025** | **30 / 42 — Silver medal** | GPT-5.6-sol | step-225, `high` budget | [solutions.csv](https://github.com/fieldsmodelorg/AIMO-Proof-Pilot/blob/main/benchmarks/imo-2025/solutions.csv) |
33
  | **IMO-ProofBench-V2** | **66.4%** (mean **4.65 / 7**) | GPT-5.6-sol | step-225, `medium` budget | [solutions.csv](https://github.com/fieldsmodelorg/AIMO-Proof-Pilot/blob/main/benchmarks/imo-proofbench-v2/solutions.csv) |
34
 
 
131
  step-225 at the `high` budget, graded per-problem by human experts (former IMO
132
  competitors) against broad IMO standards (no markscheme was available). The proofs are
133
  the tournament
134
+ [`solutions.csv`](https://github.com/fieldsmodelorg/AIMO-Proof-Pilot/blob/main/imo_2026_eval/raw/imo2026-step225-budget-high-tournament_submission.csv);
135
  per-problem
136
  [`scores.csv`](https://github.com/fieldsmodelorg/AIMO-Proof-Pilot/blob/main/imo_2026_eval/scores.csv)
137
  and full commentary are in