MATH data files (graphrpo)
Honest naming of the MATH eval/train parquets. See
experiments/verl/notes/reproduction_plan.md for the full analysis.
| file | what it actually is | rows | level mix (L1/L2/L3/L4/L5) | use |
|---|---|---|---|---|
train.parquet |
MATH-lighteval train split | 7500 | — | training |
test.parquet |
MATH-lighteval test split (full) | 5000 | — | source of the slice below |
test500.parquet |
test.parquet[:500] — a PREFIX SLICE, not MATH-500 |
500 | 64/82/114/123/117 | legacy eval; still read by in-flight jobs |
math_test_first500.parquet |
byte-identical copy of test500.parquet (honest name) |
500 | 64/82/114/123/117 | prefer this name going forward for the slice |
math500_official.parquet |
the TRUE MATH-500 (Lightman et al. / PRM800K subset, HuggingFaceH4/MATH-500) |
497* | 43/90/105/128/134 (of the canonical 500) | standard benchmark eval |
* The canonical MATH-500 has 500 problems with mix L1:43 L2:90 L3:105 L4:128 L5:134.
3 problems tokenize to >512 tokens under the Qwen2.5-1.5B-Instruct chat template and are
dropped by verl's filter_overlong_prompts (mirrored at build time), leaving 497 rows.
Emitted token lengths: min 52, max 504, mean 111.
Why this matters
test500.parquet is not the published MATH-500 benchmark. Its level mix is easier
(more Level-1, fewer Level-5) than canonical MATH-500, so any "MATH-500" number reported
on it is on a non-standard, non-comparable scale. Reproductions and paper numbers must use
math500_official.parquet.
Do not delete
test500.parquet is left in place because in-flight jobs read it. New work should use
math500_official.parquet (real benchmark) or math_test_first500.parquet (the slice,
honestly named).
Built by experiments/verl/repro/prep_math500.py (source HuggingFaceH4/MATH-500).