Buckets:
| pretty_name: Anonymous Reasoning Math | |
| license: mit | |
| configs: | |
| - config_name: Qwen3.6-35B-A3B | |
| data_files: | |
| - split: aime_2026 | |
| path: data/Qwen3.6-35B-A3B/aime_2026/*.parquet | |
| - split: cmimc_2025 | |
| path: data/Qwen3.6-35B-A3B/cmimc_2025/*.parquet | |
| - split: hmmt_feb_2026 | |
| path: data/Qwen3.6-35B-A3B/hmmt_feb_2026/*.parquet | |
| - split: hmmt_nov_2025 | |
| path: data/Qwen3.6-35B-A3B/hmmt_nov_2025/*.parquet | |
| - split: smt_2025 | |
| path: data/Qwen3.6-35B-A3B/smt_2025/*.parquet | |
| - config_name: gpt-oss-20b_low | |
| data_files: | |
| - split: aime_2026 | |
| path: data/gpt-oss-20b_low/aime_2026/*.parquet | |
| - split: cmimc_2025 | |
| path: data/gpt-oss-20b_low/cmimc_2025/*.parquet | |
| - split: hmmt_feb_2026 | |
| path: data/gpt-oss-20b_low/hmmt_feb_2026/*.parquet | |
| - split: hmmt_nov_2025 | |
| path: data/gpt-oss-20b_low/hmmt_nov_2025/*.parquet | |
| - split: smt_2025 | |
| path: data/gpt-oss-20b_low/smt_2025/*.parquet | |
| - config_name: gpt-oss-20b_medium | |
| data_files: | |
| - split: aime_2026 | |
| path: data/gpt-oss-20b_medium/aime_2026/*.parquet | |
| - split: cmimc_2025 | |
| path: data/gpt-oss-20b_medium/cmimc_2025/*.parquet | |
| - split: hmmt_feb_2026 | |
| path: data/gpt-oss-20b_medium/hmmt_feb_2026/*.parquet | |
| - split: hmmt_nov_2025 | |
| path: data/gpt-oss-20b_medium/hmmt_nov_2025/*.parquet | |
| - split: smt_2025 | |
| path: data/gpt-oss-20b_medium/smt_2025/*.parquet | |
| - config_name: gpt-oss-20b_high | |
| data_files: | |
| - split: aime_2026 | |
| path: data/gpt-oss-20b_high/aime_2026/*.parquet | |
| - split: cmimc_2025 | |
| path: data/gpt-oss-20b_high/cmimc_2025/*.parquet | |
| - split: hmmt_feb_2026 | |
| path: data/gpt-oss-20b_high/hmmt_feb_2026/*.parquet | |
| - split: hmmt_nov_2025 | |
| path: data/gpt-oss-20b_high/hmmt_nov_2025/*.parquet | |
| - split: smt_2025 | |
| path: data/gpt-oss-20b_high/smt_2025/*.parquet | |
| viewer: false | |
| task_categories: | |
| - text-generation | |
| language: | |
| - en | |
| tags: | |
| - reasoning | |
| - test-time-scaling | |
| - math | |
| - log-probabilities | |
| - top-k-logprobs | |
| - uncertainty | |
| - rollouts | |
| size_categories: | |
| - 10K<n<100K | |
| # Anonymous Reasoning Math | |
| This repository contains data accompanying an anonymous TMLR submission. It provides | |
| 59,520 sampled attempts from four model configurations across five competition-math | |
| benchmarks. Each model was sampled 80 times per question. | |
| ## Contents | |
| The data is stored in a Hugging Face Storage Bucket. Each Parquet file contains the 80 | |
| attempts for one model and one question, ordered by seed. | |
| ```text | |
| data/<model>/<task>/qNN.parquet | |
| ``` | |
| | Split | Questions | Attempts per model | | |
| |---|---:|---:| | |
| | `aime_2026` | 30 | 2,400 | | |
| | `cmimc_2025` | 40 | 3,200 | | |
| | `hmmt_feb_2026` | 33 | 2,640 | | |
| | `hmmt_nov_2025` | 30 | 2,400 | | |
| | `smt_2025` | 53 | 4,240 | | |
| The bucket contains 744 Parquet files totaling about 166.8 GiB. The four configuration | |
| keys are `Qwen3.6-35B-A3B`, `gpt-oss-20b_low`, `gpt-oss-20b_medium`, and | |
| `gpt-oss-20b_high`. | |
| ## Loading | |
| ```bash | |
| pip install -U datasets huggingface_hub pyarrow | |
| ``` | |
| ```python | |
| from datasets import load_dataset | |
| data_files = { | |
| "cmimc_2025": "data/gpt-oss-20b_high/cmimc_2025/*.parquet", | |
| } | |
| dataset = load_dataset( | |
| "buckets/AnonymizedTMLRSubmission/reasoning-math", | |
| data_files=data_files, | |
| split="cmimc_2025", | |
| streaming=True, | |
| ) | |
| record = next(iter(dataset)) | |
| ``` | |
| ## Data Format | |
| | Field | Description | | |
| |---|---| | |
| | `task` | Benchmark or split identifier | | |
| | `model` | Model identifier | | |
| | `model_key` | Unique configuration key | | |
| | `data_id`, `seed` | Question identifier and sampling seed | | |
| | `prompt`, `trigger` | Problem text and generation instruction | | |
| | `sampling` | Sampling configuration | | |
| | `text`, `finish_reason` | Generated response and termination reason | | |
| | `num_prompt_tokens`, `num_completion_tokens` | Token counts | | |
| | `ground_truth` | Reference answer | | |
| | `evalscope_extracted_answer`, `evalscope_is_correct` | Extracted answer and rule-based correctness | | |
| | `cv3b_*` | Three-way verifier outputs and diagnostic values | | |
| | `llmv_*` | Reference-free verifier criterion scores | | |
| | `tokens` | Aggregate token statistics and, where present, token-level arrays | | |
| Math-specific fields include `extracted_answer`, `boxed_answer_raw`, `has_box`, | |
| `has_valid_box`, `num_boxes`, `box_status`, and `boxed_is_correct`. | |
| The `tokens` struct includes prompt and completion log-probability summaries, realized | |
| token arrays, and top-20 candidate distributions. Each candidate contains `token`, | |
| `token_id`, `logprob`, and `rank`. The first completion candidate is the sampled token; | |
| remaining candidates are ordered by log probability. Use `completion_rank_list` for the | |
| sampled token's vocabulary rank. | |
| ## Reproducibility | |
| Rows are ordered by seed within each question. Length-truncated responses are graded | |
| incorrect because they do not contain a final boxed answer. Report the truncation policy | |
| used in comparisons. Verifier outputs should be treated as measurements with their own | |
| uncertainty. | |
| ## License | |
| The artifact is provided under the MIT License. Problem statements from the underlying | |
| competitions remain subject to their applicable terms. | |
| ## Citation | |
| Citation information will be added after anonymous review. | |
Xet Storage Details
- Size:
- 5.08 kB
- Xet hash:
- 282f84f5b087fbf023bd0d6f83bd4e66f055d620c584a0821461f025eb568143
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.