|
download
raw
5.08 kB
---
pretty_name: Anonymous Reasoning Math
license: mit
configs:
- config_name: Qwen3.6-35B-A3B
data_files:
- split: aime_2026
path: data/Qwen3.6-35B-A3B/aime_2026/*.parquet
- split: cmimc_2025
path: data/Qwen3.6-35B-A3B/cmimc_2025/*.parquet
- split: hmmt_feb_2026
path: data/Qwen3.6-35B-A3B/hmmt_feb_2026/*.parquet
- split: hmmt_nov_2025
path: data/Qwen3.6-35B-A3B/hmmt_nov_2025/*.parquet
- split: smt_2025
path: data/Qwen3.6-35B-A3B/smt_2025/*.parquet
- config_name: gpt-oss-20b_low
data_files:
- split: aime_2026
path: data/gpt-oss-20b_low/aime_2026/*.parquet
- split: cmimc_2025
path: data/gpt-oss-20b_low/cmimc_2025/*.parquet
- split: hmmt_feb_2026
path: data/gpt-oss-20b_low/hmmt_feb_2026/*.parquet
- split: hmmt_nov_2025
path: data/gpt-oss-20b_low/hmmt_nov_2025/*.parquet
- split: smt_2025
path: data/gpt-oss-20b_low/smt_2025/*.parquet
- config_name: gpt-oss-20b_medium
data_files:
- split: aime_2026
path: data/gpt-oss-20b_medium/aime_2026/*.parquet
- split: cmimc_2025
path: data/gpt-oss-20b_medium/cmimc_2025/*.parquet
- split: hmmt_feb_2026
path: data/gpt-oss-20b_medium/hmmt_feb_2026/*.parquet
- split: hmmt_nov_2025
path: data/gpt-oss-20b_medium/hmmt_nov_2025/*.parquet
- split: smt_2025
path: data/gpt-oss-20b_medium/smt_2025/*.parquet
- config_name: gpt-oss-20b_high
data_files:
- split: aime_2026
path: data/gpt-oss-20b_high/aime_2026/*.parquet
- split: cmimc_2025
path: data/gpt-oss-20b_high/cmimc_2025/*.parquet
- split: hmmt_feb_2026
path: data/gpt-oss-20b_high/hmmt_feb_2026/*.parquet
- split: hmmt_nov_2025
path: data/gpt-oss-20b_high/hmmt_nov_2025/*.parquet
- split: smt_2025
path: data/gpt-oss-20b_high/smt_2025/*.parquet
viewer: false
task_categories:
- text-generation
language:
- en
tags:
- reasoning
- test-time-scaling
- math
- log-probabilities
- top-k-logprobs
- uncertainty
- rollouts
size_categories:
- 10K<n<100K
---
# Anonymous Reasoning Math
This repository contains data accompanying an anonymous TMLR submission. It provides
59,520 sampled attempts from four model configurations across five competition-math
benchmarks. Each model was sampled 80 times per question.
## Contents
The data is stored in a Hugging Face Storage Bucket. Each Parquet file contains the 80
attempts for one model and one question, ordered by seed.
```text
data/<model>/<task>/qNN.parquet
```
| Split | Questions | Attempts per model |
|---|---:|---:|
| `aime_2026` | 30 | 2,400 |
| `cmimc_2025` | 40 | 3,200 |
| `hmmt_feb_2026` | 33 | 2,640 |
| `hmmt_nov_2025` | 30 | 2,400 |
| `smt_2025` | 53 | 4,240 |
The bucket contains 744 Parquet files totaling about 166.8 GiB. The four configuration
keys are `Qwen3.6-35B-A3B`, `gpt-oss-20b_low`, `gpt-oss-20b_medium`, and
`gpt-oss-20b_high`.
## Loading
```bash
pip install -U datasets huggingface_hub pyarrow
```
```python
from datasets import load_dataset
data_files = {
"cmimc_2025": "data/gpt-oss-20b_high/cmimc_2025/*.parquet",
}
dataset = load_dataset(
"buckets/AnonymizedTMLRSubmission/reasoning-math",
data_files=data_files,
split="cmimc_2025",
streaming=True,
)
record = next(iter(dataset))
```
## Data Format
| Field | Description |
|---|---|
| `task` | Benchmark or split identifier |
| `model` | Model identifier |
| `model_key` | Unique configuration key |
| `data_id`, `seed` | Question identifier and sampling seed |
| `prompt`, `trigger` | Problem text and generation instruction |
| `sampling` | Sampling configuration |
| `text`, `finish_reason` | Generated response and termination reason |
| `num_prompt_tokens`, `num_completion_tokens` | Token counts |
| `ground_truth` | Reference answer |
| `evalscope_extracted_answer`, `evalscope_is_correct` | Extracted answer and rule-based correctness |
| `cv3b_*` | Three-way verifier outputs and diagnostic values |
| `llmv_*` | Reference-free verifier criterion scores |
| `tokens` | Aggregate token statistics and, where present, token-level arrays |
Math-specific fields include `extracted_answer`, `boxed_answer_raw`, `has_box`,
`has_valid_box`, `num_boxes`, `box_status`, and `boxed_is_correct`.
The `tokens` struct includes prompt and completion log-probability summaries, realized
token arrays, and top-20 candidate distributions. Each candidate contains `token`,
`token_id`, `logprob`, and `rank`. The first completion candidate is the sampled token;
remaining candidates are ordered by log probability. Use `completion_rank_list` for the
sampled token's vocabulary rank.
## Reproducibility
Rows are ordered by seed within each question. Length-truncated responses are graded
incorrect because they do not contain a final boxed answer. Report the truncation policy
used in comparisons. Verifier outputs should be treated as measurements with their own
uncertainty.
## License
The artifact is provided under the MIT License. Problem statements from the underlying
competitions remain subject to their applicable terms.
## Citation
Citation information will be added after anonymous review.

Xet Storage Details

Size:
5.08 kB
·
Xet hash:
282f84f5b087fbf023bd0d6f83bd4e66f055d620c584a0821461f025eb568143

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.