harimo/scorio-math / README.md
harimo's picture
|
download
raw
14.6 kB
---
pretty_name: Scorio Math
license: mit
configs:
- config_name: Qwen3.6-35B-A3B
data_files:
- split: aime_2026
path: data/Qwen3.6-35B-A3B/aime_2026/*.parquet
- split: cmimc_2025
path: data/Qwen3.6-35B-A3B/cmimc_2025/*.parquet
- split: hmmt_feb_2026
path: data/Qwen3.6-35B-A3B/hmmt_feb_2026/*.parquet
- split: hmmt_nov_2025
path: data/Qwen3.6-35B-A3B/hmmt_nov_2025/*.parquet
- split: smt_2025
path: data/Qwen3.6-35B-A3B/smt_2025/*.parquet
- config_name: gpt-oss-20b_low
data_files:
- split: aime_2026
path: data/gpt-oss-20b_low/aime_2026/*.parquet
- split: cmimc_2025
path: data/gpt-oss-20b_low/cmimc_2025/*.parquet
- split: hmmt_feb_2026
path: data/gpt-oss-20b_low/hmmt_feb_2026/*.parquet
- split: hmmt_nov_2025
path: data/gpt-oss-20b_low/hmmt_nov_2025/*.parquet
- split: smt_2025
path: data/gpt-oss-20b_low/smt_2025/*.parquet
- config_name: gpt-oss-20b_medium
data_files:
- split: aime_2026
path: data/gpt-oss-20b_medium/aime_2026/*.parquet
- split: cmimc_2025
path: data/gpt-oss-20b_medium/cmimc_2025/*.parquet
- split: hmmt_feb_2026
path: data/gpt-oss-20b_medium/hmmt_feb_2026/*.parquet
- split: hmmt_nov_2025
path: data/gpt-oss-20b_medium/hmmt_nov_2025/*.parquet
- split: smt_2025
path: data/gpt-oss-20b_medium/smt_2025/*.parquet
- config_name: gpt-oss-20b_high
data_files:
- split: aime_2026
path: data/gpt-oss-20b_high/aime_2026/*.parquet
- split: cmimc_2025
path: data/gpt-oss-20b_high/cmimc_2025/*.parquet
- split: hmmt_feb_2026
path: data/gpt-oss-20b_high/hmmt_feb_2026/*.parquet
- split: hmmt_nov_2025
path: data/gpt-oss-20b_high/hmmt_nov_2025/*.parquet
- split: smt_2025
path: data/gpt-oss-20b_high/smt_2025/*.parquet
viewer: false
task_categories:
- text-generation
language:
- en
tags:
- reasoning
- test-time-scaling
- math
- log-probabilities
- top-k-logprobs
- uncertainty
- rollouts
size_categories:
- 10K<n<100K
---
# Scorio Math
Scorio Math contains 59,520 sampled attempts from four model configurations and five
competition-math benchmarks. Each model was run 80 times on every question. One Parquet
file contains the 80 attempts for one model and one question, ordered by seed.
Each token position includes the full top-20 candidate distribution. These distributions
are needed to reproduce token-level confidence measures such as `self_certainty`,
`deepconf_confidence`, `token_entropy`, `varentropy`, `max_softmax_probability`, and
`logprob_margin`. If these distributions are not needed, use
[`harimo/scorio-lite`](https://huggingface.co/datasets/harimo/scorio-lite), which contains
the same attempts without the top-20 lists.
| | |
|---|---:|
| Attempts | 59,520 |
| Questions | 186 |
| Parquet files | 744 |
| Download size | 166.8 GiB |
## Quick start
The data is stored in a Hugging Face Storage Bucket. Install current versions of
`datasets`, `huggingface_hub`, and `pyarrow` before loading it.
```bash
pip install -U datasets huggingface_hub pyarrow
```
```python
from datasets import load_dataset
data_files = {
"cmimc_2025": "data/gpt-oss-20b_high/cmimc_2025/*.parquet",
}
ds = load_dataset(
"buckets/harimo/scorio-math",
data_files=data_files,
split="cmimc_2025",
streaming=True,
)
record = next(iter(ds))
candidates = record["tokens"]["completion_topk_logprobs_list"][12]
```
Rows with token distributions can be large. Use streaming, or select only the columns you
need when reading an individual file:
```python
import pyarrow.parquet as pq
t = pq.read_table(
"hf://buckets/harimo/scorio-math/data/gpt-oss-20b_high/cmimc_2025/q17.parquet",
columns=["seed", "evalscope_is_correct", "cv3b_label"],
)
acc80 = sum(t["evalscope_is_correct"].to_pylist()) / 80
```
## How to use
Install [Scorio](https://github.com/mohsenhariri/scorio) to evaluate models, rank them,
or select an answer from a candidate pool:
```bash
pip install scorio
```
Scorio uses NumPy arrays whose dimensions match the structure of this dataset. Here, `M`
is the number of questions, `N` is the number of attempts per question, and `L` is the
number of model configurations. The [dataset walkthrough](https://github.com/mohsenhariri/scorio/blob/main/notebooks/datasets/math/math.ipynb)
shows how to load candidate pools and work with the five competition-math benchmarks.
## Evaluation APIs
Use [`scorio.eval`](https://github.com/mohsenhariri/scorio/tree/main/scorio/eval) to score one model. Build an `M x N` outcome matrix from
`evalscope_is_correct`, with one row per question and one column per seed. The module
includes average accuracy, Bayes@N, credible intervals, Pass@k, Maj@k, and related metrics.
See the **[evaluation notebook](https://github.com/mohsenhariri/scorio/blob/main/notebooks/datasets/math/eval.ipynb)**
for a complete example using an AIME 2026 question window.
```python
from scorio import eval
# R has shape (questions, seeds), with entries in {0, 1}
mu, sigma, lower, upper = eval.bayes_ci(R)
pass_at_8 = eval.pass_at_k(R, 8)
```
## Ranking APIs
Use [`scorio.rank`](https://github.com/mohsenhariri/scorio/tree/main/scorio/rank) to compare models. Stack the outcome matrices into an `L x M x N`
array, ordered by model, question, and seed. Ranking methods can return both ranks and the
scores used to produce them. See the **[ranking notebook](https://github.com/mohsenhariri/scorio/blob/main/notebooks/datasets/math/rank.ipynb)**
for a complete example using all four model configurations.
```python
from scorio import rank
# R_all has shape (models, questions, seeds)
ranks, scores = rank.bayes(R_all, return_scores=True)
```
## Aggregation APIs
Use [`scorio.aggregate`](https://github.com/mohsenhariri/scorio/tree/main/scorio/aggregate), also available as `scorio.agg`, to choose one answer from each
candidate pool. Pass an `M x N` array built from `extracted_answer`. Score-based methods
take a second array of the same shape, using a verifier score or another confidence score.
See the **[aggregation notebook](https://github.com/mohsenhariri/scorio/blob/main/notebooks/datasets/math/aggregate.ipynb)**
for examples of voting and verifier-based selection.
```python
from scorio import agg
# answers and verifier_scores both have shape (questions, seeds)
majority_answers = agg.majority_vote(answers)
best_answers = agg.best_of_n(answers, verifier_scores)
```
## Layout
```text
data/<model>/<task>/qNN.parquet
```
There are four model configurations. Each configuration has five splits. Every file has 80
rows in ascending seed order.
| task | questions | attempts per model |
|---|---:|---:|
| `aime_2026` | 30 | 2,400 |
| `cmimc_2025` | 40 | 3,200 |
| `hmmt_feb_2026` | 33 | 2,640 |
| `hmmt_nov_2025` | 30 | 2,400 |
| `smt_2025` | 53 | 4,240 |
## Models
| config | `model` | `sampling.reasoning_effort` |
|---|---|---|
| `Qwen3.6-35B-A3B` | `Qwen/Qwen3.6-35B-A3B` | `null` |
| `gpt-oss-20b_low` | `openai/gpt-oss-20b` | `low` |
| `gpt-oss-20b_medium` | `openai/gpt-oss-20b` | `medium` |
| `gpt-oss-20b_high` | `openai/gpt-oss-20b` | `high` |
The three gpt-oss configurations share the same `model` value. Use `model_key` when
grouping them. The reasoning level is also stored in `sampling.reasoning_effort`.
## Schema
The columns match the per-model math configurations in
[`harimo/scorio-lite`](https://huggingface.co/datasets/harimo/scorio-lite). They include the
generation, sampling settings, rule-based grading, CompassVerifier-3B scores, and scores
from the reference-free verifier. Scorio Math adds two fields inside `tokens`:
| field | type | meaning |
|---|---|---|
| `prompt_topk_logprobs_list` | `list<list<struct>>` | candidates per prompt position |
| `completion_topk_logprobs_list` | `list<list<struct>>` | candidates per completion position |
Each entry is `{token: string, token_id: int32, logprob: float64, rank: int32}`.
Prompt-side length is `num_prompt_tokens`, completion-side `num_completion_tokens`.
Eight source fields are not stored because they are exact functions of retained columns:
`cv3b_prob_label`, `cv3b_reward`, `evalscope_acc`, the three `llmv_<c>_reward` fields,
`llmv_pointwise_reward`, and `llmv_mean_expected_raw_score_1_to_20`. They can be recovered
with:
```python
criteria = ["problem_understanding", "reasoning_validity", "conclusion_support"]
rewards = [(record[f"llmv_{c}_expected"] - 1) / 19 for c in criteria]
pointwise_reward = sum(rewards) / 3
cv3b_reward = float(record["cv3b_label"] == "A")
evalscope_acc = float(record["evalscope_is_correct"])
```
See [`harimo/scorio-lite`](https://huggingface.co/datasets/harimo/scorio-lite) for the full table.
### OpenCompass verifier scores
The `cv3b_*` fields come from `opencompass/CompassVerifier-3B`. It sees the question,
reference answer, and candidate response, then judges the final answer as A (correct), B
(incorrect), or C (invalid). `cv3b_label` is the highest-logprob A/B/C choice,
`cv3b_prob` is its probability in the full vocabulary, and `cv3b_abc_A/B/C` renormalize
the three label probabilities. The separate `cv3b_ctx_A/B/C` diagnostic subtracts the
model's A/B/C log probabilities on a null `[N/A]` prompt from those on the real prompt,
divides by 1.5, and applies softmax.
### LLM-as-a-verifier
The separate reference-free verifier uses `Qwen/Qwen3.6-35B-A3B`. It sees only the prompt
and response, and scores problem understanding, reasoning validity, and conclusion
support. For each criterion, it produces an A-to-T score-token distribution (A=20, ...,
T=1) and stores its expected value in `llmv_<criterion>_expected`. A criterion reward is
`(expected_score - 1) / 19`, and the overall pointwise reward is the mean of the three
criterion rewards. These derived rewards are not stored. This is a graded quality score,
not a calibrated probability of correctness.
## Top-k list conventions
The prompt and completion candidate lists follow different conventions inherited from the
source files.
| | `prompt_topk_logprobs_list` | `completion_topk_logprobs_list` |
|---|---|---|
| Ordering | Descending by `logprob` | Sampled token first; remaining entries sorted by `logprob` |
| `rank` | Vocabulary rank | Position from 1 to 20 |
| First entry | Argmax token | Sampled token |
| Width | 20, or 21 when the realized token is appended | 20 |
The following measurements were made on `aime_2026/q00`. They describe one question and
show how often the completion list starts with a non-argmax sampled token.
| model | unsorted rows | prompt rows carrying a 21st element |
|---|---:|---:|
| `Qwen3.6-35B-A3B` | 6.9% | 5.4% |
| `gpt-oss-20b_low` | 7.2% | 34.6% |
| `gpt-oss-20b_medium` | 10.5% | 34.2% |
| `gpt-oss-20b_high` | 16.6% | 35.9% |
Use these conventions as follows:
- For the sampled token's vocabulary rank, use `completion_rank_list`, not
`completion_topk_logprobs_list[i][0]["rank"]`, which is always `1`;
- For the model's ranking at position `i`, read from index 1 onward, or sort the
row by `logprob`;
- The entries from index 1 onward are sorted. `row[1]` is the argmax whenever `row[0]` is not.
## Accuracy
Mean `evalscope_is_correct` over all 80 seeds of every question.
| model | `aime_2026` | `cmimc_2025` | `hmmt_feb_2026` | `hmmt_nov_2025` | `smt_2025` |
|---|---:|---:|---:|---:|---:|
| `Qwen3.6-35B-A3B` | 0.922 | 0.816 | 0.795 | 0.850 | 0.803 |
| `gpt-oss-20b_low` | 0.425 | 0.242 | 0.266 | 0.311 | 0.389 |
| `gpt-oss-20b_medium` | 0.777 | 0.542 | 0.539 | 0.645 | 0.655 |
| `gpt-oss-20b_high` | 0.876 | 0.704 | 0.625 | 0.720 | 0.740 |
## Truncation (`finish_reason == "length"`)
| model | `aime_2026` | `cmimc_2025` | `hmmt_feb_2026` | `hmmt_nov_2025` | `smt_2025` |
|---|---:|---:|---:|---:|---:|
| `Qwen3.6-35B-A3B` | 3.6% | 5.8% | 7.1% | 6.5% | 1.5% |
| `gpt-oss-20b_low` | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% |
| `gpt-oss-20b_medium` | 0.2% | 0.3% | 0.5% | 0.2% | 0.0% |
| `gpt-oss-20b_high` | 10.0% | 21.3% | 23.9% | 17.4% | 10.1% |
A length-truncated attempt has no final boxed answer and is graded incorrect. Report the
truncation policy when comparing reasoning levels.
## Notes
### gpt-oss response text
For gpt-oss, `text` contains only the Harmony final channel. For Qwen it contains the full
generation. Both verifiers saw only `text`, so cross-model and cross-effort comparisons of
their scores are affected by this difference. The full generation can be recovered from
`tokens.completion_token_list`.
### CompassVerifier input length
CompassVerifier-3B has a 32k context limit. It scored 8,750 attempts on truncated responses.
The dataset does not contain a flag for this condition.
### Reference-free verifier
Repeated scoring of byte-identical prompt-response pairs produced pointwise reward
differences with a median of 0.040 and a maximum of 0.796. Keep this variation in mind when
comparing small reward differences. CompassVerifier did not show the same variation.
### Row size
Median uncompressed row size ranges from 0.3 MB for `gpt-oss-20b_low` to 18.3 MB for
`gpt-oss-20b_high`; the largest measured row is 22.6 MB. Use streaming or read selected
columns from individual Parquet files.
## Provenance
The data was converted from question-major gzipped JSONL to Parquet. Source grouping and
ascending seed order were preserved. Nested grading and verifier results were flattened,
while token data remains in the `tokens` struct. Every written record, including every
top-20 entry, was compared field by field with its source record.
The gzipped source corpus was also compared with the inference output using record-level
BLAKE2b digests. Parquet serialization changes the bytes, so the field-by-field comparison
is the relevant check for this release.
## Citation
```bibtex
@article{hariri2026test,
title={Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility},
author={Hariri, Mohsen and Chen, Weicong and Shahini, Nahal and Singh, Vikash and Ye, Kai
and Samandar, Amirhossein and Ganguly, Debargha and Sankar, Sreehari and Zhang, Yanyan
and Wang, Shouren and others},
journal={arXiv preprint arXiv:2608.04001},
year={2026}
}
```
Related: [arXiv:2510.04265](https://arxiv.org/abs/2510.04265) ·
[arXiv:2603.10960](https://arxiv.org/abs/2603.10960) ·
[arXiv:2603.14103](https://arxiv.org/abs/2603.14103).
Tooling: [`scorio`](https://github.com/mohsenhariri/scorio).
## Contact
For questions about the dataset, contact [Mohsen Hariri](mailto:mohsen.hariri@case.edu).
## License
This dataset is released under the [MIT License](LICENSE). Problem statements from AIME,
HMMT, SMT, and CMIMC remain subject to their original terms.

Xet Storage Details

Size:
14.6 kB
·
Xet hash:
445c1c36c0d3422c4fa727446e07247a4edd2ce2a45ae7a85c78912dd68bd8b2

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.