Buckets:
pretty_name: Scorio GPQA
license: mit
configs:
- config_name: Qwen3.6-35B-A3B
data_files:
- split: super_gpqa
path: data/Qwen3.6-35B-A3B/super_gpqa/*.parquet
- config_name: gpt-oss-20b_low
data_files:
- split: super_gpqa
path: data/gpt-oss-20b_low/super_gpqa/*.parquet
- config_name: gpt-oss-20b_medium
data_files:
- split: super_gpqa
path: data/gpt-oss-20b_medium/super_gpqa/*.parquet
- config_name: gpt-oss-20b_high
data_files:
- split: super_gpqa
path: data/gpt-oss-20b_high/super_gpqa/*.parquet
viewer: false
task_categories:
- question-answering
- text-generation
language:
- en
tags:
- reasoning
- test-time-scaling
- supergpqa
- log-probabilities
- top-k-logprobs
- uncertainty
- rollouts
size_categories:
- 1M<n<10M
Scorio GPQA
Scorio GPQA contains 1,152,000 sampled attempts from four model configurations on a fixed, field-balanced sample of superGPQA. The sample has 3,600 questions: 50 questions from each of 72 fields. Each model was run 80 times on every question.
Each token position includes the full top-20 candidate distribution. These distributions
are needed to reproduce token-level confidence measures such as self_certainty,
deepconf_confidence, token_entropy, varentropy, max_softmax_probability, and
logprob_margin. If these distributions are not needed, use
harimo/scorio-lite, which contains
the same attempts without the top-20 lists.
| Attempts | 1,152,000 |
| Questions | 3,600 |
| Fields | 72 |
| Parquet files | 14,400 |
| Download size | 679.4 GiB |
Quick start
The data is stored in a Hugging Face Storage Bucket. Install current versions of
datasets, huggingface_hub, and pyarrow before loading it.
pip install -U datasets huggingface_hub pyarrow
from datasets import load_dataset
data_files = {
"super_gpqa": "data/gpt-oss-20b_high/super_gpqa/*.parquet",
}
ds = load_dataset(
"buckets/harimo/scorio-gpqa",
data_files=data_files,
split="super_gpqa",
streaming=True,
)
record = next(iter(ds))
candidates = record["tokens"]["completion_topk_logprobs_list"][12]
One file contains the 80 attempts for one model and one question. To read a candidate pool without loading the token lists, select only the columns you need:
import pyarrow.parquet as pq
t = pq.read_table(
"hf://buckets/harimo/scorio-gpqa/data/gpt-oss-20b_high/super_gpqa/q0000.parquet",
columns=["seed", "evalscope_is_correct", "cv3b_label", "field"],
)
acc80 = sum(t["evalscope_is_correct"].to_pylist()) / 80
How to use
Install Scorio to evaluate models, rank them, or select an answer from a candidate pool:
pip install scorio
Scorio uses NumPy arrays whose dimensions match the structure of this dataset. Here, M
is the number of questions, N is the number of attempts per question, and L is the
number of model configurations. The dataset walkthrough
shows how to load candidate pools and work with the field-balanced question sample.
Evaluation APIs
Use scorio.eval to score one model. Build an M x N outcome matrix from
evalscope_is_correct, with one row per question and one column per seed. The module
includes average accuracy, Bayes@N, credible intervals, Pass@k, Maj@k, and related metrics.
See the evaluation notebook
for a complete example using ten questions from one field.
from scorio import eval
# R has shape (questions, seeds), with entries in {0, 1}
mu, sigma, lower, upper = eval.bayes_ci(R)
pass_at_8 = eval.pass_at_k(R, 8)
Ranking APIs
Use scorio.rank to compare models. Stack the outcome matrices into an L x M x N
array, ordered by model, question, and seed. Ranking methods can return both ranks and the
scores used to produce them. See the ranking notebook
for a complete example using all four model configurations.
from scorio import rank
# R_all has shape (models, questions, seeds)
ranks, scores = rank.bayes(R_all, return_scores=True)
Aggregation APIs
Use scorio.aggregate, also available as scorio.agg, to choose one answer from each
candidate pool. Pass an M x N array built from evalscope_extracted_answer. Score-based
methods take a second array of the same shape, using a verifier score or another confidence
score. See the aggregation notebook
for examples of voting and verifier-based selection.
from scorio import agg
# answers and verifier_scores both have shape (questions, seeds)
majority_answers = agg.majority_vote(answers)
best_answers = agg.best_of_n(answers, verifier_scores)
Layout
data/<model>/super_gpqa/qNNNN.parquet
NNNN is full_data_id, from 0000 through 3599. Every file has 80 rows ordered by seed.
Each model configuration contains 288,000 attempts.
Models
| config | model |
sampling.reasoning_effort |
accuracy | truncated |
|---|---|---|---|---|
Qwen3.6-35B-A3B |
Qwen/Qwen3.6-35B-A3B |
null |
0.637 | 0.1% |
gpt-oss-20b_low |
openai/gpt-oss-20b |
low |
0.350 | 0.0% |
gpt-oss-20b_medium |
openai/gpt-oss-20b |
medium |
0.421 | 0.0% |
gpt-oss-20b_high |
openai/gpt-oss-20b |
high |
0.450 | 1.1% |
Accuracy is mean evalscope_is_correct over all 288,000 attempts. Truncated is the
percentage with finish_reason == "length". The three gpt-oss configurations share the
same model value; use model_key when grouping them.
Question structure
A fixed, field-balanced sample of 72 fields with 50 questions from each field. It spans 13 disciplines and 278 subfields and is identical across all four models. Each field occupies a contiguous block of 50 question keys, so:
field_index = full_data_id // 50 # all 72 blocks present, no gaps
| discipline | fields | questions |
|---|---|---|
| Engineering | 30 | 1,500 |
| Science | 12 | 600 |
| Medicine | 6 | 300 |
| Agronomy | 5 | 250 |
| Literature and Arts | 4 | 200 |
| Management | 4 | 200 |
| Education | 3 | 150 |
| Economics | 2 | 100 |
| Law | 2 | 100 |
| History | 1 | 50 |
| Military Science | 1 | 50 |
| Philosophy | 1 | 50 |
| Sociology | 1 | 50 |
Difficulty: middle 1,607, easy 1,447, hard 546 questions.
Every attempt also carries discipline, field, subfield, difficulty and
is_calculation directly.
data_id is local to a collection stage and is not unique across the 3,600 questions. Use
full_data_id, uuid, selection_hash, or (stage, data_id) as a question identifier.
The filename is based on full_data_id.
Schema
The columns match the per-model GPQA configurations in
harimo/scorio-lite. They include the
generation, sampling settings, rule-based grading, superGPQA metadata,
CompassVerifier-3B scores, and scores from the reference-free verifier. Scorio GPQA adds
prompt_topk_logprobs_list and completion_topk_logprobs_list inside tokens. Each is a
list<list<struct<token: string, token_id: int32, logprob: float64, rank: int32>>>.
superGPQA is multiple choice, graded on a trailing ANSWER: [LETTER] line, so the strict
\boxed{} parse columns present in harimo/scorio-math do not exist here.
Eight source fields are not stored because they are exact functions of retained columns:
cv3b_prob_label, cv3b_reward, evalscope_acc, the three llmv_<c>_reward fields,
llmv_pointwise_reward, and llmv_mean_expected_raw_score_1_to_20. They can be recovered
with:
criteria = ["problem_understanding", "reasoning_validity", "conclusion_support"]
rewards = [(record[f"llmv_{c}_expected"] - 1) / 19 for c in criteria]
pointwise_reward = sum(rewards) / 3
cv3b_reward = float(record["cv3b_label"] == "A")
evalscope_acc = float(record["evalscope_is_correct"])
See harimo/scorio-lite for the full table.
OpenCompass verifier scores
The cv3b_* fields come from opencompass/CompassVerifier-3B. It sees the question,
reference answer, and candidate response, then judges the final answer as A (correct), B
(incorrect), or C (invalid). cv3b_label is the highest-logprob A/B/C choice,
cv3b_prob is its probability in the full vocabulary, and cv3b_abc_A/B/C renormalize
the three label probabilities. The separate cv3b_ctx_A/B/C diagnostic subtracts the
model's A/B/C log probabilities on a null [N/A] prompt from those on the real prompt,
divides by 1.5, and applies softmax.
LLM-as-a-verifier
The separate reference-free verifier uses Qwen/Qwen3.6-35B-A3B. It sees only the prompt
and response, and scores problem understanding, reasoning validity, and conclusion
support. For each criterion, it produces an A-to-T score-token distribution (A=20, ...,
T=1) and stores its expected value in llmv_<criterion>_expected. A criterion reward is
(expected_score - 1) / 19, and the overall pointwise reward is the mean of the three
criterion rewards. These derived rewards are not stored. This is a graded quality score,
not a calibrated probability of correctness.
Top-k list conventions
The prompt and completion candidate lists follow different conventions inherited from the source files.
prompt_topk_logprobs_list |
completion_topk_logprobs_list |
|
|---|---|---|
| Ordering | Descending by logprob |
Sampled token first; remaining entries sorted by logprob |
rank |
Vocabulary rank | Position from 1 to 20 |
| First entry | Argmax token | Sampled token |
| Width | 20, or 21 when the realized token is appended | 20 |
| model | unsorted completion rows | prompt rows with a 21st element | tail sorted from index 1 |
|---|---|---|---|
Qwen3.6-35B-A3B |
8.6% | 6.5% | 100.0% |
gpt-oss-20b_low |
14.5% | 38.1% | 100.0% |
gpt-oss-20b_medium |
21.8% | 38.1% | 100.0% |
gpt-oss-20b_high |
27.8% | 39.1% | 100.0% |
These measurements were made on q0000 for each model. They describe one question and
should not be treated as corpus-wide estimates.
Use completion_rank_list for the sampled token's vocabulary rank, not
completion_topk_logprobs_list[i][0]["rank"], which is always 1. For the model's ranking,
read from index 1 onward or sort the row by logprob. The entries from index 1 onward are
sorted, so row[1] is the argmax whenever row[0] is not.
Notes
gpt-oss response text
For gpt-oss, text contains only the Harmony final channel. For Qwen it contains the full
generation. Both verifiers saw only text, so cross-model and cross-effort comparisons of
their scores are affected by this difference. The full generation can be recovered from
tokens.completion_token_list.
Aborted attempts
Seven gpt-oss-20b_medium attempts finish with abort: one seed for full_data_id 2915
and six seeds for 2916. They are structurally complete and graded incorrect, but text
contains an analysis dump. Filter on finish_reason if these records are not suitable for
the analysis.
Reference-free verifier
Repeated scoring of byte-identical prompt-response pairs produced pointwise reward differences with a median of 0.040 and a maximum of 0.796. Keep this variation in mind when comparing small reward differences. CompassVerifier did not show the same variation.
Citation
@article{hariri2026test,
title={Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility},
author={Hariri, Mohsen and Chen, Weicong and Shahini, Nahal and Singh, Vikash and Ye, Kai
and Samandar, Amirhossein and Ganguly, Debargha and Sankar, Sreehari and Zhang, Yanyan
and Wang, Shouren and others},
journal={arXiv preprint arXiv:2608.04001},
year={2026}
}
superGPQA questions come from the superGPQA benchmark; please cite it as well.
Tooling: scorio.
Contact
For questions about the dataset, contact Mohsen Hariri.
License
This dataset is released under the MIT License. Questions reproduced from superGPQA remain subject to the benchmark's original terms.
Xet Storage Details
- Size:
- 12.7 kB
- Xet hash:
- a1c19739c64d91d7bdc7a03a4e94c6f4f79129e39af812ed455aa978d0079433
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.