|
download
raw
4.69 kB
metadata
pretty_name: Anonymous Reasoning GPQA
license: mit
configs:
  - config_name: Qwen3.6-35B-A3B
    data_files:
      - split: super_gpqa
        path: data/Qwen3.6-35B-A3B/super_gpqa/*.parquet
  - config_name: gpt-oss-20b_low
    data_files:
      - split: super_gpqa
        path: data/gpt-oss-20b_low/super_gpqa/*.parquet
  - config_name: gpt-oss-20b_medium
    data_files:
      - split: super_gpqa
        path: data/gpt-oss-20b_medium/super_gpqa/*.parquet
  - config_name: gpt-oss-20b_high
    data_files:
      - split: super_gpqa
        path: data/gpt-oss-20b_high/super_gpqa/*.parquet
viewer: false
task_categories:
  - question-answering
  - text-generation
language:
  - en
tags:
  - reasoning
  - test-time-scaling
  - supergpqa
  - log-probabilities
  - top-k-logprobs
  - uncertainty
  - rollouts
size_categories:
  - 1M<n<10M

Anonymous Reasoning GPQA

This repository contains data accompanying an anonymous TMLR submission. It provides 1,152,000 sampled attempts from four model configurations on a fixed, field-balanced sample of 3,600 multiple-choice questions. Each model was sampled 80 times per question.

Contents

The data is stored in a Hugging Face Storage Bucket. Each Parquet file contains the 80 attempts for one model and one question, ordered by seed.

data/<model>/super_gpqa/qNNNN.parquet

There are 14,400 Parquet files totaling about 679.4 GiB. The four configuration keys are Qwen3.6-35B-A3B, gpt-oss-20b_low, gpt-oss-20b_medium, and gpt-oss-20b_high.

Loading

pip install -U datasets huggingface_hub pyarrow
from datasets import load_dataset

data_files = {
    "super_gpqa": "data/gpt-oss-20b_high/super_gpqa/*.parquet",
}
dataset = load_dataset(
    "buckets/AnonymizedTMLRSubmission/reasoning-gpqa",
    data_files=data_files,
    split="super_gpqa",
    streaming=True,
)
record = next(iter(dataset))

For direct access to one candidate pool:

import pyarrow.parquet as pq

table = pq.read_table(
    "data/gpt-oss-20b_high/super_gpqa/q0000.parquet",
    columns=["seed", "evalscope_is_correct", "cv3b_label", "field"],
)

Question Structure

The sample contains 72 fields with 50 questions per field. It spans 13 disciplines and 278 subfields. Every attempt includes discipline, field, subfield, difficulty, and is_calculation. Use full_data_id, uuid, selection_hash, or (stage, data_id) as a question identifier because data_id alone is not globally unique.

Data Format

Field Description
task Benchmark or split identifier
model Model identifier
model_key Unique configuration key
data_id, seed Question identifier and sampling seed
prompt, trigger Problem text and generation instruction
sampling Sampling configuration
text, finish_reason Generated response and termination reason
num_prompt_tokens, num_completion_tokens Token counts
ground_truth Reference answer
evalscope_extracted_answer, evalscope_is_correct Extracted answer and rule-based correctness
cv3b_* Three-way verifier outputs and diagnostic values
llmv_* Reference-free verifier criterion scores
tokens Aggregate token statistics and, where present, token-level arrays

Additional question metadata includes stage, full_data_id, uuid, field_rank, source_line, selection_hash, discipline, field, subfield, difficulty, and is_calculation.

The tokens struct includes prompt and completion log-probability summaries, token IDs, realized-token log probabilities and ranks, and top-20 candidate distributions. Each candidate entry contains token, token_id, logprob, and rank.

The first completion candidate is the sampled token. Remaining candidates are ordered by log probability. Use completion_rank_list for the sampled token's vocabulary rank.

Reproducibility

Question pools contain exactly 80 attempts in seed order. Report filtering of aborted or length-truncated attempts when comparing configurations. Derived verifier rewards can be reconstructed from the retained expected scores.

License

The artifact is provided under the MIT License. Questions from the underlying benchmark remain subject to the benchmark's applicable terms. The SuperGPQA question content is provided under ODC-BY 1.0 and requires upstream credit.

Upstream Attribution

The questions are derived from the SuperGPQA benchmark, published by the M-A-P Team. Users should credit that benchmark and its paper, "SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines" (2025), when using the question content.

Citation

Citation information will be added after anonymous review.

Xet Storage Details

Size:
4.69 kB
·
Xet hash:
225c2424b4b38436d38ab1d92b74463b427b8b021e266d6c4ce59de70fa01b23

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.