Buckets:
| license: cc-by-4.0 | |
| size_categories: | |
| - 1M<n<10M | |
| task_categories: | |
| - other | |
| pretty_name: ARC Intermediate Solving Steps | |
| tags: | |
| - arc | |
| - arc-agi | |
| - abstract-reasoning | |
| - grid-puzzles | |
| - intermediate-steps | |
| - chain-of-thought | |
| - synthetic | |
| - procedural-generation | |
| configs: | |
| - config_name: arcgen_v1 | |
| data_files: | |
| - split: train | |
| path: data/arcgen_v1/*.jsonl | |
| - config_name: arcgen_v2 | |
| data_files: | |
| - split: train | |
| path: data/arcgen_v2/*.jsonl | |
| - config_name: rearc | |
| data_files: | |
| - split: train | |
| path: data/rearc/*.jsonl | |
| # ARC Intermediate Solving Steps (`arc-steps`) | |
| This dataset accompanies the paper [TraceViT: Grounded Trace Supervision for Visual Abstract Reasoning](https://huggingface.co/papers/2607.29586). | |
| **1,286,952 procedurally generated ARC-style records — 1,062,561 of them | |
| with intermediate *solving steps*.** Each record is an `{input, steps, output}` triple: | |
| `steps` is a sequence of intermediate grids tracing a semantically meaningful solution | |
| path from the input to the output, captured at human-annotated checkpoints of the | |
| program that produces each example. | |
| **Code:** [github.com/LiuBinnan/TraceViT](https://github.com/LiuBinnan/TraceViT) | |
|  | |
| *One record from each config — top: `arcgen_v1` task `007bbfb7` (fractal composition); | |
| middle: `arcgen_v2` task `e5790162` (path tracing); bottom: `rearc` task `1e0a9b12` | |
| (gravity, resolved per column). Leftmost frame = `input`; the frames after it are | |
| `steps`, and the last one equals `output`.* | |
| | config | tasks | records | records with steps | implementation | | |
| |---|---:|---:|---:|---| | |
| | `arcgen_v1` | 400 | 399,990 | 305,616 | ARC-GEN generators, ARC-AGI-1 training tasks | | |
| | `arcgen_v2` | 500 | 487,091 | 486,091 | ARC-GEN generators, ARC-AGI-2 training tasks | | |
| | `rearc` | 400 | 399,871 | 270,854 | RE-ARC generators + step-instrumented solver programs, ARC-AGI-1 training tasks | | |
| Together the three configs cover **900 distinct ARC tasks**. `rearc` and `arcgen_v1` | |
| share the same 400 ARC-AGI-1 task ids — two independent implementations of the same | |
| puzzles, usable as a cross-implementation contrast set. | |
| **No official ARC grids are included.** All records are synthetic; the official | |
| train/test examples were removed from every config. No validation/test split is | |
| provided — evaluate on the official ARC-AGI benchmarks. | |
| ## Schema | |
| One JSON object per line (plain `*.jsonl`, one file per config): | |
| | field | type | description | | |
| |---|---|---| | |
| | `task_id` | str | 8-hex ARC task id | | |
| | `source` | str | `"arcgen_v1"` / `"arcgen_v2"` / `"rearc"` (matches the config) | | |
| | `category` | str | `"selectable"` (task has annotated steps) or `"skip"` (no meaningful intermediates) | | |
| | `input` | int[][] | input grid, values 0-9, at most 30×30 | | |
| | `output` | int[][] | output grid, same constraints | | |
| | `steps` | int[][][] | 0..N intermediate grids; **when non-empty, the last step equals `output`**; intermediate frames may differ in size from the output | | |
| | `seed` | int \| null | generation seed (arcgen only) | | |
| | `gen_kwargs` | str \| null | generator parameters as a JSON string (arcgen only; keys vary per task) | | |
| | `palette` | int[10] \| null | global palette permutation applied at synthesis, `palette[0] == 0` (rare; arcgen only) | | |
| | `background` | int \| null | background recolor layer; `0` means untouched (arcgen only) | | |
| | `foreground_map` | str \| null | foreground color-permutation as a JSON string; `"{}"` means untouched (arcgen only) | | |
| | `difficulty_rng` | float \| null | RE-ARC native difficulty knob (rearc only) | | |
| | `difficulty_pso` | float \| null | RE-ARC native difficulty knob (rearc only) | | |
| | `rule` | str \| null | model-written natural-language description of the task's transformation (`arcgen_v2` only — see "Task rule texts" below) | | |
| Identity check for arcgen records: `background == 0 and foreground_map == "{}"` means | |
| the record kept the generator's own colors. For tasks whose colors carry rule | |
| semantics, one or both recolor layers were disabled entirely during human review; | |
| per-task `bg_layer` / `fg_layer` flags are recorded in `stats/*.per_task.json`. | |
| ## Task rule texts (`arcgen_v2`) | |
| Every `arcgen_v2` record carries a `rule` string: a natural-language description of | |
| the task's transformation, written by a language model during the program- | |
| instrumentation campaign (identical for all records of a task). Two caveats, stated | |
| plainly: | |
| - The texts are **model-generated and not individually human-verified** — treat them | |
| as auxiliary annotations, not ground truth. | |
| - They describe the task in its **canonical (official) colors**. Records whose | |
| recolor layers are non-identity (`background != 0` or `foreground_map != "{}"`) | |
| will not match the rule's color words, though the underlying transformation is the | |
| same up to the recorded color permutation. | |
| `arcgen_v1` and `rearc` records have `rule: null`. | |
| ## Usage | |
| ```python | |
| from datasets import load_dataset | |
| ds = load_dataset("lbn32/arc-steps", "arcgen_v2", split="train") | |
| # only records that carry intermediate steps | |
| with_steps = ds.filter(lambda r: len(r["steps"]) > 0) | |
| # ±steps ablation: same records, just ignore the steps column | |
| ``` | |
| Per-task record counts, per-task with-steps counts, and the drop ledger are in | |
| `stats/*.per_task.json`. | |
| ## Limitations | |
| - Per-task record counts vary: sampling runs under per-attempt and per-task time | |
| budgets, and a small number of hard-to-sample tasks fall short of the per-task | |
| target (exact counts in `stats/`). | |
| - Records are procedurally generated and validated against structural invariants, | |
| not against a per-instance solvability proof. | |
| - `rearc` excludes 129 generated records (all from task `e26a3af2`) | |
| whose grids exceeded the 30×30 ARC limit. | |
| - Steps are annotated at the task level (checkpoint positions in the program), not | |
| per-instance. | |
| ## License and attribution | |
| The dataset is released under **CC-BY-4.0** (see `LICENSE`). | |
| It is generated with, and gratefully builds on, the following projects (no upstream | |
| code and no official ARC grids are redistributed here; license copies in `licenses/`): | |
| | project | role | license | | |
| |---|---|---| | |
| | [ARC-AGI](https://github.com/fchollet/ARC-AGI) / [ARC-AGI-2](https://github.com/arcprize/ARC-AGI-2) | task definitions the generators mimic | Apache-2.0 | | |
| | [ARC-GEN](https://github.com/google/ARC-GEN) | procedural generators (`arcgen_v1`, `arcgen_v2`) | Apache-2.0 | | |
| | [RE-ARC](https://github.com/michaelhodel/re-arc) | procedural generators + solver programs (`rearc`) | MIT | |
Xet Storage Details
- Size:
- 6.58 kB
- Xet hash:
- 2bf48de01b97ce39fb842c25758df09ba529194de937a9eaae10ccd01576abd7
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.