Buckets:
wikiparse 2606 data
Large-file delivery for the wikiparse 2606 technical milestone, derived from English Wikipedia
2026-06-01.
The versioned robbd/wikiparse-2606
dataset is the catalog: ten compact Parquet tables, checksums, schemas, provenance, and five small
analysis cards. This bucket carries the larger context surfaces.
Contents
| Path | Rows | Contents |
|---|---|---|
context/index/v1/ |
7,073,577 | Small article index: source key, categories, detector signals, priority, and reason. |
context/content/v1/ |
7,073,577 | Exact retained article Markdown plus the same source-bound assignment fields. |
The content release is partitioned into 1,003 disjoint 256-stream shards, with one shorter terminal shard. Each shard has one Parquet file per priority and an exact manifest; the root manifest and receipt bind every shard and file.
Priorities
| Priority | Articles | Intended use |
|---|---|---|
p0 |
908,270 | Science closure and taxa; first-pass research context. |
p1 |
4,665,480 | General retained articles. |
p2 |
1,208,553 | Useful secondary context: settlements, films, music releases, and sports people. |
p3 |
291,274 | Serial or incidental context: lists, sports seasons, and disambiguation pages. |
These labels are mutually exclusive delivery hints. They are not quality judgments, exclusion
authority, or a claim that p2/p3 are useless. Categories and detector signals remain attached so
consumers can write a different policy.
Download
hf buckets sync \
hf://buckets/robbd/wikiparse-2606-data/context/index/v1 \
./wikiparse-2606-index
hf buckets sync \
hf://buckets/robbd/wikiparse-2606-data/context/content/v1 \
./wikiparse-2606-content
# Or fetch only the science-first content slice and its receipts.
hf buckets sync \
hf://buckets/robbd/wikiparse-2606-data/context/content/v1 \
./wikiparse-2606-p0 \
--include "shards/*/p0.parquet" \
--include "shards/*/manifest.json" \
--include "manifest.json" \
--include "receipt.json"
The context files copy retained Markdown from the existing compact corpus. They do not rerun the parser, call a model, or alter the source store.
Boundary
- Technical milestone only; announcement and outreach remain held.
- Citation records are article-level and deduplicated, not inline claim markers.
- Parser-quality and detector fields are operational evidence, not adjudicated truth.
- No model weights are published here.
- Data is distributed under CC BY-SA 4.0; attribution and modification notes are in the versioned dataset.
Xet Storage Details
- Size:
- 3.13 kB
- Xet hash:
- bd150b070cee83f293b5da2c644bc180f432b4b883a88dd1128c4e02322fb368
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.