Buckets:
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| context | 5,023 items | ||
| README.md | 3.13 kB xet | bd150b07 |
wikiparse 2606 data
Large-file delivery for the wikiparse 2606 technical milestone, derived from English Wikipedia
2026-06-01.
The versioned robbd/wikiparse-2606
dataset is the catalog: ten compact Parquet tables, checksums, schemas, provenance, and five small
analysis cards. This bucket carries the larger context surfaces.
Contents
| Path | Rows | Contents |
|---|---|---|
context/index/v1/ |
7,073,577 | Small article index: source key, categories, detector signals, priority, and reason. |
context/content/v1/ |
7,073,577 | Exact retained article Markdown plus the same source-bound assignment fields. |
The content release is partitioned into 1,003 disjoint 256-stream shards, with one shorter terminal shard. Each shard has one Parquet file per priority and an exact manifest; the root manifest and receipt bind every shard and file.
Priorities
| Priority | Articles | Intended use |
|---|---|---|
p0 |
908,270 | Science closure and taxa; first-pass research context. |
p1 |
4,665,480 | General retained articles. |
p2 |
1,208,553 | Useful secondary context: settlements, films, music releases, and sports people. |
p3 |
291,274 | Serial or incidental context: lists, sports seasons, and disambiguation pages. |
These labels are mutually exclusive delivery hints. They are not quality judgments, exclusion
authority, or a claim that p2/p3 are useless. Categories and detector signals remain attached so
consumers can write a different policy.
Download
hf buckets sync \
hf://buckets/robbd/wikiparse-2606-data/context/index/v1 \
./wikiparse-2606-index
hf buckets sync \
hf://buckets/robbd/wikiparse-2606-data/context/content/v1 \
./wikiparse-2606-content
# Or fetch only the science-first content slice and its receipts.
hf buckets sync \
hf://buckets/robbd/wikiparse-2606-data/context/content/v1 \
./wikiparse-2606-p0 \
--include "shards/*/p0.parquet" \
--include "shards/*/manifest.json" \
--include "manifest.json" \
--include "receipt.json"
The context files copy retained Markdown from the existing compact corpus. They do not rerun the parser, call a model, or alter the source store.
Boundary
- Technical milestone only; announcement and outreach remain held.
- Citation records are article-level and deduplicated, not inline claim markers.
- Parser-quality and detector fields are operational evidence, not adjudicated truth.
- No model weights are published here.
- Data is distributed under CC BY-SA 4.0; attribution and modification notes are in the versioned dataset.
- Total size
- 18.7 GB
- Files
- 5,024
- Last updated
- Aug 12
- Pre-warmed CDN
- US EU US EU