Data Inventory: Dolma 6T Corpus and Working Samples
This document describes where all data artifacts live. HuggingFace under HCAI-Lab/ is the primary surface for everything except the raw 6T .jsonl.zst source corpus (R2-only, ~5 TB, intentional). Modal volumes are transient caches that mirror what is already on R2/HF. PACE storage no longer holds anything sole-source — see PRESERVATION_GAP_REPORT.md for the verification audit.
It covers the deduplicated corpus, quality labels, manifests, materialized working samples, and the TrackStar attribution artifacts.
Source corpus
The source data is the Dolma 3 corpus, deduplicated to ~6T tokens across ~58K shards.
Original upstream
| Property | Value |
|---|---|
| Pool dataset | allenai/dolma3_pool (~9T tokens, 2732 shards) |
| Curated mix | allenai/dolma3_mix-6T-1025-7B |
| 150B sample | allenai/dolma3_mix-150B-1025 |
Deduplicated corpus on R2
The deduplicated corpus lives on Cloudflare R2. It was produced in two phases (SOC-127).
| Property | Value |
|---|---|
| R2 bucket | soc127-dedup |
| R2 endpoint | https://0934ab8e84ac8f4e81decaf3eb121337.r2.cloudflarestorage.com |
| Total shards | ~58,621 |
| Shard format | .jsonl.zst (zstandard-compressed JSONL) |
| Shard size | ~50-100 MB each |
| Total size | ~5 TB |
Source shard prefixes:
soc127/phase1_pool_shared/ # pool data (common_crawl, olmocr, etc.)
soc127/phase2_nonpool_final/ # non-pool sources
Each shard contains Dolma documents with fields: id, text, metadata.
Quality labels on R2 and on HF
WebOrganizer topic and format labels (SOC-91) are stored as parquet sidecars alongside the source shards. As of 2026-05-23 they are also mirrored to HF Datasets for HF-only workflows.
| Property | Value |
|---|---|
| R2 prefix | soc91-labels/ (58,465 per-shard parquets + .stats.json sidecars) |
| HF Dataset | HCAI-Lab/soc91-labels (private; 2,719 parquet chunks, 169.94 GB; 58,465 source files concatenated into ~60 MB chunks; each row carries source_shard_path) |
| Format | .parquet per source shard |
| Labels | 24 topics x 24 formats (576 bins) |
| Stats on R2 | soc91-stats/ prefix |
| Stats on HF | HCAI-Lab/soc91-stats (private; raw EDA/audit JSON outputs) |
Topic categories: adult_content, art_and_design, crime_and_law, education_and_jobs,
electronics_and_hardware, entertainment, fashion_and_beauty, finance_and_business,
food_and_dining, games, health, history_and_geography, home_and_hobbies,
industrial, literature, politics, religion, science_math_and_technology,
social_life, software, software_development, sports_and_fitness, transportation,
travel_and_tourism.
Format categories: academic_writing, content_listing, creative_writing,
customer_support, comment_section, faq, truncated, knowledge_article,
legal_notices, listicle, news_article, nonfiction_writing, about_org,
news_org, about_pers, personal_blog, product_page, q_a_forum, spam_ads,
structured_data, documentation, audio_transcript, tutorial, user_review.
Quality sidecars on R2 and on HF
Per-document quality classifier scores (SOC-139) — separate from the WebOrganizer
topic/format labels above. Includes the SOC-142 inverted-label fix (commit 3342baf).
| Property | Value |
|---|---|
| R2 prefix | soc139-quality-sidecars/ (58,621 per-shard parquets + .label-fix.done + .stats.json) |
| HF Dataset | HCAI-Lab/soc139-quality-sidecars (private; 1.26B rows, 41.8 GB, 80 parquet chunks) |
| HF columns | doc_id, quality_label_id, quality_score, quality_high_prob, quality_low_prob, quality_confidence, source_shard_path |
Unified corpus manifest on R2 and on HF
The full corpus manifest (SOC-95) contains one row per document with doc_id, token count, shard path, and WebOrganizer labels. This is the index used to draw stratified samples.
| Property | Value |
|---|---|
| R2 prefix | soc95-manifest/data/ (per-shard parquets) |
| HF Dataset (primary) | HCAI-Lab/dolma3-olmo3-corpus-manifest (1.1B rows, 32-col schema with topic+format+quality+token+shard) |
| HF Dataset (alt) | HCAI-Lab/dolma3-6t-corpus-manifest |
| Format | Parquet files (~58K files, one per shard) |
| Columns | doc_id, token_count, shard_path, weborganizer_topic, weborganizer_format (plus quality fields in the unified HF version) |
| Total size | ~213 GB |
Modal persistent volumes (transient caches)
The following Modal volumes cache data for fast access from Modal workers.
These are in the eilab-gt Modal workspace and mirror R2/HF sources — none is sole-source.
Safe to delete if Modal compute is no longer needed.
| Volume | Contents | Size |
|---|---|---|
soc134-corpus-cache |
All ~58K source shards from R2 (mirrors R2 key structure) | ~5 TB |
soc134-manifest-cache |
All ~58K manifest parquet files from R2 | ~213 GB |
soc134-output-sample-500-docs |
Materialized docs for sample_500_docs | varies |
soc134-output-sample-1000-docs |
Materialized docs for sample_1000_docs | varies |
soc134-output-sample-5000-docs |
Materialized docs for sample_5000_docs | varies |
soc134-output-sample-10000-docs |
Materialized docs for sample_10000_docs | varies |
soc134-output-dolma3-6t-preconditioner-100k |
Materialized docs for preconditioner_100k (uniform random) | varies |
Corpus cache volume layout
soc134-corpus-cache/
soc127/phase1_pool_shared/common_crawl/part_000/data__...__shard_00000001.jsonl.zst
soc127/phase1_pool_shared/common_crawl/part_000/data__...__shard_00000002.jsonl.zst
soc127/phase2_nonpool_final/...
...
Preserves the R2 key structure. shard_path values from manifests resolve directly as paths.
Output volume layout
Each per-sample output volume has the same structure:
soc134-output-sample-{N}-docs/
{run_id}/
manifest.parquet
sample_contract.json
bin_summary.csv
worker_0000/
soc127__phase1_pool_shared__...jsonl.zst # materialized doc subset
soc127__phase1_pool_shared__...jsonl.zst.stats.json
soc127__phase1_pool_shared__...jsonl.zst.done
worker_0001/
...
worker_0127/
...
.hf_upload/
metadata.done
batch_0000.done
batch_0001.done
...
Run IDs for current materialization:
| Sample | Run ID |
|---|---|
| sample_500_docs | soc134_materialize_20260320_150348 |
| sample_1000_docs | soc134_materialize_20260320_151802 |
| sample_5000_docs | soc134_materialize_20260320_152205 |
| sample_10000_docs | soc134_materialize_20260320_152236 |
HuggingFace datasets
Working samples are published to HuggingFace under the HCAI-Lab organization. All six stratified samples are complete; each is published as both an HF Dataset (viewer + snapshot_download) and an HF Bucket of the same name (S3-style access via hf buckets sync); bit-identical data on both surfaces.
| HF repo | Sample | Docs | Tokens | Underfilled bins | Status |
|---|---|---|---|---|---|
HCAI-Lab/dolma3-6t-sample-500-docs |
500 docs/bin | 287,936 | 538.8M | 1 | Complete (43,129 files) |
HCAI-Lab/dolma3-6t-sample-1000-docs |
1,000 docs/bin | 575,187 | 1.08B | 3 | Complete (46,808 files) |
HCAI-Lab/dolma3-6t-sample-5000-docs |
5,000 docs/bin | 2,855,446 | 5.30B | 7 | Complete (56,043 files) |
HCAI-Lab/dolma3-6t-sample-10000-docs |
10,000 docs/bin | 5,678,621 | 10.5B | 17 | Complete (57,845 files) |
HCAI-Lab/dolma3-6t-sample-50000-docs |
50,000 docs/bin | 26,249,124 | 62.8B | 79 | Complete (58,246 files) |
HCAI-Lab/dolma3-6t-sample-100000-docs |
100,000 docs/bin | 49,709,294 | 118.4B | 130 | Complete (58,264 files, dataset twin added 2026-05-25) |
Preconditioner sample (uniform random, no stratification):
| Location | Repo | Docs | Tokens | Min tokens | Files | Status |
|---|---|---|---|---|---|---|
| HF Dataset | HCAI-Lab/dolma3-6t-preconditioner-100k |
100,000 | 251.5M | 512 | 38,277 | Complete |
| HF Bucket | HCAI-Lab/dolma3-6t-preconditioner-100k |
100,000 | 251.5M | 512 | 38,260 | Complete |
The HF Dataset repo is the primary access path. Use snapshot_download or hf download to fetch it.
HF repo layout
HCAI-Lab/dolma3-6t-sample-{N}-docs/
README.md
sample_contract.json
bin_summary.csv
working_sample_manifest.parquet
data/
part_000/
soc127__phase1_pool_shared__...jsonl.zst
soc127__phase1_pool_shared__...jsonl.zst
... # up to 5,000 files per part
part_001/
...
part_NNN/
...
Data files are partitioned into data/part_NNN/ subdirectories with max 5,000 files each
(HuggingFace enforces a 10,000 files per directory limit).
HF repo metadata files
working_sample_manifest.parquet: One row per sampled document. Columns:doc_id,token_count,shard_path,bin_id,bin_topic,bin_format.sample_contract.json: Aggregate sampling parameters and realized counts.bin_summary.csv: Per-bin fill rates across 576 bins (24 topics x 24 formats).README.md: Data card with parameters, counts, and format description.
Working sample design
Samples are drawn by stratified sampling across the 576 topic x format bins.
Each sample config specifies a number of documents per bin (docs_per_bin).
Within each bin, documents are selected by deterministic priority scoring
(blake2b hash of doc_id:seed), ensuring reproducibility.
| Sample | docs/bin | Total docs | Total tokens | Unique shards | Underfilled bins |
|---|---|---|---|---|---|
| sample_500_docs | 500 | 287,936 | 538.8M | ~42,461 | 1 |
| sample_1000_docs | 1,000 | 575,187 | 1.08B | ~45,926 | 3 |
| sample_5000_docs | 5,000 | 2,855,446 | 5.30B | ~55,347 | 7 |
| sample_10000_docs | 10,000 | 5,678,621 | 10.5B | ~57,622 | 17 |
| sample_50000_docs | 50,000 | 26,249,124 | 62.8B | ~58,200 | 79 |
| sample_100000_docs | 100,000 | 49,709,294 | 118.4B | ~58,264 | 130 |
All samples use seed 42. Bins are "underfilled" when the source corpus has fewer documents in that topic x format combination than the requested docs_per_bin.
Sample manifest columns
| Column | Type | Description |
|---|---|---|
doc_id |
string | Unique document identifier from the Dolma corpus |
token_count |
int | Token count for this document |
shard_path |
string | R2 key of the source shard containing this document |
bin_id |
int | 1-indexed bin number (1-576) |
bin_topic |
string | WebOrganizer topic label |
bin_format |
string | WebOrganizer format label |
TrackStar attribution artifacts in cloud
All TrackStar artifacts produced by SOC-156, SOC-161, SOC-170, SOC-171 have been migrated to HuggingFace as part of the 2026-05-22 PACE-to-cloud handoff. See PRESERVATION_GAP_REPORT.md for the migration audit.
Bergson training gradient index
The reusable per-document gradient index built with Bergson from sample_10000_docs (5.68M docs).
| Property | Value |
|---|---|
| HF Bucket | HCAI-Lab/trackstar-gradient-index-base (private) |
| Total size | 1.30 TB |
| Total files | 3,792 |
| Shards | 316 (shard_0000 .. shard_0315) |
| Per-shard size | ~3.9 GB (gradients.bin) + metadata |
| Model | allenai/Olmo-3-1025-7B (base) |
| Projection dim | 16, fp32 |
Per-shard layout: data.hf/ (HF Dataset with length/loss), gradients.bin, info.json, index_config.json, normalizers.pth, preconditioners.pth, preconditioners_eigen.pth, preprocess_config.json, processor_config.json, total_processed.pt.
TrackStar training shards (positional doc IDs)
Plain JSONL with {id, text} mapping shard_NNNN:INDEX positional doc IDs (used in score matrices) back to source document UUIDs.
| Property | Value |
|---|---|
| HF Dataset | HCAI-Lab/dolma3-6t-sample-10000-docs-trackstar-shards (private) |
| Total size | 41 GB (HF accounting: 46 GB) |
| Files | 316 shard_NNNN.jsonl (~145 MB each, ~18K docs per shard) |
Required if you load any of the score matrices below — without these, positional IDs cannot be resolved to text.
Query gradient indices
Per-benchmark query gradients, three model variants (base, instruct_base, instruct_cot).
| Property | Value |
|---|---|
| HF Bucket | HCAI-Lab/trackstar-query-gradients-base (private) |
| Total size | 17.9 GB |
| Files | 336 (across SOC-156, SOC-161, SOC-170, SOC-171 runs) |
Canonical SOC-156 builds live under base/20260326T163642Z_1102443/queries_{gsm8k,mmlu_social_science,mmlu_stem,socialiqa}/.
Per-query score matrices
Influence(doc_i, query_j) for the four OLMES benchmarks. Multiple model and benchmark variants:
| HF Bucket | Run | Files | Size |
|---|---|---|---|
HCAI-Lab/trackstar-scores-base-olmes-4bench |
OLMo-3-7B base, 4 OLMES benchmarks | 2,532 | 396 GB |
HCAI-Lab/trackstar-scores-instruct-cot-olmes-4bench |
OLMo-3-7B instruct-cot, 4 OLMES benchmarks | 2,532 | 396 GB |
HCAI-Lab/trackstar-scores-base-bbh |
OLMo-3-7B base, BBH | 1,899 | 14 GB |
HCAI-Lab/trackstar-scores-instruct-base-bbh |
OLMo-3-7B instruct, BBH | 1,899 | 14 GB |
HCAI-Lab/trackstar-scores-base-gsm8k-arc |
OLMo-3-7B base, GSM8K+ARC | 1,899 | 79 GB |
HCAI-Lab/trackstar-scores-instruct-base-gsm8k-arc |
OLMo-3-7B instruct, GSM8K+ARC | 1,899 | 79 GB |
HCAI-Lab/trackstar-scores-instruct-cot-noprecond-mmlu-ss |
OLMo-3-7B instruct-cot, MMLU social science only, no preconditioner (ablation pair to soc161) | 633 | 70 GB |
Format: per-benchmark dir containing 316 shard_NNNN.npy + 316 shard_NNNN_doc_ids.json + 1 query_ids.json (633 files × N benchmarks per bucket). See docs/TRACKSTAR_DATA_ARTIFACTS.md §1 for details.
Preconditioners
Pre-built TrackStar mixed preconditioners for three model variants.
| Property | Value |
|---|---|
| HF Bucket | HCAI-Lab/trackstar-preconditioners (public) |
| Total size | 885 MB |
| Files | 78 across olmo-3-1025-7b/, olmo-3-7b-instruct/, olmo-3-7b-think/ |
Note: the preconditioner must match the model used to build the gradients (SOC-162 finding).
Top-K influence results
Ranked top-K influence lists, much smaller than full score matrices.
| HF Dataset | Contents |
|---|---|
HCAI-Lab/dolma3-trackstar-influence-scores (private) |
influence_scores_full.parquet + top2k_{gsm8k,mmlu_socsci,mmlu_stem,socialiqa}.{csv,parquet} |
HCAI-Lab/trackstar-top2k-base-gsm8k-arc (bucket) |
SOC-171 top-2K for base variant |
HCAI-Lab/trackstar-top2k-instruct-base-gsm8k-arc (bucket) |
SOC-171 top-2K for instruct variant |
Job archive (Tier 3)
Tarballs of SLURM logs, sbatch artifacts, dedup work state, SOC-149 / SOC-159 / SOC-167 outputs, miscellaneous analysis. For reproducibility audits, not for day-to-day use.
| HF Dataset | Size | Files |
|---|---|---|
HCAI-Lab/dolma3-attribution-job-archive (private) |
12.87 GB | 17 (16 tarballs + README) |
Data home
A consolidated lookup for "where do I find X?" — README + machine-readable JSON inventory.
| HF Dataset | Mirror in repo |
|---|---|
HCAI-Lab/dolma3-data-attribution-index (private) |
docs/data_home/ |
Also published as an HF Collection: "Dolma3 Data Attribution" under HCAI-Lab/.
Local repo data
Sample manifests and contracts are checked into the repo:
data/samples/
sample_500_docs/
working_sample_manifest.parquet (14 MB)
sample_contract.json
bin_summary.csv
sample_1000_docs/
working_sample_manifest.parquet (29 MB)
sample_contract.json
bin_summary.csv
sample_5000_docs/
working_sample_manifest.parquet (141 MB)
sample_contract.json
bin_summary.csv
sample_10000_docs/
working_sample_manifest.parquet (280 MB)
sample_contract.json
bin_summary.csv
Document format
Each materialized JSONL.zst file contains a subset of Dolma documents. Each line is a JSON object with at minimum:
{
"id": "dolma-document-id",
"text": "full document text...",
"metadata": { ... }
}
The id field corresponds to doc_id in the manifest. The text field contains
the full document content. The metadata object varies by source.
Access patterns
Reading a specific sample
To read materialized documents for a sample:
From HuggingFace (once upload completes):
from datasets import load_dataset ds = load_dataset("HCAI-Lab/dolma3-6t-sample-500-docs")From Modal volume (for Modal workers):
volume = modal.Volume.from_name("soc134-output-sample-500-docs") # mount at /data in your function, read worker dirsFrom the manifest (to look up which docs are in a sample without downloading):
import pandas as pd manifest = pd.read_parquet("data/samples/sample_500_docs/working_sample_manifest.parquet")
Reading source shards
For the materialized stratified samples, prefer the HF dataset directly — hf_hub_download or snapshot_download against HCAI-Lab/dolma3-6t-sample-{N}-docs already exposes the per-doc JSONL files. Reading raw upstream shards is only required when you want documents outside any sample.
From HF dataset (materialized sample) — recommended for sample work:
from huggingface_hub import hf_hub_download path = hf_hub_download( "HCAI-Lab/dolma3-6t-sample-500-docs", "data/part_000/soc127__phase1_pool_shared__...jsonl.zst", repo_type="dataset", )From Modal corpus cache (fastest, for Modal workers):
# mounted at /corpus shard_path = "soc127/phase1_pool_shared/common_crawl/part_000/data__...__shard_00000001.jsonl.zst" data = Path("/corpus") / shard_pathFrom R2 direct — only for raw upstream shards outside the materialized samples:
from dolma.quality.r2 import R2Config, create_r2_client, download_object_bytes config = R2Config.from_env(output_prefix="any") client = create_r2_client(config) data = download_object_bytes(client, bucket="soc127-dedup", key=shard_path)
R2 credentials (only needed for raw source shards or per-shard sidecar parquets)
R2 hosts the raw 6T .jsonl.zst source corpus and the original per-shard sidecars. Most workflows can stay HF-only since the sidecar mirrors are on HF (HCAI-Lab/soc91-labels, soc91-stats, soc139-quality-sidecars). R2 credentials are stored in 1Password. For local scripts:
bash scripts/bootstrap/with_r2_credentials.sh <command>
For Modal workers: credentials are in the r2-credentials Modal secret.
Cleanup notes
soc134-materialize-cache: Legacy shared volume that hit the 500K inode limit. Contains partial data from the failed multi-sample run. Safe to delete after per-sample volumes are verified.soc134-corpus-cache: Keep as long as materialization runs are needed. Delete when the project phase is complete (~$750/month for 5TB).soc134-manifest-cache: Keep for future sample draws. Small (~213GB).
Xet Storage Details
- Size:
- 19.4 kB
- Xet hash:
- 2f796d0ed617df28faa750b31ee643f5ce629c9ce5a75cc63d0156d42f11467b
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.