Download guide/02-data.md from Cion-lab/ounce100m-code: direct link, hf CLI and curl.
- Browser
- Download file 6.77 kB
-
https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/guide/02-data.md
- Command line
-
hf download hf://Cion-lab/ounce100m-code/guide/02-data.md
-
curl -L -o 02-data.md https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/guide/02-data.md
02 β Data
The mix is the part you cannot re-run cheaply after training starts. Build it once, publish it reproducibly, and prove it is clean before a single GPU second is spent on it.
1. Design for abilities, not for items
Sources: Phase 1 of the constitution Β· D-009
Decide what the finished model must do β commonsense and physical inference, science QA, breadth of factual knowledge, factuality, arithmetic, multi-step reasoning β then reason from what kinds of text build those abilities to proportions. Write a rationale per source, citing the corpus property you relied on.
The line that keeps this honest: proportions justified by corpus properties are legitimate; proportions justified by looking at evaluation questions are contamination. If a source's justification reads like a benchmark name, that's the tell.
- Prefer few large, well-attributed sources over many small scraped ones. Every source is a maintenance cost and a contamination surface.
- Check what a source's maintainers already exclude. Some curated math sets decontaminate against test splits only β which silently leaves validation in your train data if you don't look.
- Question every modality you are told to include. This project's constitution said "general English plus math and code", and code was dropped (D-009): the corpus exposed whole-repository rows behind a nested struct, honest per-file token accounting was a subproject, and at 100M parameters with GSM8K being word problems, those tokens bought more elsewhere. An instruction in a plan is not a measurement.
2. Size the mix with headroom, then let dedup eat it
Target N tokens in the final mix; assemble 1.1β1.25Γ that, because dedup, exclusion masks and the
held-out split all subtract and you cannot top up after publication. Reserve validation (~2 %) before
sharding, with the same tokenizer, and keep it out of every training shard.
3. Tokenise once, pack, shard for exact positional resumption
Emit raw uint16/uint32 token arrays per shard plus a manifest.
| Decision | What this project chose | Why, and the alternative |
|---|---|---|
| Document boundaries | Contiguous packing, no separator token, no block-diagonal mask (D-008) | The seam cost was measured at ~0.16 % of windows before it was accepted. A separator plus attention mask is the textbook answer and costs throughput; the wrong move is to pick either without measuring the seam rate. |
| Sample length | Fixed multiple of seq_len; tail dropped |
Record the dropped count β it belongs in the reproducibility proof. |
| Shard size | Small enough that a job reads a few at a time | Select shards by step range and stream. A training job must never materialise the whole mix on an instance disk. |
4. Exact positional resumption is a data-structure problem, not a state problem
sample k must be a pure function of (dataset fingerprint, shuffle_seed, k).
| Layout | Verdict |
|---|---|
Shuffle the list of samples with a recorded seed; persist samples_consumed |
Works β recompute everything else |
Shard-major, offset-major traversal; cursor = (shard, offset) |
Works |
Anything relying on Python iteration order, an open file handle, unsaved DataLoader worker state, or the set of files on disk at the time |
Fails β and fails quietly, which is worse |
Persist in the checkpoint, all of them: seq_len, samples_consumed, shuffle_seed, shuffle_perm_sha,
dataset_files_sha, step, tokens_consumed. The hashes let a resume detect that it is resuming into a
different dataset than it was trained on β raise, never silently re-shuffle.
Prove it three ways: recorded index/hash, token-offset check, and loss continuity across the break within noise. Then resume twice in sequence β one interruption is a fluke. Position must never drift backwards or re-read consumed data.
5. The contamination audit is mechanical, and its scope is narrow
Audit against validation/dev material only. test splits stay untouched until the benchmark phase.
This is a hard boundary, not a conservatism: reading test items β even to count overlap against them β is the
one thing that can quietly invalidate the entire result set. Compute overlap without ever printing an item.
Method: normalise (lowercase, strip punctuation/whitespace), count exact long-n-gram matches (13-grams are the
usual standard) between mix documents and each task's non-test reference split, report per-source with an
overlap_total. Drop or clean any source that fails; re-run to zero.
Three traps, all of which produce a passing audit that means nothing:
- Zero denominator. If extraction pulled 0 reference strings,
overlap_total 0is true and worthless. Assert a minimum row count per task, and print denominators beside numerators. - Wrong split. Verify per task which split the harness actually scores. In current lm-eval several standard
academic tasks are scored on
validation, and some datasets expose notestat all. Auditingtestbecause that's the intuitive name measures the wrong bytes and breaks the rule above. - The audit's own artefacts leak benchmark text. Exception strings, sample matches and debug dumps get
serialised into
audit.jsonand published with the mix, quoting real cells. Publish counts, and scrub every string that could carry an item β this happened here before it was caught.
Keep the re-runnable script and the row-count-bearing report in the repo. "We decontaminated" is a claim; a script another party can run against the published mix is evidence.
6. Publish so a stranger can rebuild it
manifest.json: per-source token counts, train/val totals, shard counts and sizes, tokenizer identity (hub repo, revision, vocab size, tokenizer file hashes), the packing rule, the seed.- The exclusion masks and dedup keys, not just the result β otherwise the build is unreviewable.
- A content fingerprint over the ordered file set (hash of concatenated per-file sha256s) plus the per-file
sha256s. Training asserts it matches; that is what makes
dataset_files_shaabove meaningful. - Self-consistency: totals equal the sum of parts, shard counts equal the file listing, no empty shards. Write the checker; don't eyeball a table. A manifest whose numbers don't add up is worse than none, because it reads like a guarantee.
Then verify the published artefact: re-download the manifest anonymously, re-hash shards, and feed the downloaded bytes to the same reader the trainer uses. The dataset is done when an instance with an empty cache can start from the repo id alone and read the right tokens.