|
Download guide/02-data.md from Cion-lab/ounce100m-code: direct link, hf CLI and curl.
- Browser
- Download file 6.77 kB
-
https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/guide/02-data.md
- Command line
-
hf download hf://Cion-lab/ounce100m-code/guide/02-data.md
-
curl -L -o 02-data.md https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/guide/02-data.md
6.77 kB
| # 02 β Data | |
| The mix is the part you cannot re-run cheaply after training starts. Build it once, publish it reproducibly, | |
| and prove it is clean before a single GPU second is spent on it. | |
| ## 1. Design for abilities, not for items | |
| <sub>Sources: Phase 1 of the constitution Β· D-009</sub> | |
| Decide what the finished model must do β commonsense and physical inference, science QA, breadth of factual | |
| knowledge, factuality, arithmetic, multi-step reasoning β then reason from *what kinds of text build those | |
| abilities* to proportions. Write a rationale per source, citing the corpus property you relied on. | |
| **The line that keeps this honest:** proportions justified by corpus properties are legitimate; proportions | |
| justified by looking at evaluation questions are contamination. If a source's justification reads like a | |
| benchmark name, that's the tell. | |
| - Prefer few large, well-attributed sources over many small scraped ones. Every source is a maintenance cost | |
| and a contamination surface. | |
| - Check what a source's maintainers *already* exclude. Some curated math sets decontaminate against **test** | |
| splits only β which silently leaves validation in your train data if you don't look. | |
| - Question every modality you are told to include. This project's constitution said "general English plus math | |
| and code", and code was **dropped** (D-009): the corpus exposed whole-repository rows behind a nested struct, | |
| honest per-file token accounting was a subproject, and at 100M parameters with GSM8K being word problems, | |
| those tokens bought more elsewhere. An instruction in a plan is not a measurement. | |
| ## 2. Size the mix with headroom, then let dedup eat it | |
| Target `N` tokens in the *final* mix; assemble **1.1β1.25Γ** that, because dedup, exclusion masks and the | |
| held-out split all subtract and you cannot top up after publication. Reserve validation (~2 %) *before* | |
| sharding, with the same tokenizer, and keep it out of every training shard. | |
| ## 3. Tokenise once, pack, shard for exact positional resumption | |
| Emit raw `uint16`/`uint32` token arrays per shard plus a manifest. | |
| | Decision | What this project chose | Why, and the alternative | | |
| |---|---|---| | |
| | Document boundaries | **Contiguous packing, no separator token, no block-diagonal mask** (D-008) | The seam cost was *measured* at ~0.16 % of windows before it was accepted. A separator plus attention mask is the textbook answer and costs throughput; the wrong move is to pick either without measuring the seam rate. | | |
| | Sample length | Fixed multiple of `seq_len`; tail dropped | **Record the dropped count** β it belongs in the reproducibility proof. | | |
| | Shard size | Small enough that a job reads a few at a time | Select shards by step range and stream. A training job must never materialise the whole mix on an instance disk. | | |
| ## 4. Exact positional resumption is a data-structure problem, not a state problem | |
| `sample k` must be a pure function of `(dataset fingerprint, shuffle_seed, k)`. | |
| | Layout | Verdict | | |
| |---|---| | |
| | Shuffle the *list* of samples with a recorded seed; persist `samples_consumed` | **Works** β recompute everything else | | |
| | Shard-major, offset-major traversal; cursor = `(shard, offset)` | **Works** | | |
| | Anything relying on Python iteration order, an open file handle, unsaved `DataLoader` worker state, or *the set of files on disk at the time* | **Fails** β and fails quietly, which is worse | | |
| Persist in the checkpoint, all of them: `seq_len`, `samples_consumed`, `shuffle_seed`, `shuffle_perm_sha`, | |
| `dataset_files_sha`, `step`, `tokens_consumed`. The hashes let a resume **detect** that it is resuming into a | |
| different dataset than it was trained on β raise, never silently re-shuffle. | |
| Prove it three ways: recorded index/hash, token-offset check, and loss continuity across the break within | |
| noise. Then **resume twice in sequence** β one interruption is a fluke. Position must never drift backwards or | |
| re-read consumed data. | |
| ## 5. The contamination audit is mechanical, and its scope is narrow | |
| **Audit against `validation`/`dev` material only. `test` splits stay untouched until the benchmark phase.** | |
| This is a hard boundary, not a conservatism: reading test items β even to count overlap against them β is the | |
| one thing that can quietly invalidate the entire result set. Compute overlap without ever printing an item. | |
| Method: normalise (lowercase, strip punctuation/whitespace), count exact long-n-gram matches (13-grams are the | |
| usual standard) between mix documents and each task's **non-test** reference split, report per-source with an | |
| `overlap_total`. Drop or clean any source that fails; re-run to zero. | |
| Three traps, all of which produce a *passing* audit that means nothing: | |
| 1. **Zero denominator.** If extraction pulled 0 reference strings, `overlap_total 0` is true and worthless. | |
| Assert a minimum row count per task, and print denominators beside numerators. | |
| 2. **Wrong split.** Verify per task which split the harness actually scores. In current lm-eval several standard | |
| academic tasks are scored on `validation`, and some datasets expose no `test` at all. Auditing `test` because | |
| that's the intuitive name measures the wrong bytes *and* breaks the rule above. | |
| 3. **The audit's own artefacts leak benchmark text.** Exception strings, sample matches and debug dumps get | |
| serialised into `audit.json` and published with the mix, quoting real cells. Publish **counts**, and scrub | |
| every string that could carry an item β this happened here before it was caught. | |
| Keep the re-runnable script and the row-count-bearing report in the repo. "We decontaminated" is a claim; a | |
| script another party can run against the published mix is evidence. | |
| ## 6. Publish so a stranger can rebuild it | |
| - `manifest.json`: per-source token counts, train/val totals, shard counts and sizes, tokenizer identity (hub | |
| repo, revision, vocab size, tokenizer file hashes), the packing rule, the seed. | |
| - **The exclusion masks and dedup keys**, not just the result β otherwise the build is unreviewable. | |
| - A **content fingerprint** over the ordered file set (hash of concatenated per-file sha256s) plus the per-file | |
| sha256s. Training asserts it matches; that is what makes `dataset_files_sha` above meaningful. | |
| - Self-consistency: totals equal the sum of parts, shard counts equal the file listing, no empty shards. Write | |
| the checker; don't eyeball a table. A manifest whose numbers don't add up is worse than none, because it | |
| reads like a guarantee. | |
| Then verify the **published** artefact: re-download the manifest anonymously, re-hash shards, and feed the | |
| downloaded bytes to the same reader the trainer uses. The dataset is done when an instance with an empty cache | |
| can start from the repo id alone and read the right tokens. | |