File size: 6,767 Bytes
6302710
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
# 02 β€” Data

The mix is the part you cannot re-run cheaply after training starts. Build it once, publish it reproducibly,
and prove it is clean before a single GPU second is spent on it.

## 1. Design for abilities, not for items

<sub>Sources: Phase 1 of the constitution Β· D-009</sub>

Decide what the finished model must do β€” commonsense and physical inference, science QA, breadth of factual
knowledge, factuality, arithmetic, multi-step reasoning β€” then reason from *what kinds of text build those
abilities* to proportions. Write a rationale per source, citing the corpus property you relied on.

**The line that keeps this honest:** proportions justified by corpus properties are legitimate; proportions
justified by looking at evaluation questions are contamination. If a source's justification reads like a
benchmark name, that's the tell.

- Prefer few large, well-attributed sources over many small scraped ones. Every source is a maintenance cost
  and a contamination surface.
- Check what a source's maintainers *already* exclude. Some curated math sets decontaminate against **test**
  splits only β€” which silently leaves validation in your train data if you don't look.
- Question every modality you are told to include. This project's constitution said "general English plus math
  and code", and code was **dropped** (D-009): the corpus exposed whole-repository rows behind a nested struct,
  honest per-file token accounting was a subproject, and at 100M parameters with GSM8K being word problems,
  those tokens bought more elsewhere. An instruction in a plan is not a measurement.

## 2. Size the mix with headroom, then let dedup eat it

Target `N` tokens in the *final* mix; assemble **1.1–1.25Γ—** that, because dedup, exclusion masks and the
held-out split all subtract and you cannot top up after publication. Reserve validation (~2 %) *before*
sharding, with the same tokenizer, and keep it out of every training shard.

## 3. Tokenise once, pack, shard for exact positional resumption

Emit raw `uint16`/`uint32` token arrays per shard plus a manifest.

| Decision | What this project chose | Why, and the alternative |
|---|---|---|
| Document boundaries | **Contiguous packing, no separator token, no block-diagonal mask** (D-008) | The seam cost was *measured* at ~0.16 % of windows before it was accepted. A separator plus attention mask is the textbook answer and costs throughput; the wrong move is to pick either without measuring the seam rate. |
| Sample length | Fixed multiple of `seq_len`; tail dropped | **Record the dropped count** β€” it belongs in the reproducibility proof. |
| Shard size | Small enough that a job reads a few at a time | Select shards by step range and stream. A training job must never materialise the whole mix on an instance disk. |

## 4. Exact positional resumption is a data-structure problem, not a state problem

`sample k` must be a pure function of `(dataset fingerprint, shuffle_seed, k)`.

| Layout | Verdict |
|---|---|
| Shuffle the *list* of samples with a recorded seed; persist `samples_consumed` | **Works** β€” recompute everything else |
| Shard-major, offset-major traversal; cursor = `(shard, offset)` | **Works** |
| Anything relying on Python iteration order, an open file handle, unsaved `DataLoader` worker state, or *the set of files on disk at the time* | **Fails** β€” and fails quietly, which is worse |

Persist in the checkpoint, all of them: `seq_len`, `samples_consumed`, `shuffle_seed`, `shuffle_perm_sha`,
`dataset_files_sha`, `step`, `tokens_consumed`. The hashes let a resume **detect** that it is resuming into a
different dataset than it was trained on β€” raise, never silently re-shuffle.

Prove it three ways: recorded index/hash, token-offset check, and loss continuity across the break within
noise. Then **resume twice in sequence** β€” one interruption is a fluke. Position must never drift backwards or
re-read consumed data.

## 5. The contamination audit is mechanical, and its scope is narrow

**Audit against `validation`/`dev` material only. `test` splits stay untouched until the benchmark phase.**
This is a hard boundary, not a conservatism: reading test items β€” even to count overlap against them β€” is the
one thing that can quietly invalidate the entire result set. Compute overlap without ever printing an item.

Method: normalise (lowercase, strip punctuation/whitespace), count exact long-n-gram matches (13-grams are the
usual standard) between mix documents and each task's **non-test** reference split, report per-source with an
`overlap_total`. Drop or clean any source that fails; re-run to zero.

Three traps, all of which produce a *passing* audit that means nothing:

1. **Zero denominator.** If extraction pulled 0 reference strings, `overlap_total 0` is true and worthless.
   Assert a minimum row count per task, and print denominators beside numerators.
2. **Wrong split.** Verify per task which split the harness actually scores. In current lm-eval several standard
   academic tasks are scored on `validation`, and some datasets expose no `test` at all. Auditing `test` because
   that's the intuitive name measures the wrong bytes *and* breaks the rule above.
3. **The audit's own artefacts leak benchmark text.** Exception strings, sample matches and debug dumps get
   serialised into `audit.json` and published with the mix, quoting real cells. Publish **counts**, and scrub
   every string that could carry an item β€” this happened here before it was caught.

Keep the re-runnable script and the row-count-bearing report in the repo. "We decontaminated" is a claim; a
script another party can run against the published mix is evidence.

## 6. Publish so a stranger can rebuild it

- `manifest.json`: per-source token counts, train/val totals, shard counts and sizes, tokenizer identity (hub
  repo, revision, vocab size, tokenizer file hashes), the packing rule, the seed.
- **The exclusion masks and dedup keys**, not just the result β€” otherwise the build is unreviewable.
- A **content fingerprint** over the ordered file set (hash of concatenated per-file sha256s) plus the per-file
  sha256s. Training asserts it matches; that is what makes `dataset_files_sha` above meaningful.
- Self-consistency: totals equal the sum of parts, shard counts equal the file listing, no empty shards. Write
  the checker; don't eyeball a table. A manifest whose numbers don't add up is worse than none, because it
  reads like a guarantee.

Then verify the **published** artefact: re-download the manifest anonymously, re-hash shards, and feed the
downloaded bytes to the same reader the trainer uses. The dataset is done when an instance with an empty cache
can start from the repo id alone and read the right tokens.