File size: 23,664 Bytes
f345921
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
# 00 β€” Platform notes (Phase 0)

Observed mechanics of the two execution environments this project spans: this workspace and the
Kaggle image. Everything here was **measured on 2026-09-19**, not recalled. Where a claim contradicts
prior documentation (including the Kaggle skill's own notes), the contradiction is called out, because
a future session will otherwise trust the older statement.

Raw probe JSON lives in the run logs of the kernels listed in `memory/ASSETS.md`.

---

## 1. The two environments are not the same machine

| | this workspace | Kaggle CPU session | Kaggle GPU session (2xT4) |
|---|---|---|---|
| OS | Windows 11, git-bash | Linux 6.12.90+, glibc 2.35 | Linux 6.12.90+, glibc 2.35 |
| Python | 3.11 (and 3.14) | **3.12.13** | **3.12.13** |
| Docker image | n/a | `gcr.io/kaggle-images/python@sha256:dafd4ce5…c40b9` | `gcr.io/kaggle-gpu-images/python@sha256:37c64f7d…7d461` |
| torch | β€” | **2.10.0+cpu** | **2.10.0+cu128** |
| internet | β€” | yes (needs `enableInternet: true`) | yes (needs `enableInternet: true`) |
| GPU | none | `device_count()==0` | **2x Tesla T4, cc 7.5, 40 SM, 14.56 GiB usable each** |

**Consequence:** CPU and GPU sessions are *different images*, not one image with a GPU attached. The
CPU image's `torch` is a `+cpu` build, so a script that imports `torch.cuda` will work on the GPU
shape and silently do nothing useful on the CPU shape. Pin these digests in any job that must be
reproducible; a "latest" Kaggle image can move under the run.

`huggingface_hub` is **1.32.0 locally** and reported as `huggingface-hub 1.11.0` inside the Kaggle
image. Local code and in-job code are therefore on different minor versions β€” an API that exists
locally may not exist in the job (`create_repo(tags=…)` is one such example; it raised
`TypeError` locally). Check the target environment's version before using a new API.

## 2. Job mechanics that the main run will depend on

- **Submission is one call.** `save_notebook` with `kernelExecutionType: "SaveAndRunAll"` creates
  *and* runs. `kernelType: "script"` and `machineShape` must both be present or the call fails with a
  message-less error. `machineShape: "GPU"` + `enableGpu: true` is accepted and yields 2xT4 β€” no
  T4-specific shape string is needed, and none should be invented.
- **`newTitle` rewrites the slug.** Kept `newTitle` equal to the slug suffix in every probe so the
  returned `ref` matched what was requested. Still: **poll the returned slug, never the requested one.**
- **Latency.** A ~30 s CPU script went submit β†’ `COMPLETE` in ~85 s; a ~58 s GPU script in ~145 s.
  So budget **~60–90 s of queue+boot per job**, and don't poll sooner than ~60 s. Both figures are
  single samples from an uncongested account at ~15:20 UTC.
- **Where output lands.** `list_notebook_session_output` returns `files[]` (signed URLs, expiry
  unknown β€” download promptly) and `log`, a JSON array of `{stream_name, time, data}` where `time` is
  **seconds since session start**, one array element per line. The log carries the whole stdout, so a
  probe that prints its result as JSON needs no file download. **Only files written under

  `/kaggle/working` are returned** β€” the probes deliberately wrote nothing to `/tmp`, because `/tmp`
  looks enormous (Β§4) and is therefore the easiest place to silently lose an artifact.
- **Kernel logs are not secret-safe.** Anything printed is retained in the kernel revision. No token
  may be printed, and by extension no token may be *in the kernel source*, since source and logs are
  both Kaggle-held.
- **`CUDA_VISIBLE_DEVICES` is unset in both shapes** β€” it is absent from the environment entirely, not
  set to the string `"None"`. The Kaggle skill documents `CUDA_VISIBLE_DEVICES=None` for CPU runs;
  that is wrong for this image (E-002 in `memory/ERRORS.md`), and code that branches on
  `== "None"` will take the GPU branch on a CPU box. Gate on `torch.cuda.is_available()` or
  `torch.cuda.device_count()` instead.

## 3. Quota accounting

- Readout is **seconds**: `total_time_allowed: "108000s"` = 30 h, `quota_refresh_time:
  2026-09-26T00:00:00Z`. Weekly reset is **Saturday 00:00 UTC**.
- CPU sessions cost **zero** GPU quota, confirmed: two CPU probes ran and `time_used` stayed `0s`.
- A 2xT4 session charged **66.411 s** for **58.4 s** of script time β†’ **~1x wall-clock, not 2x**.
  Full reasoning, the caveat, and the second sample pending in `memory/QUOTA.md`. This is the single
  most schedule-relevant number measured in Phase 0: at 1x, 30 h/week is 30 wall-clock hours with
  both cards, not the 15 that Β§2's "shared across both cards" phrasing suggests.

## 4. Disk and memory β€” the assumption in Β§2 is only half right

| Mount | Total | Free | Notes |
|---|---|---|---|
| `/kaggle/working`, `/kaggle/input` | **19.5 GB** | 19.5 GB | the small Kaggle disk; **only this is retrieved as output** |
| `/dev/loop1` β†’ `/kaggle/src` | 20 GB | 20 GB | separate loop device |
| `/`, `/tmp`, `/usr`, `/root/.cache` (overlay) | 7.9 TB | **1.1 TB** | **shared host filesystem, 88 % used by other tenants** |
| `/dev/shm` | 14 GB | 14 GB | useful for DataLoader worker traffic |
| RAM | 32 GB (`cgroup memory.max` = 30 GiB) | 31.2 GB available | `SwapTotal: 0` |
| CPU | 4 vCPU (`cpu.max` = `400000 100000`), Intel Xeon @ 2.20 GHz | | |

**Read this carefully before planning the checkpoint cycle.** Β§2 says "Kaggle disk is small", and for
`/kaggle/working` that is exactly true: 19.5 GB. But the root overlay reports ~1 TB free, and a naive
optimisation would write caches or staged shards there. Two reasons not to:

1. The overlay is a **shared host volume already at 88 %**, so "1 TB free" is other tenants' headroom
   and can disappear mid-run. Β§3.13's "never let it fill the disk" must be enforced against a number
   we do not control.
2. Nothing outside `/kaggle/working` is retrievable when the session ends, so it is not storage in
   any sense that matters.

Default `HF_HOME`/datasets cache resolves under `/root/.cache`, i.e. the overlay, not the 19.5 GB
volume. So **a training job's apparent free space can be large while `/kaggle/working` fills up** β€”
the failure mode Β§3.13 warns about is invisible to a naive `disk_usage("/")` check. **Every job must

measure the specific filesystem it writes to, and log free space for that path.** On the CPU shape the
effective budget is ~19.5 GB; treat that as the design number for shards resident at once.

No swap: a tokenizer build or a large in-memory dedup on a 30 GiB cgroup must be sized, not streamed
optimistically.

## 5. Network behaviour, including one thing that does not resolve

Working, measured from inside a GPU and a CPU session:
`huggingface.co/api/models` β†’ 200 in 149 ms; `pypi.org/simple/` β†’ 200 in 73 ms;
`datasets/…/resolve/main/README.md` β†’ 200 via `api/resolve-cache/…`;
`datasets.load_dataset("wikimedia/wikipedia", "20231101.en", streaming=True)` β†’ first row in **7.4 s**,
unauthenticated. `pip install` from PyPI works: `xxhash` in **5.8 s**, import verified.

**Does not resolve:** `cdn-lfs.huggingface.co`, `cdn-lfs-us-1.huggingface.co`, `xethub.huggingface.co`
β†’ `gaierror`. That looked alarming, so probe D tested it directly by pulling 300 MB from two large
public parquet files.

**Sustained Hub download throughput from inside a session: 72.8 MB/s** (`wikimedia/wikipedia`, 745 MB
file) **and 88.9 MB/s** (`HuggingFaceFW/fineweb-edu`, 2334 MB file), with 0.8–0.94 s time-to-first-byte.
So the non-resolving LFS hostnames do not block bulk transfer β€” `resolve/` redirects land on hosts that
do resolve. **Ingest of the mix is not network-bound:** at a conservative 70 MB/s, one billion tokens
(β‰ˆ3.6 GB of raw text at ~3.6 chars/token, or ~4–6 GB as parquet) downloads in single-digit minutes.
Phase 2 should plan shard staging around CPU and disk limits, not bandwidth.

Also measured in the same probe: `datasets` streaming at **139.5 rows/s β‰ˆ 3.0 MB of text per second**
single-process, unauthenticated (83 rows/s / 1.8 MB/s cold, ~3 MB/s warm). Unauthenticated Hub access
works but the API warns about rate limits (`Please set a HF_TOKEN`); at multi-hundred-shard scale that
is a plausible 429 source, which is an argument for the token-in-job path (D-002) for the *data plane*
even if the *checkpoint plane* stays local.

Enumerating repo files without a client library: `GET https://huggingface.co/api/datasets/<id>/tree/main?recursive=true`
returns `path`/`size`/`type` and works with plain `urllib`. It appeared to cap at 1000 entries
(`fineweb-edu` returned exactly 1000), so **assume pagination is required** for a large repo.

### CPU-side ingest throughput, and what it rules out

Measured on the Kaggle CPU shape (4 vCPU Xeon @ 2.2 GHz) with the gpt2 tokenizer over 6 MB of real
Wikipedia English, `tokenizers 0.22.2` `Tokenizer.encode_batch`:

| cores | tokens/s | hours to tokenize 1B tokens |
|---|---|---|
| 1 | 245,660 | 1.13 |
| 2 | 492,191 | 0.56 |
| 4 | **660,943** | **0.42** |

`chars_per_token` 4.53 on this corpus. Scaling is near-linear to 2 cores and then sub-linear
(2.7x at 4 cores), which is what a GIL-bound parent feeding rayon workers looks like β€” so a Phase 2
pipeline that shards by **process**, not by thread, should recover the rest. `datasets` streaming ran
at 151.6 rows/s β‰ˆ **3.26 MB of text per second** single-process, unauthenticated.

**The conclusion is a negative result, which is the useful kind:** tokenizing 1B tokens costs ~0.4 h of
the 4-core budget, and at 35–89 MB/s the raw bytes arrive faster than `datasets` can stream them.
**Phase 2 is therefore not CPU-bound and not bandwidth-bound β€” it is bound by dedup memory and by the

30 GiB cgroup with no swap.** Budget the design accordingly: the dedup/audit stage must be written to
stream and shard against fixed RAM, not to hold a billion sketches, and the only real wall-clock cost
in building the mix is the exact-duplicate/n-gram overlap pass. Note also that download throughput
varied by 2.5x between identical runs (36.9 β†’ 72.8 β†’ 63.2 MB/s), so it should be measured inside the
actual ingest job rather than treated as a constant.


## 6. Accelerator capability envelope (Turing, cc 7.5)

Established inside a GPU session, not assumed:

- **`torch.compile` works** (`warm + 20 steps` of an MLP in 6.0 s, incl. first-call compile).
- **Flash-attention is not available.** `SDPBackend.FLASH_ATTENTION` raises
  `No available kernel. Aborting execution.`, preceded by
  `Flash attention only supports gpu architectures in the range [sm80, sm121]. Attempting to run on
  a sm 7.5 gpu.` β€” so this is architectural, not a missing dependency, and installing `flash-attn`
  cannot fix it. **Β§3's "never touch Hopper-only kernels" now has a concrete instance.**
- **`EFFICIENT_ATTENTION` and `MATH` SDPA backends both run.** Memory-efficient attention is the

  viable fast path; note torch warns "Memory efficient attention has been runtime disabled" *when

  flash is force-requested*, so backend selection must be explicit.

- `torch.cuda.get_arch_list()` includes `sm_75` (also 70, 80, 86, 90, 100, 120) β€” this torch build

  does have kernels for the card.

- CUDA 12.8, cuDNN 9.10.2, **NCCL 2.27.5** present; `triton 3.6.0` present.
- Absent from the image: `flash-attn`, `xformers`, `trl`, `deepspeed`, `bitsandbytes`, `liger-kernel`,
  `evaluate`. Present: `transformers 5.0.0`, `datasets 5.0.0`, `accelerate 1.13.0`,
  `tokenizers 0.22.2`, `safetensors 0.7.0`, `numpy 2.0.2`, `pyarrow 24.0.0`, `peft 0.19.1`.
- bf16 is a hardware question, not a torch one: `torch.bfloat16` *tensors* exist, and bf16 on cc 7.5
  is at best emulated. Probe C records whether an autocast-bf16 forward even completes and what it
  costs, so the Phase 1 precision decision is evidence-backed rather than a recitation of "no bf16".
- `nvidia-smi` has no `multiprocessor_count` query field (rc=2); use `torch.cuda.…
get_device_properties(i).multi_processor_count` (40/SM) instead.

## 7. fp16 training trap found by accident

Probe B ran a small Llama with **weights cast to fp16** (`model.cuda().to(torch.float16)`) under
autocast with a plain AdamW and **no GradScaler**: every loss came back `NaN` in 65 steps. That is the
classic fp16 underflow/overflow signature, and it is exactly the kind of bug that would otherwise be
discovered 10 hours into the main run. Probe C re-runs the throughput measurements with the correct
recipe β€” **fp32 master weights + `torch.autocast(fp16)` + `torch.amp.GradScaler("cuda")` + grad

clipping** β€” and asserts finiteness, because on T4 there is no bf16 fallback: fp16 with loss scaling
is the only mixed-precision option, so getting it right is not optional.

## 8. What Phase 0 has *not* yet proven

Be honest about the boundary of this gate:

- **Hub writes from inside a Kaggle job** β€” never attempted. No credential path chosen (D-002 open).
- **NCCL across these two T4s** β€” still unproven; probe B's attempt was defeated by its own parser,
  and one failed `all_reduce` is not evidence either way. Probe C is fixing that now.
- **Checkpoint upload β†’ independent verification β†’ local delete β†’ free-space-confirmed (Β§3.13)** β€”
  untouched; that is a Phase 3 gate, and it needs the real 100M-model sizing arithmetic first.
- **Cold resume on a fresh instance with an empty disk** β€” Phase 3.
- **Real session caps**: `sessionTimeoutSeconds` was accepted and no session ran long enough to hit
  any limit. The actual wall-clock ceiling, idle-timeout behaviour, and preemption policy are
  **unknown**, and the multi-week resume plan in Phase 4 depends on them. Probe E
  (`ounce100m-p0e-session-cap`) is a ~12.5 h CPU heartbeat running now specifically to find this;
  if Kaggle kills it, the last heartbeat is the answer. **Update while writing: the session was still

  `RUNNING` at 108 minutes** (submitted 15:38Z, checked 17:26Z), so no CPU cap exists below ~1.8 h. That
  already rules out the worst case for Phase 4 β€” a session that dies inside an hour β€” and the build is
  sized so one source (~5–40 min) is the unit of lost work either way.

## 9. GPU probe C results β€” the numbers the schedule is built on

Kernel `dodosoomro/ounce100m-p0c-fp16-ddp` v2, GPU image
`gcr.io/kaggle-gpu-images/python@sha256:37c64f7d…7d461`, 2x Tesla T4, charged 157.704 s of quota.

Probe B's two void measurements are now replaced. Model shape used β€” deliberately near the likely
main-run shape so the arithmetic is about the real run and not a toy:
`hidden 768, layers 12, heads 12, GQA kv_heads 3, FFN 2048, vocab 50257, tied embeddings, seq 1024`
β†’ **112,934,400 parameters**, of which 38,597,376 are the (tied) embedding.

| measurement | result |
|---|---|
| fp32 master weights + `autocast(fp16)` + `GradScaler` + clip | **loss 9.988, finite** β€” the recipe works; probe B's NaN was the recipe, not the hardware |
| bf16 `autocast` on cc 7.5 | **runs** (loss 10.966, first fwd 0.36 s) β€” i.e. *emulated*, not tensor-cored |
| tokens/s, 1 GPU, sdpa, bs4Γ—1024 | 6,478 |
| tokens/s, 1 GPU, **eager**, bs4Γ—1024 | **7,239** |
| tokens/s per rank, DDP 2 ranks, sdpa, bs4Γ—1024 | 5,531 β†’ **11,062 aggregate** |
| DDP scaling efficiency | 11,062 / 6,478 = **1.71x** (85 % of ideal) |
| NCCL across these two T4s | **works** β€” `rc 0`, `backend nccl`, `world 2`, both ranks finite, 30 steps |
| peak CUDA memory, bs4Γ—1024, 1 GPU | **11.51 GB** (sdpa) / 13.79 GB (eager) of 14.56 GB usable |
| peak CUDA memory, bs4Γ—1024, DDP | 8.68 GB per rank |
| micro-batch ceiling at this shape | **bs 8 OOMs** (`Tried to allocate 1.54 GiB… 520 MiB free`) |

### What each line forces

- **`eager` beat `sdpa` here (7,239 vs 6,478 tok/s).** At `seq_len 1024` the memory-efficient kernel's
  overhead is not repaid, and flash-attention is unavailable on sm 7.5 anyway. Do not assume "flash is
  the fast one" β€” Phase 3 must re-measure at the frozen sequence length, because the crossover is
  length-dependent.
- **Memory, not compute, is the binding constraint on batch size.** bs4 fits in 11.5 GB of 14.56 GB;
  bs8 does not fit at all. So global batch must be built with **gradient accumulation**, not larger
  micro-batches β€” which also means the usual "increase batch until full" advice is unavailable here.
  The fp32 AdamW state for a 100M model (β‰ˆ1.2 GB) plus fp32 master weights is a fixed floor that
  does not shrink with batch size.
- **85 % DDP efficiency is good enough to plan on**, and is the number the ETA below uses. It was
  measured on 30 steps with NCCL over PCIe between two T4s; treat it as optimistic-typical, not a
  guaranteed constant.
- **bf16 running is a trap, not a permission.** Β§2 and Β§8 both rule it out. On Turing it is emulated,
  so a config that silently selects bf16 will "work" and be slow. Precision must be asserted to be
  fp16 in the frozen config, not left to a library default.

### First ETA, from measurement rather than hope

`1e9 tokens Γ· 11,062 tok/s = 90,400 s β‰ˆ 25.1 wall-clock hours` at the measured aggregate rate.

Against a 30 h weekly allowance accruing at 1x, **the main run fits inside one week of quota on this

shape β€” but with only ~5 h of headroom**, which is not enough to survive both checkpoint upload time
and the multi-session restarts that Β§4/Phase 4 assume. Three honest caveats on that number:

1. It is a 30-step measurement on one shape that is **over the parameter budget** (112.9 M > 110 M),
   so the frozen config will differ and throughput with it.
2. It excludes checkpoint write/upload/verify cycles, `DataLoader` stalls, and any
   tokenisation-time-vs-disk tradeoff in the input pipeline.
3. It assumes sessions can run long enough to be efficient; the cap is unknown until probe E reports.

Planning stance for `docs/01-plan.md`: treat **~9–11k tok/s aggregate** as the working band, size the
token target from the *low* end (~28–31 h for 1 B tokens β†’ exceeds a single week, so plan for a
two-week run), and revisit once Gate 3 measures the frozen configuration. Do not let the 25 h figure
become the plan's basis, because it leaves no room for the interruption budget that Β§3.1 explicitly
expects to be spent.

## 10. Credentials and the Kaggle-side secret store

`dodosoomro/ounce100m-p1-credential-path` (CPU, read-only) established:

- Internet-enabled Kaggle sessions **arrive pre-authenticated against the Kaggle API**. The `kaggle`
  CLI **2.0.2 is preinstalled**, and `kaggle config view` reports `username: dodosoomro`,
  `auth_method: ACCESS_TOKEN`, config from `/root/.config/kaggle` β€” with no `kaggle.json` written by
  us. `kaggle kernels list --mine` returns real rows, so the principal is genuinely usable, not merely
  present. Injected env vars: `KAGGLE_API_V1_TOKEN` (32 ch), `KAGGLE_DATA_PROXY_TOKEN` (473),
  `KAGGLE_USER_SECRETS_TOKEN` (245).
- **No Hugging Face credential exists in the container**: `HF_TOKEN` and `HUGGING_FACE_HUB_TOKEN` are
  both absent, `HF_HOME` unset, `~/.cache/huggingface` does not exist. An anonymous Hub *write* is
  correctly rejected β€” `POST /api/models/…/commit/main` β†’ **401 Unauthorized** β€” while anonymous reads
  work. So jobs can pull data freely and cannot push anything until given a token.
- `/kaggle/input` is empty and `/kaggle/input/.secrets/` **does not exist**, so Kaggle's UI-configured
  "user input / secrets" facility is unavailable: enabling it is a web-UI action, and there are no
  locally usable Kaggle credentials to do it with (`~/.kaggle` absent on this machine; Kaggle auth lives
  in the MCP server and inside the container). That rules out option 1 of D-002 for a concrete reason.
- **Consequence: the in-job authenticated Kaggle API is the usable secret store.** A job can create a
  **private** Kaggle dataset holding the HF token; later jobs mount it via `datasetDataSources` and read
  it from `/kaggle/input/…`, so the token appears in exactly one private kernel revision instead of in
  every job's source. Β§2 explicitly sanctions "a secret store", so this is the intended mechanism.

**Ordering used, because the failure mode is unrecoverable.** `kaggle datasets create` derives
visibility from a metadata field whose name has moved between CLI versions, and an unrecognised key is
**silently ignored** β€” which would create a *public* dataset containing the token. So the bootstrap
worker (a) creates the dataset with an inert placeholder only, (b) proves privacy by confirming an
**anonymous** download fails, and (c) uploads the credential as a later version only if (b) passes,
then re-proves. If anonymous read ever succeeds, the secret is never written and the job says so loudly.

**Residual risk, stated:** kernel `p1-secret-bootstrap` v1 necessarily contains the token inline, and
Kaggle retains private kernel revisions β€” overwrite-after-use does not erase history, and there is no
delete-kernel operation in CLI 2.0.2 or in the MCP toolset. This is a strictly smaller exposure than
putting the token in every training job and in anything that mirrors kernel source.

## 11. Pre-existing Kaggle kernels, not ours

`kaggle kernels list --mine` also returned five kernels predating this project
(`quick-python-cpu-smoke-test`, `monte-carlo-cpu-smoke-test`, `cpu-test`, `exam-80-monte-carlo`,
`workbuddy-cpu-probe`; last run 2026-09-17 β†’ 2026-09-19). They are **not touched** (Β§3.10). This also
explains E-002: `search_notebooks` returning `{}` for this account was the search tool failing, not the
account being empty β€” so an empty search must never be read as "no prior work".



## 12. Measured notebook session cap: 12.0 h survived, the instance reported 45,000 s

`dodosoomro/ounce100m-p0e-session-cap` β€” a CPU kernel that printed one heartbeat every two minutes and
nothing else, costing zero GPU quota, left running since 2026-09-19T15:38Z purely as a measurement.

```
CAP_START {"epoch": 1789832305.41, "host": "096ef365b02e", "pid": 7, "limit": 45000}
HB 1   elapsed=0      free_GB=19.5 mem_avail_MB=31204 load=0.41 net=204
HB 361 elapsed=43204  free_GB=19.5 mem_avail_MB=31288 load=0.00 net=204
```

**361 heartbeats, 43,204 s = 12.0 h of continuous running, then the platform took it**
(`CANCEL_ACKNOWLEDGED`, not a crash β€” the last heartbeat is clean and disk/RAM are unchanged at 19.5 GB
free and 31.3 GB available). The instance itself reported `limit: 45000` s, i.e. **12.5 h**, so the kill
arrived at ~96 % of its own stated budget.

What this does and does not license:

- It is a **CPU** measurement. The GPU accelerator quota is tracked separately, and nothing here proves a
  GPU session is allowed to run 12.5 h. So Phase 4 plans `SESSION_GPU_HOURS=6.9` for session 1 β€” well
  inside every plausible cap β€” and can extend later sessions only on evidence from earlier ones.
- It does settle the question that the plan had been hedging since Β§2 of `04-run-log.md` ("the session cap
  is unmeasured"): an interruption at 6 h is not the platform's limit, so a long session is not doomed, and
  a `--stop-after-steps` schedule sized for 6-7 h is conservative rather than necessary.
- It confirms the working volume is stable for the whole duration β€” `free_GB` never moved from 19.5, so an
  11 GB drift over a session is not a thing to plan for, and the disk floor in the launcher is about the
  mix and checkpoints, not about slow leaks.