# 00 — Platform notes (Phase 0) Observed mechanics of the two execution environments this project spans: this workspace and the Kaggle image. Everything here was **measured on 2026-09-19**, not recalled. Where a claim contradicts prior documentation (including the Kaggle skill's own notes), the contradiction is called out, because a future session will otherwise trust the older statement. Raw probe JSON lives in the run logs of the kernels listed in `memory/ASSETS.md`. --- ## 1. The two environments are not the same machine | | this workspace | Kaggle CPU session | Kaggle GPU session (2xT4) | |---|---|---|---| | OS | Windows 11, git-bash | Linux 6.12.90+, glibc 2.35 | Linux 6.12.90+, glibc 2.35 | | Python | 3.11 (and 3.14) | **3.12.13** | **3.12.13** | | Docker image | n/a | `gcr.io/kaggle-images/python@sha256:dafd4ce5…c40b9` | `gcr.io/kaggle-gpu-images/python@sha256:37c64f7d…7d461` | | torch | — | **2.10.0+cpu** | **2.10.0+cu128** | | internet | — | yes (needs `enableInternet: true`) | yes (needs `enableInternet: true`) | | GPU | none | `device_count()==0` | **2x Tesla T4, cc 7.5, 40 SM, 14.56 GiB usable each** | **Consequence:** CPU and GPU sessions are *different images*, not one image with a GPU attached. The CPU image's `torch` is a `+cpu` build, so a script that imports `torch.cuda` will work on the GPU shape and silently do nothing useful on the CPU shape. Pin these digests in any job that must be reproducible; a "latest" Kaggle image can move under the run. `huggingface_hub` is **1.32.0 locally** and reported as `huggingface-hub 1.11.0` inside the Kaggle image. Local code and in-job code are therefore on different minor versions — an API that exists locally may not exist in the job (`create_repo(tags=…)` is one such example; it raised `TypeError` locally). Check the target environment's version before using a new API. ## 2. Job mechanics that the main run will depend on - **Submission is one call.** `save_notebook` with `kernelExecutionType: "SaveAndRunAll"` creates *and* runs. `kernelType: "script"` and `machineShape` must both be present or the call fails with a message-less error. `machineShape: "GPU"` + `enableGpu: true` is accepted and yields 2xT4 — no T4-specific shape string is needed, and none should be invented. - **`newTitle` rewrites the slug.** Kept `newTitle` equal to the slug suffix in every probe so the returned `ref` matched what was requested. Still: **poll the returned slug, never the requested one.** - **Latency.** A ~30 s CPU script went submit → `COMPLETE` in ~85 s; a ~58 s GPU script in ~145 s. So budget **~60–90 s of queue+boot per job**, and don't poll sooner than ~60 s. Both figures are single samples from an uncongested account at ~15:20 UTC. - **Where output lands.** `list_notebook_session_output` returns `files[]` (signed URLs, expiry unknown — download promptly) and `log`, a JSON array of `{stream_name, time, data}` where `time` is **seconds since session start**, one array element per line. The log carries the whole stdout, so a probe that prints its result as JSON needs no file download. **Only files written under `/kaggle/working` are returned** — the probes deliberately wrote nothing to `/tmp`, because `/tmp` looks enormous (§4) and is therefore the easiest place to silently lose an artifact. - **Kernel logs are not secret-safe.** Anything printed is retained in the kernel revision. No token may be printed, and by extension no token may be *in the kernel source*, since source and logs are both Kaggle-held. - **`CUDA_VISIBLE_DEVICES` is unset in both shapes** — it is absent from the environment entirely, not set to the string `"None"`. The Kaggle skill documents `CUDA_VISIBLE_DEVICES=None` for CPU runs; that is wrong for this image (E-002 in `memory/ERRORS.md`), and code that branches on `== "None"` will take the GPU branch on a CPU box. Gate on `torch.cuda.is_available()` or `torch.cuda.device_count()` instead. ## 3. Quota accounting - Readout is **seconds**: `total_time_allowed: "108000s"` = 30 h, `quota_refresh_time: 2026-09-26T00:00:00Z`. Weekly reset is **Saturday 00:00 UTC**. - CPU sessions cost **zero** GPU quota, confirmed: two CPU probes ran and `time_used` stayed `0s`. - A 2xT4 session charged **66.411 s** for **58.4 s** of script time → **~1x wall-clock, not 2x**. Full reasoning, the caveat, and the second sample pending in `memory/QUOTA.md`. This is the single most schedule-relevant number measured in Phase 0: at 1x, 30 h/week is 30 wall-clock hours with both cards, not the 15 that §2's "shared across both cards" phrasing suggests. ## 4. Disk and memory — the assumption in §2 is only half right | Mount | Total | Free | Notes | |---|---|---|---| | `/kaggle/working`, `/kaggle/input` | **19.5 GB** | 19.5 GB | the small Kaggle disk; **only this is retrieved as output** | | `/dev/loop1` → `/kaggle/src` | 20 GB | 20 GB | separate loop device | | `/`, `/tmp`, `/usr`, `/root/.cache` (overlay) | 7.9 TB | **1.1 TB** | **shared host filesystem, 88 % used by other tenants** | | `/dev/shm` | 14 GB | 14 GB | useful for DataLoader worker traffic | | RAM | 32 GB (`cgroup memory.max` = 30 GiB) | 31.2 GB available | `SwapTotal: 0` | | CPU | 4 vCPU (`cpu.max` = `400000 100000`), Intel Xeon @ 2.20 GHz | | | **Read this carefully before planning the checkpoint cycle.** §2 says "Kaggle disk is small", and for `/kaggle/working` that is exactly true: 19.5 GB. But the root overlay reports ~1 TB free, and a naive optimisation would write caches or staged shards there. Two reasons not to: 1. The overlay is a **shared host volume already at 88 %**, so "1 TB free" is other tenants' headroom and can disappear mid-run. §3.13's "never let it fill the disk" must be enforced against a number we do not control. 2. Nothing outside `/kaggle/working` is retrievable when the session ends, so it is not storage in any sense that matters. Default `HF_HOME`/datasets cache resolves under `/root/.cache`, i.e. the overlay, not the 19.5 GB volume. So **a training job's apparent free space can be large while `/kaggle/working` fills up** — the failure mode §3.13 warns about is invisible to a naive `disk_usage("/")` check. **Every job must measure the specific filesystem it writes to, and log free space for that path.** On the CPU shape the effective budget is ~19.5 GB; treat that as the design number for shards resident at once. No swap: a tokenizer build or a large in-memory dedup on a 30 GiB cgroup must be sized, not streamed optimistically. ## 5. Network behaviour, including one thing that does not resolve Working, measured from inside a GPU and a CPU session: `huggingface.co/api/models` → 200 in 149 ms; `pypi.org/simple/` → 200 in 73 ms; `datasets/…/resolve/main/README.md` → 200 via `api/resolve-cache/…`; `datasets.load_dataset("wikimedia/wikipedia", "20231101.en", streaming=True)` → first row in **7.4 s**, unauthenticated. `pip install` from PyPI works: `xxhash` in **5.8 s**, import verified. **Does not resolve:** `cdn-lfs.huggingface.co`, `cdn-lfs-us-1.huggingface.co`, `xethub.huggingface.co` → `gaierror`. That looked alarming, so probe D tested it directly by pulling 300 MB from two large public parquet files. **Sustained Hub download throughput from inside a session: 72.8 MB/s** (`wikimedia/wikipedia`, 745 MB file) **and 88.9 MB/s** (`HuggingFaceFW/fineweb-edu`, 2334 MB file), with 0.8–0.94 s time-to-first-byte. So the non-resolving LFS hostnames do not block bulk transfer — `resolve/` redirects land on hosts that do resolve. **Ingest of the mix is not network-bound:** at a conservative 70 MB/s, one billion tokens (≈3.6 GB of raw text at ~3.6 chars/token, or ~4–6 GB as parquet) downloads in single-digit minutes. Phase 2 should plan shard staging around CPU and disk limits, not bandwidth. Also measured in the same probe: `datasets` streaming at **139.5 rows/s ≈ 3.0 MB of text per second** single-process, unauthenticated (83 rows/s / 1.8 MB/s cold, ~3 MB/s warm). Unauthenticated Hub access works but the API warns about rate limits (`Please set a HF_TOKEN`); at multi-hundred-shard scale that is a plausible 429 source, which is an argument for the token-in-job path (D-002) for the *data plane* even if the *checkpoint plane* stays local. Enumerating repo files without a client library: `GET https://huggingface.co/api/datasets//tree/main?recursive=true` returns `path`/`size`/`type` and works with plain `urllib`. It appeared to cap at 1000 entries (`fineweb-edu` returned exactly 1000), so **assume pagination is required** for a large repo. ### CPU-side ingest throughput, and what it rules out Measured on the Kaggle CPU shape (4 vCPU Xeon @ 2.2 GHz) with the gpt2 tokenizer over 6 MB of real Wikipedia English, `tokenizers 0.22.2` `Tokenizer.encode_batch`: | cores | tokens/s | hours to tokenize 1B tokens | |---|---|---| | 1 | 245,660 | 1.13 | | 2 | 492,191 | 0.56 | | 4 | **660,943** | **0.42** | `chars_per_token` 4.53 on this corpus. Scaling is near-linear to 2 cores and then sub-linear (2.7x at 4 cores), which is what a GIL-bound parent feeding rayon workers looks like — so a Phase 2 pipeline that shards by **process**, not by thread, should recover the rest. `datasets` streaming ran at 151.6 rows/s ≈ **3.26 MB of text per second** single-process, unauthenticated. **The conclusion is a negative result, which is the useful kind:** tokenizing 1B tokens costs ~0.4 h of the 4-core budget, and at 35–89 MB/s the raw bytes arrive faster than `datasets` can stream them. **Phase 2 is therefore not CPU-bound and not bandwidth-bound — it is bound by dedup memory and by the 30 GiB cgroup with no swap.** Budget the design accordingly: the dedup/audit stage must be written to stream and shard against fixed RAM, not to hold a billion sketches, and the only real wall-clock cost in building the mix is the exact-duplicate/n-gram overlap pass. Note also that download throughput varied by 2.5x between identical runs (36.9 → 72.8 → 63.2 MB/s), so it should be measured inside the actual ingest job rather than treated as a constant. ## 6. Accelerator capability envelope (Turing, cc 7.5) Established inside a GPU session, not assumed: - **`torch.compile` works** (`warm + 20 steps` of an MLP in 6.0 s, incl. first-call compile). - **Flash-attention is not available.** `SDPBackend.FLASH_ATTENTION` raises `No available kernel. Aborting execution.`, preceded by `Flash attention only supports gpu architectures in the range [sm80, sm121]. Attempting to run on a sm 7.5 gpu.` — so this is architectural, not a missing dependency, and installing `flash-attn` cannot fix it. **§3's "never touch Hopper-only kernels" now has a concrete instance.** - **`EFFICIENT_ATTENTION` and `MATH` SDPA backends both run.** Memory-efficient attention is the viable fast path; note torch warns "Memory efficient attention has been runtime disabled" *when flash is force-requested*, so backend selection must be explicit. - `torch.cuda.get_arch_list()` includes `sm_75` (also 70, 80, 86, 90, 100, 120) — this torch build does have kernels for the card. - CUDA 12.8, cuDNN 9.10.2, **NCCL 2.27.5** present; `triton 3.6.0` present. - Absent from the image: `flash-attn`, `xformers`, `trl`, `deepspeed`, `bitsandbytes`, `liger-kernel`, `evaluate`. Present: `transformers 5.0.0`, `datasets 5.0.0`, `accelerate 1.13.0`, `tokenizers 0.22.2`, `safetensors 0.7.0`, `numpy 2.0.2`, `pyarrow 24.0.0`, `peft 0.19.1`. - bf16 is a hardware question, not a torch one: `torch.bfloat16` *tensors* exist, and bf16 on cc 7.5 is at best emulated. Probe C records whether an autocast-bf16 forward even completes and what it costs, so the Phase 1 precision decision is evidence-backed rather than a recitation of "no bf16". - `nvidia-smi` has no `multiprocessor_count` query field (rc=2); use `torch.cuda.… get_device_properties(i).multi_processor_count` (40/SM) instead. ## 7. fp16 training trap found by accident Probe B ran a small Llama with **weights cast to fp16** (`model.cuda().to(torch.float16)`) under autocast with a plain AdamW and **no GradScaler**: every loss came back `NaN` in 65 steps. That is the classic fp16 underflow/overflow signature, and it is exactly the kind of bug that would otherwise be discovered 10 hours into the main run. Probe C re-runs the throughput measurements with the correct recipe — **fp32 master weights + `torch.autocast(fp16)` + `torch.amp.GradScaler("cuda")` + grad clipping** — and asserts finiteness, because on T4 there is no bf16 fallback: fp16 with loss scaling is the only mixed-precision option, so getting it right is not optional. ## 8. What Phase 0 has *not* yet proven Be honest about the boundary of this gate: - **Hub writes from inside a Kaggle job** — never attempted. No credential path chosen (D-002 open). - **NCCL across these two T4s** — still unproven; probe B's attempt was defeated by its own parser, and one failed `all_reduce` is not evidence either way. Probe C is fixing that now. - **Checkpoint upload → independent verification → local delete → free-space-confirmed (§3.13)** — untouched; that is a Phase 3 gate, and it needs the real 100M-model sizing arithmetic first. - **Cold resume on a fresh instance with an empty disk** — Phase 3. - **Real session caps**: `sessionTimeoutSeconds` was accepted and no session ran long enough to hit any limit. The actual wall-clock ceiling, idle-timeout behaviour, and preemption policy are **unknown**, and the multi-week resume plan in Phase 4 depends on them. Probe E (`ounce100m-p0e-session-cap`) is a ~12.5 h CPU heartbeat running now specifically to find this; if Kaggle kills it, the last heartbeat is the answer. **Update while writing: the session was still `RUNNING` at 108 minutes** (submitted 15:38Z, checked 17:26Z), so no CPU cap exists below ~1.8 h. That already rules out the worst case for Phase 4 — a session that dies inside an hour — and the build is sized so one source (~5–40 min) is the unit of lost work either way. ## 9. GPU probe C results — the numbers the schedule is built on Kernel `dodosoomro/ounce100m-p0c-fp16-ddp` v2, GPU image `gcr.io/kaggle-gpu-images/python@sha256:37c64f7d…7d461`, 2x Tesla T4, charged 157.704 s of quota. Probe B's two void measurements are now replaced. Model shape used — deliberately near the likely main-run shape so the arithmetic is about the real run and not a toy: `hidden 768, layers 12, heads 12, GQA kv_heads 3, FFN 2048, vocab 50257, tied embeddings, seq 1024` → **112,934,400 parameters**, of which 38,597,376 are the (tied) embedding. | measurement | result | |---|---| | fp32 master weights + `autocast(fp16)` + `GradScaler` + clip | **loss 9.988, finite** — the recipe works; probe B's NaN was the recipe, not the hardware | | bf16 `autocast` on cc 7.5 | **runs** (loss 10.966, first fwd 0.36 s) — i.e. *emulated*, not tensor-cored | | tokens/s, 1 GPU, sdpa, bs4×1024 | 6,478 | | tokens/s, 1 GPU, **eager**, bs4×1024 | **7,239** | | tokens/s per rank, DDP 2 ranks, sdpa, bs4×1024 | 5,531 → **11,062 aggregate** | | DDP scaling efficiency | 11,062 / 6,478 = **1.71x** (85 % of ideal) | | NCCL across these two T4s | **works** — `rc 0`, `backend nccl`, `world 2`, both ranks finite, 30 steps | | peak CUDA memory, bs4×1024, 1 GPU | **11.51 GB** (sdpa) / 13.79 GB (eager) of 14.56 GB usable | | peak CUDA memory, bs4×1024, DDP | 8.68 GB per rank | | micro-batch ceiling at this shape | **bs 8 OOMs** (`Tried to allocate 1.54 GiB… 520 MiB free`) | ### What each line forces - **`eager` beat `sdpa` here (7,239 vs 6,478 tok/s).** At `seq_len 1024` the memory-efficient kernel's overhead is not repaid, and flash-attention is unavailable on sm 7.5 anyway. Do not assume "flash is the fast one" — Phase 3 must re-measure at the frozen sequence length, because the crossover is length-dependent. - **Memory, not compute, is the binding constraint on batch size.** bs4 fits in 11.5 GB of 14.56 GB; bs8 does not fit at all. So global batch must be built with **gradient accumulation**, not larger micro-batches — which also means the usual "increase batch until full" advice is unavailable here. The fp32 AdamW state for a 100M model (≈1.2 GB) plus fp32 master weights is a fixed floor that does not shrink with batch size. - **85 % DDP efficiency is good enough to plan on**, and is the number the ETA below uses. It was measured on 30 steps with NCCL over PCIe between two T4s; treat it as optimistic-typical, not a guaranteed constant. - **bf16 running is a trap, not a permission.** §2 and §8 both rule it out. On Turing it is emulated, so a config that silently selects bf16 will "work" and be slow. Precision must be asserted to be fp16 in the frozen config, not left to a library default. ### First ETA, from measurement rather than hope `1e9 tokens ÷ 11,062 tok/s = 90,400 s ≈ 25.1 wall-clock hours` at the measured aggregate rate. Against a 30 h weekly allowance accruing at 1x, **the main run fits inside one week of quota on this shape — but with only ~5 h of headroom**, which is not enough to survive both checkpoint upload time and the multi-session restarts that §4/Phase 4 assume. Three honest caveats on that number: 1. It is a 30-step measurement on one shape that is **over the parameter budget** (112.9 M > 110 M), so the frozen config will differ and throughput with it. 2. It excludes checkpoint write/upload/verify cycles, `DataLoader` stalls, and any tokenisation-time-vs-disk tradeoff in the input pipeline. 3. It assumes sessions can run long enough to be efficient; the cap is unknown until probe E reports. Planning stance for `docs/01-plan.md`: treat **~9–11k tok/s aggregate** as the working band, size the token target from the *low* end (~28–31 h for 1 B tokens → exceeds a single week, so plan for a two-week run), and revisit once Gate 3 measures the frozen configuration. Do not let the 25 h figure become the plan's basis, because it leaves no room for the interruption budget that §3.1 explicitly expects to be spent. ## 10. Credentials and the Kaggle-side secret store `dodosoomro/ounce100m-p1-credential-path` (CPU, read-only) established: - Internet-enabled Kaggle sessions **arrive pre-authenticated against the Kaggle API**. The `kaggle` CLI **2.0.2 is preinstalled**, and `kaggle config view` reports `username: dodosoomro`, `auth_method: ACCESS_TOKEN`, config from `/root/.config/kaggle` — with no `kaggle.json` written by us. `kaggle kernels list --mine` returns real rows, so the principal is genuinely usable, not merely present. Injected env vars: `KAGGLE_API_V1_TOKEN` (32 ch), `KAGGLE_DATA_PROXY_TOKEN` (473), `KAGGLE_USER_SECRETS_TOKEN` (245). - **No Hugging Face credential exists in the container**: `HF_TOKEN` and `HUGGING_FACE_HUB_TOKEN` are both absent, `HF_HOME` unset, `~/.cache/huggingface` does not exist. An anonymous Hub *write* is correctly rejected — `POST /api/models/…/commit/main` → **401 Unauthorized** — while anonymous reads work. So jobs can pull data freely and cannot push anything until given a token. - `/kaggle/input` is empty and `/kaggle/input/.secrets/` **does not exist**, so Kaggle's UI-configured "user input / secrets" facility is unavailable: enabling it is a web-UI action, and there are no locally usable Kaggle credentials to do it with (`~/.kaggle` absent on this machine; Kaggle auth lives in the MCP server and inside the container). That rules out option 1 of D-002 for a concrete reason. - **Consequence: the in-job authenticated Kaggle API is the usable secret store.** A job can create a **private** Kaggle dataset holding the HF token; later jobs mount it via `datasetDataSources` and read it from `/kaggle/input/…`, so the token appears in exactly one private kernel revision instead of in every job's source. §2 explicitly sanctions "a secret store", so this is the intended mechanism. **Ordering used, because the failure mode is unrecoverable.** `kaggle datasets create` derives visibility from a metadata field whose name has moved between CLI versions, and an unrecognised key is **silently ignored** — which would create a *public* dataset containing the token. So the bootstrap worker (a) creates the dataset with an inert placeholder only, (b) proves privacy by confirming an **anonymous** download fails, and (c) uploads the credential as a later version only if (b) passes, then re-proves. If anonymous read ever succeeds, the secret is never written and the job says so loudly. **Residual risk, stated:** kernel `p1-secret-bootstrap` v1 necessarily contains the token inline, and Kaggle retains private kernel revisions — overwrite-after-use does not erase history, and there is no delete-kernel operation in CLI 2.0.2 or in the MCP toolset. This is a strictly smaller exposure than putting the token in every training job and in anything that mirrors kernel source. ## 11. Pre-existing Kaggle kernels, not ours `kaggle kernels list --mine` also returned five kernels predating this project (`quick-python-cpu-smoke-test`, `monte-carlo-cpu-smoke-test`, `cpu-test`, `exam-80-monte-carlo`, `workbuddy-cpu-probe`; last run 2026-09-17 → 2026-09-19). They are **not touched** (§3.10). This also explains E-002: `search_notebooks` returning `{}` for this account was the search tool failing, not the account being empty — so an empty search must never be read as "no prior work". ## 12. Measured notebook session cap: 12.0 h survived, the instance reported 45,000 s `dodosoomro/ounce100m-p0e-session-cap` — a CPU kernel that printed one heartbeat every two minutes and nothing else, costing zero GPU quota, left running since 2026-09-19T15:38Z purely as a measurement. ``` CAP_START {"epoch": 1789832305.41, "host": "096ef365b02e", "pid": 7, "limit": 45000} HB 1 elapsed=0 free_GB=19.5 mem_avail_MB=31204 load=0.41 net=204 HB 361 elapsed=43204 free_GB=19.5 mem_avail_MB=31288 load=0.00 net=204 ``` **361 heartbeats, 43,204 s = 12.0 h of continuous running, then the platform took it** (`CANCEL_ACKNOWLEDGED`, not a crash — the last heartbeat is clean and disk/RAM are unchanged at 19.5 GB free and 31.3 GB available). The instance itself reported `limit: 45000` s, i.e. **12.5 h**, so the kill arrived at ~96 % of its own stated budget. What this does and does not license: - It is a **CPU** measurement. The GPU accelerator quota is tracked separately, and nothing here proves a GPU session is allowed to run 12.5 h. So Phase 4 plans `SESSION_GPU_HOURS=6.9` for session 1 — well inside every plausible cap — and can extend later sessions only on evidence from earlier ones. - It does settle the question that the plan had been hedging since §2 of `04-run-log.md` ("the session cap is unmeasured"): an interruption at 6 h is not the platform's limit, so a long session is not doomed, and a `--stop-after-steps` schedule sized for 6-7 h is conservative rather than necessary. - It confirms the working volume is stable for the whole duration — `free_GB` never moved from 19.5, so an 11 GB drift over a session is not a thing to plan for, and the disk floor in the launcher is about the mix and checkpoints, not about slow leaks.