ounce100m-code / docs /00-platform-notes.md
Cion-lab's picture
publish docs/: plan, mix rationale, preflight report, run log, frozen eval protocol, final report
f345921 verified
|
Raw History Blame Contribute Delete
23.7 kB
# 00 β€” Platform notes (Phase 0)
Observed mechanics of the two execution environments this project spans: this workspace and the
Kaggle image. Everything here was **measured on 2026-09-19**, not recalled. Where a claim contradicts
prior documentation (including the Kaggle skill's own notes), the contradiction is called out, because
a future session will otherwise trust the older statement.
Raw probe JSON lives in the run logs of the kernels listed in `memory/ASSETS.md`.
---
## 1. The two environments are not the same machine
| | this workspace | Kaggle CPU session | Kaggle GPU session (2xT4) |
|---|---|---|---|
| OS | Windows 11, git-bash | Linux 6.12.90+, glibc 2.35 | Linux 6.12.90+, glibc 2.35 |
| Python | 3.11 (and 3.14) | **3.12.13** | **3.12.13** |
| Docker image | n/a | `gcr.io/kaggle-images/python@sha256:dafd4ce5…c40b9` | `gcr.io/kaggle-gpu-images/python@sha256:37c64f7d…7d461` |
| torch | β€” | **2.10.0+cpu** | **2.10.0+cu128** |
| internet | β€” | yes (needs `enableInternet: true`) | yes (needs `enableInternet: true`) |
| GPU | none | `device_count()==0` | **2x Tesla T4, cc 7.5, 40 SM, 14.56 GiB usable each** |
**Consequence:** CPU and GPU sessions are *different images*, not one image with a GPU attached. The
CPU image's `torch` is a `+cpu` build, so a script that imports `torch.cuda` will work on the GPU
shape and silently do nothing useful on the CPU shape. Pin these digests in any job that must be
reproducible; a "latest" Kaggle image can move under the run.
`huggingface_hub` is **1.32.0 locally** and reported as `huggingface-hub 1.11.0` inside the Kaggle
image. Local code and in-job code are therefore on different minor versions β€” an API that exists
locally may not exist in the job (`create_repo(tags=…)` is one such example; it raised
`TypeError` locally). Check the target environment's version before using a new API.
## 2. Job mechanics that the main run will depend on
- **Submission is one call.** `save_notebook` with `kernelExecutionType: "SaveAndRunAll"` creates
*and* runs. `kernelType: "script"` and `machineShape` must both be present or the call fails with a
message-less error. `machineShape: "GPU"` + `enableGpu: true` is accepted and yields 2xT4 β€” no
T4-specific shape string is needed, and none should be invented.
- **`newTitle` rewrites the slug.** Kept `newTitle` equal to the slug suffix in every probe so the
returned `ref` matched what was requested. Still: **poll the returned slug, never the requested one.**
- **Latency.** A ~30 s CPU script went submit β†’ `COMPLETE` in ~85 s; a ~58 s GPU script in ~145 s.
So budget **~60–90 s of queue+boot per job**, and don't poll sooner than ~60 s. Both figures are
single samples from an uncongested account at ~15:20 UTC.
- **Where output lands.** `list_notebook_session_output` returns `files[]` (signed URLs, expiry
unknown β€” download promptly) and `log`, a JSON array of `{stream_name, time, data}` where `time` is
**seconds since session start**, one array element per line. The log carries the whole stdout, so a
probe that prints its result as JSON needs no file download. **Only files written under
`/kaggle/working` are returned** β€” the probes deliberately wrote nothing to `/tmp`, because `/tmp`
looks enormous (Β§4) and is therefore the easiest place to silently lose an artifact.
- **Kernel logs are not secret-safe.** Anything printed is retained in the kernel revision. No token
may be printed, and by extension no token may be *in the kernel source*, since source and logs are
both Kaggle-held.
- **`CUDA_VISIBLE_DEVICES` is unset in both shapes** β€” it is absent from the environment entirely, not
set to the string `"None"`. The Kaggle skill documents `CUDA_VISIBLE_DEVICES=None` for CPU runs;
that is wrong for this image (E-002 in `memory/ERRORS.md`), and code that branches on
`== "None"` will take the GPU branch on a CPU box. Gate on `torch.cuda.is_available()` or
`torch.cuda.device_count()` instead.
## 3. Quota accounting
- Readout is **seconds**: `total_time_allowed: "108000s"` = 30 h, `quota_refresh_time:
2026-09-26T00:00:00Z`. Weekly reset is **Saturday 00:00 UTC**.
- CPU sessions cost **zero** GPU quota, confirmed: two CPU probes ran and `time_used` stayed `0s`.
- A 2xT4 session charged **66.411 s** for **58.4 s** of script time β†’ **~1x wall-clock, not 2x**.
Full reasoning, the caveat, and the second sample pending in `memory/QUOTA.md`. This is the single
most schedule-relevant number measured in Phase 0: at 1x, 30 h/week is 30 wall-clock hours with
both cards, not the 15 that Β§2's "shared across both cards" phrasing suggests.
## 4. Disk and memory β€” the assumption in Β§2 is only half right
| Mount | Total | Free | Notes |
|---|---|---|---|
| `/kaggle/working`, `/kaggle/input` | **19.5 GB** | 19.5 GB | the small Kaggle disk; **only this is retrieved as output** |
| `/dev/loop1` β†’ `/kaggle/src` | 20 GB | 20 GB | separate loop device |
| `/`, `/tmp`, `/usr`, `/root/.cache` (overlay) | 7.9 TB | **1.1 TB** | **shared host filesystem, 88 % used by other tenants** |
| `/dev/shm` | 14 GB | 14 GB | useful for DataLoader worker traffic |
| RAM | 32 GB (`cgroup memory.max` = 30 GiB) | 31.2 GB available | `SwapTotal: 0` |
| CPU | 4 vCPU (`cpu.max` = `400000 100000`), Intel Xeon @ 2.20 GHz | | |
**Read this carefully before planning the checkpoint cycle.** Β§2 says "Kaggle disk is small", and for
`/kaggle/working` that is exactly true: 19.5 GB. But the root overlay reports ~1 TB free, and a naive
optimisation would write caches or staged shards there. Two reasons not to:
1. The overlay is a **shared host volume already at 88 %**, so "1 TB free" is other tenants' headroom
and can disappear mid-run. Β§3.13's "never let it fill the disk" must be enforced against a number
we do not control.
2. Nothing outside `/kaggle/working` is retrievable when the session ends, so it is not storage in
any sense that matters.
Default `HF_HOME`/datasets cache resolves under `/root/.cache`, i.e. the overlay, not the 19.5 GB
volume. So **a training job's apparent free space can be large while `/kaggle/working` fills up** β€”
the failure mode Β§3.13 warns about is invisible to a naive `disk_usage("/")` check. **Every job must
measure the specific filesystem it writes to, and log free space for that path.** On the CPU shape the
effective budget is ~19.5 GB; treat that as the design number for shards resident at once.
No swap: a tokenizer build or a large in-memory dedup on a 30 GiB cgroup must be sized, not streamed
optimistically.
## 5. Network behaviour, including one thing that does not resolve
Working, measured from inside a GPU and a CPU session:
`huggingface.co/api/models` β†’ 200 in 149 ms; `pypi.org/simple/` β†’ 200 in 73 ms;
`datasets/…/resolve/main/README.md` β†’ 200 via `api/resolve-cache/…`;
`datasets.load_dataset("wikimedia/wikipedia", "20231101.en", streaming=True)` β†’ first row in **7.4 s**,
unauthenticated. `pip install` from PyPI works: `xxhash` in **5.8 s**, import verified.
**Does not resolve:** `cdn-lfs.huggingface.co`, `cdn-lfs-us-1.huggingface.co`, `xethub.huggingface.co`
β†’ `gaierror`. That looked alarming, so probe D tested it directly by pulling 300 MB from two large
public parquet files.
**Sustained Hub download throughput from inside a session: 72.8 MB/s** (`wikimedia/wikipedia`, 745 MB
file) **and 88.9 MB/s** (`HuggingFaceFW/fineweb-edu`, 2334 MB file), with 0.8–0.94 s time-to-first-byte.
So the non-resolving LFS hostnames do not block bulk transfer β€” `resolve/` redirects land on hosts that
do resolve. **Ingest of the mix is not network-bound:** at a conservative 70 MB/s, one billion tokens
(β‰ˆ3.6 GB of raw text at ~3.6 chars/token, or ~4–6 GB as parquet) downloads in single-digit minutes.
Phase 2 should plan shard staging around CPU and disk limits, not bandwidth.
Also measured in the same probe: `datasets` streaming at **139.5 rows/s β‰ˆ 3.0 MB of text per second**
single-process, unauthenticated (83 rows/s / 1.8 MB/s cold, ~3 MB/s warm). Unauthenticated Hub access
works but the API warns about rate limits (`Please set a HF_TOKEN`); at multi-hundred-shard scale that
is a plausible 429 source, which is an argument for the token-in-job path (D-002) for the *data plane*
even if the *checkpoint plane* stays local.
Enumerating repo files without a client library: `GET https://huggingface.co/api/datasets/<id>/tree/main?recursive=true`
returns `path`/`size`/`type` and works with plain `urllib`. It appeared to cap at 1000 entries
(`fineweb-edu` returned exactly 1000), so **assume pagination is required** for a large repo.
### CPU-side ingest throughput, and what it rules out
Measured on the Kaggle CPU shape (4 vCPU Xeon @ 2.2 GHz) with the gpt2 tokenizer over 6 MB of real
Wikipedia English, `tokenizers 0.22.2` `Tokenizer.encode_batch`:
| cores | tokens/s | hours to tokenize 1B tokens |
|---|---|---|
| 1 | 245,660 | 1.13 |
| 2 | 492,191 | 0.56 |
| 4 | **660,943** | **0.42** |
`chars_per_token` 4.53 on this corpus. Scaling is near-linear to 2 cores and then sub-linear
(2.7x at 4 cores), which is what a GIL-bound parent feeding rayon workers looks like β€” so a Phase 2
pipeline that shards by **process**, not by thread, should recover the rest. `datasets` streaming ran
at 151.6 rows/s β‰ˆ **3.26 MB of text per second** single-process, unauthenticated.
**The conclusion is a negative result, which is the useful kind:** tokenizing 1B tokens costs ~0.4 h of
the 4-core budget, and at 35–89 MB/s the raw bytes arrive faster than `datasets` can stream them.
**Phase 2 is therefore not CPU-bound and not bandwidth-bound β€” it is bound by dedup memory and by the
30 GiB cgroup with no swap.** Budget the design accordingly: the dedup/audit stage must be written to
stream and shard against fixed RAM, not to hold a billion sketches, and the only real wall-clock cost
in building the mix is the exact-duplicate/n-gram overlap pass. Note also that download throughput
varied by 2.5x between identical runs (36.9 β†’ 72.8 β†’ 63.2 MB/s), so it should be measured inside the
actual ingest job rather than treated as a constant.
## 6. Accelerator capability envelope (Turing, cc 7.5)
Established inside a GPU session, not assumed:
- **`torch.compile` works** (`warm + 20 steps` of an MLP in 6.0 s, incl. first-call compile).
- **Flash-attention is not available.** `SDPBackend.FLASH_ATTENTION` raises
`No available kernel. Aborting execution.`, preceded by
`Flash attention only supports gpu architectures in the range [sm80, sm121]. Attempting to run on
a sm 7.5 gpu.` β€” so this is architectural, not a missing dependency, and installing `flash-attn`
cannot fix it. **Β§3's "never touch Hopper-only kernels" now has a concrete instance.**
- **`EFFICIENT_ATTENTION` and `MATH` SDPA backends both run.** Memory-efficient attention is the
viable fast path; note torch warns "Memory efficient attention has been runtime disabled" *when
flash is force-requested*, so backend selection must be explicit.
- `torch.cuda.get_arch_list()` includes `sm_75` (also 70, 80, 86, 90, 100, 120) β€” this torch build
does have kernels for the card.
- CUDA 12.8, cuDNN 9.10.2, **NCCL 2.27.5** present; `triton 3.6.0` present.
- Absent from the image: `flash-attn`, `xformers`, `trl`, `deepspeed`, `bitsandbytes`, `liger-kernel`,
`evaluate`. Present: `transformers 5.0.0`, `datasets 5.0.0`, `accelerate 1.13.0`,
`tokenizers 0.22.2`, `safetensors 0.7.0`, `numpy 2.0.2`, `pyarrow 24.0.0`, `peft 0.19.1`.
- bf16 is a hardware question, not a torch one: `torch.bfloat16` *tensors* exist, and bf16 on cc 7.5
is at best emulated. Probe C records whether an autocast-bf16 forward even completes and what it
costs, so the Phase 1 precision decision is evidence-backed rather than a recitation of "no bf16".
- `nvidia-smi` has no `multiprocessor_count` query field (rc=2); use `torch.cuda.…
get_device_properties(i).multi_processor_count` (40/SM) instead.
## 7. fp16 training trap found by accident
Probe B ran a small Llama with **weights cast to fp16** (`model.cuda().to(torch.float16)`) under
autocast with a plain AdamW and **no GradScaler**: every loss came back `NaN` in 65 steps. That is the
classic fp16 underflow/overflow signature, and it is exactly the kind of bug that would otherwise be
discovered 10 hours into the main run. Probe C re-runs the throughput measurements with the correct
recipe β€” **fp32 master weights + `torch.autocast(fp16)` + `torch.amp.GradScaler("cuda")` + grad
clipping** β€” and asserts finiteness, because on T4 there is no bf16 fallback: fp16 with loss scaling
is the only mixed-precision option, so getting it right is not optional.
## 8. What Phase 0 has *not* yet proven
Be honest about the boundary of this gate:
- **Hub writes from inside a Kaggle job** β€” never attempted. No credential path chosen (D-002 open).
- **NCCL across these two T4s** β€” still unproven; probe B's attempt was defeated by its own parser,
and one failed `all_reduce` is not evidence either way. Probe C is fixing that now.
- **Checkpoint upload β†’ independent verification β†’ local delete β†’ free-space-confirmed (Β§3.13)** β€”
untouched; that is a Phase 3 gate, and it needs the real 100M-model sizing arithmetic first.
- **Cold resume on a fresh instance with an empty disk** β€” Phase 3.
- **Real session caps**: `sessionTimeoutSeconds` was accepted and no session ran long enough to hit
any limit. The actual wall-clock ceiling, idle-timeout behaviour, and preemption policy are
**unknown**, and the multi-week resume plan in Phase 4 depends on them. Probe E
(`ounce100m-p0e-session-cap`) is a ~12.5 h CPU heartbeat running now specifically to find this;
if Kaggle kills it, the last heartbeat is the answer. **Update while writing: the session was still
`RUNNING` at 108 minutes** (submitted 15:38Z, checked 17:26Z), so no CPU cap exists below ~1.8 h. That
already rules out the worst case for Phase 4 β€” a session that dies inside an hour β€” and the build is
sized so one source (~5–40 min) is the unit of lost work either way.
## 9. GPU probe C results β€” the numbers the schedule is built on
Kernel `dodosoomro/ounce100m-p0c-fp16-ddp` v2, GPU image
`gcr.io/kaggle-gpu-images/python@sha256:37c64f7d…7d461`, 2x Tesla T4, charged 157.704 s of quota.
Probe B's two void measurements are now replaced. Model shape used β€” deliberately near the likely
main-run shape so the arithmetic is about the real run and not a toy:
`hidden 768, layers 12, heads 12, GQA kv_heads 3, FFN 2048, vocab 50257, tied embeddings, seq 1024`
β†’ **112,934,400 parameters**, of which 38,597,376 are the (tied) embedding.
| measurement | result |
|---|---|
| fp32 master weights + `autocast(fp16)` + `GradScaler` + clip | **loss 9.988, finite** β€” the recipe works; probe B's NaN was the recipe, not the hardware |
| bf16 `autocast` on cc 7.5 | **runs** (loss 10.966, first fwd 0.36 s) β€” i.e. *emulated*, not tensor-cored |
| tokens/s, 1 GPU, sdpa, bs4Γ—1024 | 6,478 |
| tokens/s, 1 GPU, **eager**, bs4Γ—1024 | **7,239** |
| tokens/s per rank, DDP 2 ranks, sdpa, bs4Γ—1024 | 5,531 β†’ **11,062 aggregate** |
| DDP scaling efficiency | 11,062 / 6,478 = **1.71x** (85 % of ideal) |
| NCCL across these two T4s | **works** β€” `rc 0`, `backend nccl`, `world 2`, both ranks finite, 30 steps |
| peak CUDA memory, bs4Γ—1024, 1 GPU | **11.51 GB** (sdpa) / 13.79 GB (eager) of 14.56 GB usable |
| peak CUDA memory, bs4Γ—1024, DDP | 8.68 GB per rank |
| micro-batch ceiling at this shape | **bs 8 OOMs** (`Tried to allocate 1.54 GiB… 520 MiB free`) |
### What each line forces
- **`eager` beat `sdpa` here (7,239 vs 6,478 tok/s).** At `seq_len 1024` the memory-efficient kernel's
overhead is not repaid, and flash-attention is unavailable on sm 7.5 anyway. Do not assume "flash is
the fast one" β€” Phase 3 must re-measure at the frozen sequence length, because the crossover is
length-dependent.
- **Memory, not compute, is the binding constraint on batch size.** bs4 fits in 11.5 GB of 14.56 GB;
bs8 does not fit at all. So global batch must be built with **gradient accumulation**, not larger
micro-batches β€” which also means the usual "increase batch until full" advice is unavailable here.
The fp32 AdamW state for a 100M model (β‰ˆ1.2 GB) plus fp32 master weights is a fixed floor that
does not shrink with batch size.
- **85 % DDP efficiency is good enough to plan on**, and is the number the ETA below uses. It was
measured on 30 steps with NCCL over PCIe between two T4s; treat it as optimistic-typical, not a
guaranteed constant.
- **bf16 running is a trap, not a permission.** Β§2 and Β§8 both rule it out. On Turing it is emulated,
so a config that silently selects bf16 will "work" and be slow. Precision must be asserted to be
fp16 in the frozen config, not left to a library default.
### First ETA, from measurement rather than hope
`1e9 tokens Γ· 11,062 tok/s = 90,400 s β‰ˆ 25.1 wall-clock hours` at the measured aggregate rate.
Against a 30 h weekly allowance accruing at 1x, **the main run fits inside one week of quota on this
shape β€” but with only ~5 h of headroom**, which is not enough to survive both checkpoint upload time
and the multi-session restarts that Β§4/Phase 4 assume. Three honest caveats on that number:
1. It is a 30-step measurement on one shape that is **over the parameter budget** (112.9 M > 110 M),
so the frozen config will differ and throughput with it.
2. It excludes checkpoint write/upload/verify cycles, `DataLoader` stalls, and any
tokenisation-time-vs-disk tradeoff in the input pipeline.
3. It assumes sessions can run long enough to be efficient; the cap is unknown until probe E reports.
Planning stance for `docs/01-plan.md`: treat **~9–11k tok/s aggregate** as the working band, size the
token target from the *low* end (~28–31 h for 1 B tokens β†’ exceeds a single week, so plan for a
two-week run), and revisit once Gate 3 measures the frozen configuration. Do not let the 25 h figure
become the plan's basis, because it leaves no room for the interruption budget that Β§3.1 explicitly
expects to be spent.
## 10. Credentials and the Kaggle-side secret store
`dodosoomro/ounce100m-p1-credential-path` (CPU, read-only) established:
- Internet-enabled Kaggle sessions **arrive pre-authenticated against the Kaggle API**. The `kaggle`
CLI **2.0.2 is preinstalled**, and `kaggle config view` reports `username: dodosoomro`,
`auth_method: ACCESS_TOKEN`, config from `/root/.config/kaggle` β€” with no `kaggle.json` written by
us. `kaggle kernels list --mine` returns real rows, so the principal is genuinely usable, not merely
present. Injected env vars: `KAGGLE_API_V1_TOKEN` (32 ch), `KAGGLE_DATA_PROXY_TOKEN` (473),
`KAGGLE_USER_SECRETS_TOKEN` (245).
- **No Hugging Face credential exists in the container**: `HF_TOKEN` and `HUGGING_FACE_HUB_TOKEN` are
both absent, `HF_HOME` unset, `~/.cache/huggingface` does not exist. An anonymous Hub *write* is
correctly rejected β€” `POST /api/models/…/commit/main` β†’ **401 Unauthorized** β€” while anonymous reads
work. So jobs can pull data freely and cannot push anything until given a token.
- `/kaggle/input` is empty and `/kaggle/input/.secrets/` **does not exist**, so Kaggle's UI-configured
"user input / secrets" facility is unavailable: enabling it is a web-UI action, and there are no
locally usable Kaggle credentials to do it with (`~/.kaggle` absent on this machine; Kaggle auth lives
in the MCP server and inside the container). That rules out option 1 of D-002 for a concrete reason.
- **Consequence: the in-job authenticated Kaggle API is the usable secret store.** A job can create a
**private** Kaggle dataset holding the HF token; later jobs mount it via `datasetDataSources` and read
it from `/kaggle/input/…`, so the token appears in exactly one private kernel revision instead of in
every job's source. Β§2 explicitly sanctions "a secret store", so this is the intended mechanism.
**Ordering used, because the failure mode is unrecoverable.** `kaggle datasets create` derives
visibility from a metadata field whose name has moved between CLI versions, and an unrecognised key is
**silently ignored** β€” which would create a *public* dataset containing the token. So the bootstrap
worker (a) creates the dataset with an inert placeholder only, (b) proves privacy by confirming an
**anonymous** download fails, and (c) uploads the credential as a later version only if (b) passes,
then re-proves. If anonymous read ever succeeds, the secret is never written and the job says so loudly.
**Residual risk, stated:** kernel `p1-secret-bootstrap` v1 necessarily contains the token inline, and
Kaggle retains private kernel revisions β€” overwrite-after-use does not erase history, and there is no
delete-kernel operation in CLI 2.0.2 or in the MCP toolset. This is a strictly smaller exposure than
putting the token in every training job and in anything that mirrors kernel source.
## 11. Pre-existing Kaggle kernels, not ours
`kaggle kernels list --mine` also returned five kernels predating this project
(`quick-python-cpu-smoke-test`, `monte-carlo-cpu-smoke-test`, `cpu-test`, `exam-80-monte-carlo`,
`workbuddy-cpu-probe`; last run 2026-09-17 β†’ 2026-09-19). They are **not touched** (Β§3.10). This also
explains E-002: `search_notebooks` returning `{}` for this account was the search tool failing, not the
account being empty β€” so an empty search must never be read as "no prior work".
## 12. Measured notebook session cap: 12.0 h survived, the instance reported 45,000 s
`dodosoomro/ounce100m-p0e-session-cap` β€” a CPU kernel that printed one heartbeat every two minutes and
nothing else, costing zero GPU quota, left running since 2026-09-19T15:38Z purely as a measurement.
```
CAP_START {"epoch": 1789832305.41, "host": "096ef365b02e", "pid": 7, "limit": 45000}
HB 1 elapsed=0 free_GB=19.5 mem_avail_MB=31204 load=0.41 net=204
HB 361 elapsed=43204 free_GB=19.5 mem_avail_MB=31288 load=0.00 net=204
```
**361 heartbeats, 43,204 s = 12.0 h of continuous running, then the platform took it**
(`CANCEL_ACKNOWLEDGED`, not a crash β€” the last heartbeat is clean and disk/RAM are unchanged at 19.5 GB
free and 31.3 GB available). The instance itself reported `limit: 45000` s, i.e. **12.5 h**, so the kill
arrived at ~96 % of its own stated budget.
What this does and does not license:
- It is a **CPU** measurement. The GPU accelerator quota is tracked separately, and nothing here proves a
GPU session is allowed to run 12.5 h. So Phase 4 plans `SESSION_GPU_HOURS=6.9` for session 1 β€” well
inside every plausible cap β€” and can extend later sessions only on evidence from earlier ones.
- It does settle the question that the plan had been hedging since Β§2 of `04-run-log.md` ("the session cap
is unmeasured"): an interruption at 6 h is not the platform's limit, so a long session is not doomed, and
a `--stop-after-steps` schedule sized for 6-7 h is conservative rather than necessary.
- It confirms the working volume is stable for the whole duration β€” `free_GB` never moved from 19.5, so an
11 GB drift over a session is not a thing to plan for, and the disk floor in the launcher is about the
mix and checkpoints, not about slow leaks.