ounce100m-code / docs /00-platform-notes.md
Cion-lab's picture
publish docs/: plan, mix rationale, preflight report, run log, frozen eval protocol, final report
f345921 verified
|
Raw History Blame Contribute Delete
23.7 kB

00 β€” Platform notes (Phase 0)

Observed mechanics of the two execution environments this project spans: this workspace and the Kaggle image. Everything here was measured on 2026-09-19, not recalled. Where a claim contradicts prior documentation (including the Kaggle skill's own notes), the contradiction is called out, because a future session will otherwise trust the older statement.

Raw probe JSON lives in the run logs of the kernels listed in memory/ASSETS.md.


1. The two environments are not the same machine

this workspace Kaggle CPU session Kaggle GPU session (2xT4)
OS Windows 11, git-bash Linux 6.12.90+, glibc 2.35 Linux 6.12.90+, glibc 2.35
Python 3.11 (and 3.14) 3.12.13 3.12.13
Docker image n/a gcr.io/kaggle-images/python@sha256:dafd4ce5…c40b9 gcr.io/kaggle-gpu-images/python@sha256:37c64f7d…7d461
torch β€” 2.10.0+cpu 2.10.0+cu128
internet β€” yes (needs enableInternet: true) yes (needs enableInternet: true)
GPU none device_count()==0 2x Tesla T4, cc 7.5, 40 SM, 14.56 GiB usable each

Consequence: CPU and GPU sessions are different images, not one image with a GPU attached. The CPU image's torch is a +cpu build, so a script that imports torch.cuda will work on the GPU shape and silently do nothing useful on the CPU shape. Pin these digests in any job that must be reproducible; a "latest" Kaggle image can move under the run.

huggingface_hub is 1.32.0 locally and reported as huggingface-hub 1.11.0 inside the Kaggle image. Local code and in-job code are therefore on different minor versions β€” an API that exists locally may not exist in the job (create_repo(tags=…) is one such example; it raised TypeError locally). Check the target environment's version before using a new API.

2. Job mechanics that the main run will depend on

  • Submission is one call. save_notebook with kernelExecutionType: "SaveAndRunAll" creates and runs. kernelType: "script" and machineShape must both be present or the call fails with a message-less error. machineShape: "GPU" + enableGpu: true is accepted and yields 2xT4 β€” no T4-specific shape string is needed, and none should be invented.
  • newTitle rewrites the slug. Kept newTitle equal to the slug suffix in every probe so the returned ref matched what was requested. Still: poll the returned slug, never the requested one.
  • Latency. A 30 s CPU script went submit β†’ COMPLETE in ~85 s; a ~58 s GPU script in ~145 s. So budget **60–90 s of queue+boot per job**, and don't poll sooner than ~60 s. Both figures are single samples from an uncongested account at ~15:20 UTC.
  • Where output lands. list_notebook_session_output returns files[] (signed URLs, expiry unknown β€” download promptly) and log, a JSON array of {stream_name, time, data} where time is seconds since session start, one array element per line. The log carries the whole stdout, so a probe that prints its result as JSON needs no file download. Only files written under /kaggle/working are returned β€” the probes deliberately wrote nothing to /tmp, because /tmp looks enormous (Β§4) and is therefore the easiest place to silently lose an artifact.
  • Kernel logs are not secret-safe. Anything printed is retained in the kernel revision. No token may be printed, and by extension no token may be in the kernel source, since source and logs are both Kaggle-held.
  • CUDA_VISIBLE_DEVICES is unset in both shapes β€” it is absent from the environment entirely, not set to the string "None". The Kaggle skill documents CUDA_VISIBLE_DEVICES=None for CPU runs; that is wrong for this image (E-002 in memory/ERRORS.md), and code that branches on == "None" will take the GPU branch on a CPU box. Gate on torch.cuda.is_available() or torch.cuda.device_count() instead.

3. Quota accounting

  • Readout is seconds: total_time_allowed: "108000s" = 30 h, quota_refresh_time: 2026-09-26T00:00:00Z. Weekly reset is Saturday 00:00 UTC.
  • CPU sessions cost zero GPU quota, confirmed: two CPU probes ran and time_used stayed 0s.
  • A 2xT4 session charged 66.411 s for 58.4 s of script time β†’ ~1x wall-clock, not 2x. Full reasoning, the caveat, and the second sample pending in memory/QUOTA.md. This is the single most schedule-relevant number measured in Phase 0: at 1x, 30 h/week is 30 wall-clock hours with both cards, not the 15 that Β§2's "shared across both cards" phrasing suggests.

4. Disk and memory β€” the assumption in Β§2 is only half right

Mount Total Free Notes
/kaggle/working, /kaggle/input 19.5 GB 19.5 GB the small Kaggle disk; only this is retrieved as output
/dev/loop1 β†’ /kaggle/src 20 GB 20 GB separate loop device
/, /tmp, /usr, /root/.cache (overlay) 7.9 TB 1.1 TB shared host filesystem, 88 % used by other tenants
/dev/shm 14 GB 14 GB useful for DataLoader worker traffic
RAM 32 GB (cgroup memory.max = 30 GiB) 31.2 GB available SwapTotal: 0
CPU 4 vCPU (cpu.max = 400000 100000), Intel Xeon @ 2.20 GHz

Read this carefully before planning the checkpoint cycle. Β§2 says "Kaggle disk is small", and for /kaggle/working that is exactly true: 19.5 GB. But the root overlay reports ~1 TB free, and a naive optimisation would write caches or staged shards there. Two reasons not to:

  1. The overlay is a shared host volume already at 88 %, so "1 TB free" is other tenants' headroom and can disappear mid-run. Β§3.13's "never let it fill the disk" must be enforced against a number we do not control.
  2. Nothing outside /kaggle/working is retrievable when the session ends, so it is not storage in any sense that matters.

Default HF_HOME/datasets cache resolves under /root/.cache, i.e. the overlay, not the 19.5 GB volume. So a training job's apparent free space can be large while /kaggle/working fills up β€” the failure mode Β§3.13 warns about is invisible to a naive disk_usage("/") check. Every job must measure the specific filesystem it writes to, and log free space for that path. On the CPU shape the effective budget is ~19.5 GB; treat that as the design number for shards resident at once.

No swap: a tokenizer build or a large in-memory dedup on a 30 GiB cgroup must be sized, not streamed optimistically.

5. Network behaviour, including one thing that does not resolve

Working, measured from inside a GPU and a CPU session: huggingface.co/api/models β†’ 200 in 149 ms; pypi.org/simple/ β†’ 200 in 73 ms; datasets/…/resolve/main/README.md β†’ 200 via api/resolve-cache/…; datasets.load_dataset("wikimedia/wikipedia", "20231101.en", streaming=True) β†’ first row in 7.4 s, unauthenticated. pip install from PyPI works: xxhash in 5.8 s, import verified.

Does not resolve: cdn-lfs.huggingface.co, cdn-lfs-us-1.huggingface.co, xethub.huggingface.co β†’ gaierror. That looked alarming, so probe D tested it directly by pulling 300 MB from two large public parquet files.

Sustained Hub download throughput from inside a session: 72.8 MB/s (wikimedia/wikipedia, 745 MB file) and 88.9 MB/s (HuggingFaceFW/fineweb-edu, 2334 MB file), with 0.8–0.94 s time-to-first-byte. So the non-resolving LFS hostnames do not block bulk transfer β€” resolve/ redirects land on hosts that do resolve. Ingest of the mix is not network-bound: at a conservative 70 MB/s, one billion tokens (β‰ˆ3.6 GB of raw text at ~3.6 chars/token, or ~4–6 GB as parquet) downloads in single-digit minutes. Phase 2 should plan shard staging around CPU and disk limits, not bandwidth.

Also measured in the same probe: datasets streaming at 139.5 rows/s β‰ˆ 3.0 MB of text per second single-process, unauthenticated (83 rows/s / 1.8 MB/s cold, ~3 MB/s warm). Unauthenticated Hub access works but the API warns about rate limits (Please set a HF_TOKEN); at multi-hundred-shard scale that is a plausible 429 source, which is an argument for the token-in-job path (D-002) for the data plane even if the checkpoint plane stays local.

Enumerating repo files without a client library: GET https://huggingface.co/api/datasets/<id>/tree/main?recursive=true returns path/size/type and works with plain urllib. It appeared to cap at 1000 entries (fineweb-edu returned exactly 1000), so assume pagination is required for a large repo.

CPU-side ingest throughput, and what it rules out

Measured on the Kaggle CPU shape (4 vCPU Xeon @ 2.2 GHz) with the gpt2 tokenizer over 6 MB of real Wikipedia English, tokenizers 0.22.2 Tokenizer.encode_batch:

cores tokens/s hours to tokenize 1B tokens
1 245,660 1.13
2 492,191 0.56
4 660,943 0.42

chars_per_token 4.53 on this corpus. Scaling is near-linear to 2 cores and then sub-linear (2.7x at 4 cores), which is what a GIL-bound parent feeding rayon workers looks like β€” so a Phase 2 pipeline that shards by process, not by thread, should recover the rest. datasets streaming ran at 151.6 rows/s β‰ˆ 3.26 MB of text per second single-process, unauthenticated.

The conclusion is a negative result, which is the useful kind: tokenizing 1B tokens costs ~0.4 h of the 4-core budget, and at 35–89 MB/s the raw bytes arrive faster than datasets can stream them. Phase 2 is therefore not CPU-bound and not bandwidth-bound β€” it is bound by dedup memory and by the 30 GiB cgroup with no swap. Budget the design accordingly: the dedup/audit stage must be written to stream and shard against fixed RAM, not to hold a billion sketches, and the only real wall-clock cost in building the mix is the exact-duplicate/n-gram overlap pass. Note also that download throughput varied by 2.5x between identical runs (36.9 β†’ 72.8 β†’ 63.2 MB/s), so it should be measured inside the actual ingest job rather than treated as a constant.

6. Accelerator capability envelope (Turing, cc 7.5)

Established inside a GPU session, not assumed:

  • torch.compile works (warm + 20 steps of an MLP in 6.0 s, incl. first-call compile).
  • Flash-attention is not available. SDPBackend.FLASH_ATTENTION raises No available kernel. Aborting execution., preceded by Flash attention only supports gpu architectures in the range [sm80, sm121]. Attempting to run on a sm 7.5 gpu. β€” so this is architectural, not a missing dependency, and installing flash-attn cannot fix it. Β§3's "never touch Hopper-only kernels" now has a concrete instance.
  • EFFICIENT_ATTENTION and MATH SDPA backends both run. Memory-efficient attention is the viable fast path; note torch warns "Memory efficient attention has been runtime disabled" when flash is force-requested, so backend selection must be explicit.
  • torch.cuda.get_arch_list() includes sm_75 (also 70, 80, 86, 90, 100, 120) β€” this torch build does have kernels for the card.
  • CUDA 12.8, cuDNN 9.10.2, NCCL 2.27.5 present; triton 3.6.0 present.
  • Absent from the image: flash-attn, xformers, trl, deepspeed, bitsandbytes, liger-kernel, evaluate. Present: transformers 5.0.0, datasets 5.0.0, accelerate 1.13.0, tokenizers 0.22.2, safetensors 0.7.0, numpy 2.0.2, pyarrow 24.0.0, peft 0.19.1.
  • bf16 is a hardware question, not a torch one: torch.bfloat16 tensors exist, and bf16 on cc 7.5 is at best emulated. Probe C records whether an autocast-bf16 forward even completes and what it costs, so the Phase 1 precision decision is evidence-backed rather than a recitation of "no bf16".
  • nvidia-smi has no multiprocessor_count query field (rc=2); use torch.cuda.… get_device_properties(i).multi_processor_count (40/SM) instead.

7. fp16 training trap found by accident

Probe B ran a small Llama with weights cast to fp16 (model.cuda().to(torch.float16)) under autocast with a plain AdamW and no GradScaler: every loss came back NaN in 65 steps. That is the classic fp16 underflow/overflow signature, and it is exactly the kind of bug that would otherwise be discovered 10 hours into the main run. Probe C re-runs the throughput measurements with the correct recipe β€” fp32 master weights + torch.autocast(fp16) + torch.amp.GradScaler("cuda") + grad clipping β€” and asserts finiteness, because on T4 there is no bf16 fallback: fp16 with loss scaling is the only mixed-precision option, so getting it right is not optional.

8. What Phase 0 has not yet proven

Be honest about the boundary of this gate:

  • Hub writes from inside a Kaggle job β€” never attempted. No credential path chosen (D-002 open).
  • NCCL across these two T4s β€” still unproven; probe B's attempt was defeated by its own parser, and one failed all_reduce is not evidence either way. Probe C is fixing that now.
  • Checkpoint upload β†’ independent verification β†’ local delete β†’ free-space-confirmed (Β§3.13) β€” untouched; that is a Phase 3 gate, and it needs the real 100M-model sizing arithmetic first.
  • Cold resume on a fresh instance with an empty disk β€” Phase 3.
  • Real session caps: sessionTimeoutSeconds was accepted and no session ran long enough to hit any limit. The actual wall-clock ceiling, idle-timeout behaviour, and preemption policy are unknown, and the multi-week resume plan in Phase 4 depends on them. Probe E (ounce100m-p0e-session-cap) is a 12.5 h CPU heartbeat running now specifically to find this; if Kaggle kills it, the last heartbeat is the answer. Update while writing: the session was still RUNNING at 108 minutes (submitted 15:38Z, checked 17:26Z), so no CPU cap exists below ~1.8 h. That already rules out the worst case for Phase 4 β€” a session that dies inside an hour β€” and the build is sized so one source (5–40 min) is the unit of lost work either way.

9. GPU probe C results β€” the numbers the schedule is built on

Kernel dodosoomro/ounce100m-p0c-fp16-ddp v2, GPU image gcr.io/kaggle-gpu-images/python@sha256:37c64f7d…7d461, 2x Tesla T4, charged 157.704 s of quota.

Probe B's two void measurements are now replaced. Model shape used β€” deliberately near the likely main-run shape so the arithmetic is about the real run and not a toy: hidden 768, layers 12, heads 12, GQA kv_heads 3, FFN 2048, vocab 50257, tied embeddings, seq 1024 β†’ 112,934,400 parameters, of which 38,597,376 are the (tied) embedding.

measurement result
fp32 master weights + autocast(fp16) + GradScaler + clip loss 9.988, finite β€” the recipe works; probe B's NaN was the recipe, not the hardware
bf16 autocast on cc 7.5 runs (loss 10.966, first fwd 0.36 s) β€” i.e. emulated, not tensor-cored
tokens/s, 1 GPU, sdpa, bs4Γ—1024 6,478
tokens/s, 1 GPU, eager, bs4Γ—1024 7,239
tokens/s per rank, DDP 2 ranks, sdpa, bs4Γ—1024 5,531 β†’ 11,062 aggregate
DDP scaling efficiency 11,062 / 6,478 = 1.71x (85 % of ideal)
NCCL across these two T4s works β€” rc 0, backend nccl, world 2, both ranks finite, 30 steps
peak CUDA memory, bs4Γ—1024, 1 GPU 11.51 GB (sdpa) / 13.79 GB (eager) of 14.56 GB usable
peak CUDA memory, bs4Γ—1024, DDP 8.68 GB per rank
micro-batch ceiling at this shape bs 8 OOMs (Tried to allocate 1.54 GiB… 520 MiB free)

What each line forces

  • eager beat sdpa here (7,239 vs 6,478 tok/s). At seq_len 1024 the memory-efficient kernel's overhead is not repaid, and flash-attention is unavailable on sm 7.5 anyway. Do not assume "flash is the fast one" β€” Phase 3 must re-measure at the frozen sequence length, because the crossover is length-dependent.
  • Memory, not compute, is the binding constraint on batch size. bs4 fits in 11.5 GB of 14.56 GB; bs8 does not fit at all. So global batch must be built with gradient accumulation, not larger micro-batches β€” which also means the usual "increase batch until full" advice is unavailable here. The fp32 AdamW state for a 100M model (β‰ˆ1.2 GB) plus fp32 master weights is a fixed floor that does not shrink with batch size.
  • 85 % DDP efficiency is good enough to plan on, and is the number the ETA below uses. It was measured on 30 steps with NCCL over PCIe between two T4s; treat it as optimistic-typical, not a guaranteed constant.
  • bf16 running is a trap, not a permission. Β§2 and Β§8 both rule it out. On Turing it is emulated, so a config that silently selects bf16 will "work" and be slow. Precision must be asserted to be fp16 in the frozen config, not left to a library default.

First ETA, from measurement rather than hope

1e9 tokens Γ· 11,062 tok/s = 90,400 s β‰ˆ 25.1 wall-clock hours at the measured aggregate rate.

Against a 30 h weekly allowance accruing at 1x, the main run fits inside one week of quota on this shape β€” but with only ~5 h of headroom, which is not enough to survive both checkpoint upload time and the multi-session restarts that Β§4/Phase 4 assume. Three honest caveats on that number:

  1. It is a 30-step measurement on one shape that is over the parameter budget (112.9 M > 110 M), so the frozen config will differ and throughput with it.
  2. It excludes checkpoint write/upload/verify cycles, DataLoader stalls, and any tokenisation-time-vs-disk tradeoff in the input pipeline.
  3. It assumes sessions can run long enough to be efficient; the cap is unknown until probe E reports.

Planning stance for docs/01-plan.md: treat ~9–11k tok/s aggregate as the working band, size the token target from the low end (~28–31 h for 1 B tokens β†’ exceeds a single week, so plan for a two-week run), and revisit once Gate 3 measures the frozen configuration. Do not let the 25 h figure become the plan's basis, because it leaves no room for the interruption budget that Β§3.1 explicitly expects to be spent.

10. Credentials and the Kaggle-side secret store

dodosoomro/ounce100m-p1-credential-path (CPU, read-only) established:

  • Internet-enabled Kaggle sessions arrive pre-authenticated against the Kaggle API. The kaggle CLI 2.0.2 is preinstalled, and kaggle config view reports username: dodosoomro, auth_method: ACCESS_TOKEN, config from /root/.config/kaggle β€” with no kaggle.json written by us. kaggle kernels list --mine returns real rows, so the principal is genuinely usable, not merely present. Injected env vars: KAGGLE_API_V1_TOKEN (32 ch), KAGGLE_DATA_PROXY_TOKEN (473), KAGGLE_USER_SECRETS_TOKEN (245).
  • No Hugging Face credential exists in the container: HF_TOKEN and HUGGING_FACE_HUB_TOKEN are both absent, HF_HOME unset, ~/.cache/huggingface does not exist. An anonymous Hub write is correctly rejected β€” POST /api/models/…/commit/main β†’ 401 Unauthorized β€” while anonymous reads work. So jobs can pull data freely and cannot push anything until given a token.
  • /kaggle/input is empty and /kaggle/input/.secrets/ does not exist, so Kaggle's UI-configured "user input / secrets" facility is unavailable: enabling it is a web-UI action, and there are no locally usable Kaggle credentials to do it with (~/.kaggle absent on this machine; Kaggle auth lives in the MCP server and inside the container). That rules out option 1 of D-002 for a concrete reason.
  • Consequence: the in-job authenticated Kaggle API is the usable secret store. A job can create a private Kaggle dataset holding the HF token; later jobs mount it via datasetDataSources and read it from /kaggle/input/…, so the token appears in exactly one private kernel revision instead of in every job's source. Β§2 explicitly sanctions "a secret store", so this is the intended mechanism.

Ordering used, because the failure mode is unrecoverable. kaggle datasets create derives visibility from a metadata field whose name has moved between CLI versions, and an unrecognised key is silently ignored β€” which would create a public dataset containing the token. So the bootstrap worker (a) creates the dataset with an inert placeholder only, (b) proves privacy by confirming an anonymous download fails, and (c) uploads the credential as a later version only if (b) passes, then re-proves. If anonymous read ever succeeds, the secret is never written and the job says so loudly.

Residual risk, stated: kernel p1-secret-bootstrap v1 necessarily contains the token inline, and Kaggle retains private kernel revisions β€” overwrite-after-use does not erase history, and there is no delete-kernel operation in CLI 2.0.2 or in the MCP toolset. This is a strictly smaller exposure than putting the token in every training job and in anything that mirrors kernel source.

11. Pre-existing Kaggle kernels, not ours

kaggle kernels list --mine also returned five kernels predating this project (quick-python-cpu-smoke-test, monte-carlo-cpu-smoke-test, cpu-test, exam-80-monte-carlo, workbuddy-cpu-probe; last run 2026-09-17 β†’ 2026-09-19). They are not touched (Β§3.10). This also explains E-002: search_notebooks returning {} for this account was the search tool failing, not the account being empty β€” so an empty search must never be read as "no prior work".

12. Measured notebook session cap: 12.0 h survived, the instance reported 45,000 s

dodosoomro/ounce100m-p0e-session-cap β€” a CPU kernel that printed one heartbeat every two minutes and nothing else, costing zero GPU quota, left running since 2026-09-19T15:38Z purely as a measurement.

CAP_START {"epoch": 1789832305.41, "host": "096ef365b02e", "pid": 7, "limit": 45000}
HB 1   elapsed=0      free_GB=19.5 mem_avail_MB=31204 load=0.41 net=204
HB 361 elapsed=43204  free_GB=19.5 mem_avail_MB=31288 load=0.00 net=204

361 heartbeats, 43,204 s = 12.0 h of continuous running, then the platform took it (CANCEL_ACKNOWLEDGED, not a crash β€” the last heartbeat is clean and disk/RAM are unchanged at 19.5 GB free and 31.3 GB available). The instance itself reported limit: 45000 s, i.e. 12.5 h, so the kill arrived at ~96 % of its own stated budget.

What this does and does not license:

  • It is a CPU measurement. The GPU accelerator quota is tracked separately, and nothing here proves a GPU session is allowed to run 12.5 h. So Phase 4 plans SESSION_GPU_HOURS=6.9 for session 1 β€” well inside every plausible cap β€” and can extend later sessions only on evidence from earlier ones.
  • It does settle the question that the plan had been hedging since Β§2 of 04-run-log.md ("the session cap is unmeasured"): an interruption at 6 h is not the platform's limit, so a long session is not doomed, and a --stop-after-steps schedule sized for 6-7 h is conservative rather than necessary.
  • It confirms the working volume is stable for the whole duration β€” free_GB never moved from 19.5, so an 11 GB drift over a session is not a thing to plan for, and the disk floor in the launcher is about the mix and checkpoints, not about slow leaks.