Download docs/00-platform-notes.md from Cion-lab/ounce100m-code: direct link, hf CLI and curl.
- Browser
- Download file 23.7 kB
-
https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/docs/00-platform-notes.md
- Command line
-
hf download hf://Cion-lab/ounce100m-code/docs/00-platform-notes.md
-
curl -L -o 00-platform-notes.md https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/docs/00-platform-notes.md
00 β Platform notes (Phase 0)
Observed mechanics of the two execution environments this project spans: this workspace and the Kaggle image. Everything here was measured on 2026-09-19, not recalled. Where a claim contradicts prior documentation (including the Kaggle skill's own notes), the contradiction is called out, because a future session will otherwise trust the older statement.
Raw probe JSON lives in the run logs of the kernels listed in memory/ASSETS.md.
1. The two environments are not the same machine
| this workspace | Kaggle CPU session | Kaggle GPU session (2xT4) | |
|---|---|---|---|
| OS | Windows 11, git-bash | Linux 6.12.90+, glibc 2.35 | Linux 6.12.90+, glibc 2.35 |
| Python | 3.11 (and 3.14) | 3.12.13 | 3.12.13 |
| Docker image | n/a | gcr.io/kaggle-images/python@sha256:dafd4ce5β¦c40b9 |
gcr.io/kaggle-gpu-images/python@sha256:37c64f7dβ¦7d461 |
| torch | β | 2.10.0+cpu | 2.10.0+cu128 |
| internet | β | yes (needs enableInternet: true) |
yes (needs enableInternet: true) |
| GPU | none | device_count()==0 |
2x Tesla T4, cc 7.5, 40 SM, 14.56 GiB usable each |
Consequence: CPU and GPU sessions are different images, not one image with a GPU attached. The
CPU image's torch is a +cpu build, so a script that imports torch.cuda will work on the GPU
shape and silently do nothing useful on the CPU shape. Pin these digests in any job that must be
reproducible; a "latest" Kaggle image can move under the run.
huggingface_hub is 1.32.0 locally and reported as huggingface-hub 1.11.0 inside the Kaggle
image. Local code and in-job code are therefore on different minor versions β an API that exists
locally may not exist in the job (create_repo(tags=β¦) is one such example; it raised
TypeError locally). Check the target environment's version before using a new API.
2. Job mechanics that the main run will depend on
- Submission is one call.
save_notebookwithkernelExecutionType: "SaveAndRunAll"creates and runs.kernelType: "script"andmachineShapemust both be present or the call fails with a message-less error.machineShape: "GPU"+enableGpu: trueis accepted and yields 2xT4 β no T4-specific shape string is needed, and none should be invented. newTitlerewrites the slug. KeptnewTitleequal to the slug suffix in every probe so the returnedrefmatched what was requested. Still: poll the returned slug, never the requested one.- Latency. A
30 s CPU script went submit β60β90 s of queue+boot per job**, and don't poll sooner than ~60 s. Both figures are single samples from an uncongested account at ~15:20 UTC.COMPLETEin ~85 s; a ~58 s GPU script in ~145 s. So budget ** - Where output lands.
list_notebook_session_outputreturnsfiles[](signed URLs, expiry unknown β download promptly) andlog, a JSON array of{stream_name, time, data}wheretimeis seconds since session start, one array element per line. The log carries the whole stdout, so a probe that prints its result as JSON needs no file download. Only files written under/kaggle/workingare returned β the probes deliberately wrote nothing to/tmp, because/tmplooks enormous (Β§4) and is therefore the easiest place to silently lose an artifact. - Kernel logs are not secret-safe. Anything printed is retained in the kernel revision. No token may be printed, and by extension no token may be in the kernel source, since source and logs are both Kaggle-held.
CUDA_VISIBLE_DEVICESis unset in both shapes β it is absent from the environment entirely, not set to the string"None". The Kaggle skill documentsCUDA_VISIBLE_DEVICES=Nonefor CPU runs; that is wrong for this image (E-002 inmemory/ERRORS.md), and code that branches on== "None"will take the GPU branch on a CPU box. Gate ontorch.cuda.is_available()ortorch.cuda.device_count()instead.
3. Quota accounting
- Readout is seconds:
total_time_allowed: "108000s"= 30 h,quota_refresh_time: 2026-09-26T00:00:00Z. Weekly reset is Saturday 00:00 UTC. - CPU sessions cost zero GPU quota, confirmed: two CPU probes ran and
time_usedstayed0s. - A 2xT4 session charged 66.411 s for 58.4 s of script time β ~1x wall-clock, not 2x.
Full reasoning, the caveat, and the second sample pending in
memory/QUOTA.md. This is the single most schedule-relevant number measured in Phase 0: at 1x, 30 h/week is 30 wall-clock hours with both cards, not the 15 that Β§2's "shared across both cards" phrasing suggests.
4. Disk and memory β the assumption in Β§2 is only half right
| Mount | Total | Free | Notes |
|---|---|---|---|
/kaggle/working, /kaggle/input |
19.5 GB | 19.5 GB | the small Kaggle disk; only this is retrieved as output |
/dev/loop1 β /kaggle/src |
20 GB | 20 GB | separate loop device |
/, /tmp, /usr, /root/.cache (overlay) |
7.9 TB | 1.1 TB | shared host filesystem, 88 % used by other tenants |
/dev/shm |
14 GB | 14 GB | useful for DataLoader worker traffic |
| RAM | 32 GB (cgroup memory.max = 30 GiB) |
31.2 GB available | SwapTotal: 0 |
| CPU | 4 vCPU (cpu.max = 400000 100000), Intel Xeon @ 2.20 GHz |
Read this carefully before planning the checkpoint cycle. Β§2 says "Kaggle disk is small", and for
/kaggle/working that is exactly true: 19.5 GB. But the root overlay reports ~1 TB free, and a naive
optimisation would write caches or staged shards there. Two reasons not to:
- The overlay is a shared host volume already at 88 %, so "1 TB free" is other tenants' headroom and can disappear mid-run. Β§3.13's "never let it fill the disk" must be enforced against a number we do not control.
- Nothing outside
/kaggle/workingis retrievable when the session ends, so it is not storage in any sense that matters.
Default HF_HOME/datasets cache resolves under /root/.cache, i.e. the overlay, not the 19.5 GB
volume. So a training job's apparent free space can be large while /kaggle/working fills up β
the failure mode Β§3.13 warns about is invisible to a naive disk_usage("/") check. Every job must
measure the specific filesystem it writes to, and log free space for that path. On the CPU shape the
effective budget is ~19.5 GB; treat that as the design number for shards resident at once.
No swap: a tokenizer build or a large in-memory dedup on a 30 GiB cgroup must be sized, not streamed optimistically.
5. Network behaviour, including one thing that does not resolve
Working, measured from inside a GPU and a CPU session:
huggingface.co/api/models β 200 in 149 ms; pypi.org/simple/ β 200 in 73 ms;
datasets/β¦/resolve/main/README.md β 200 via api/resolve-cache/β¦;
datasets.load_dataset("wikimedia/wikipedia", "20231101.en", streaming=True) β first row in 7.4 s,
unauthenticated. pip install from PyPI works: xxhash in 5.8 s, import verified.
Does not resolve: cdn-lfs.huggingface.co, cdn-lfs-us-1.huggingface.co, xethub.huggingface.co
β gaierror. That looked alarming, so probe D tested it directly by pulling 300 MB from two large
public parquet files.
Sustained Hub download throughput from inside a session: 72.8 MB/s (wikimedia/wikipedia, 745 MB
file) and 88.9 MB/s (HuggingFaceFW/fineweb-edu, 2334 MB file), with 0.8β0.94 s time-to-first-byte.
So the non-resolving LFS hostnames do not block bulk transfer β resolve/ redirects land on hosts that
do resolve. Ingest of the mix is not network-bound: at a conservative 70 MB/s, one billion tokens
(β3.6 GB of raw text at ~3.6 chars/token, or ~4β6 GB as parquet) downloads in single-digit minutes.
Phase 2 should plan shard staging around CPU and disk limits, not bandwidth.
Also measured in the same probe: datasets streaming at 139.5 rows/s β 3.0 MB of text per second
single-process, unauthenticated (83 rows/s / 1.8 MB/s cold, ~3 MB/s warm). Unauthenticated Hub access
works but the API warns about rate limits (Please set a HF_TOKEN); at multi-hundred-shard scale that
is a plausible 429 source, which is an argument for the token-in-job path (D-002) for the data plane
even if the checkpoint plane stays local.
Enumerating repo files without a client library: GET https://huggingface.co/api/datasets/<id>/tree/main?recursive=true
returns path/size/type and works with plain urllib. It appeared to cap at 1000 entries
(fineweb-edu returned exactly 1000), so assume pagination is required for a large repo.
CPU-side ingest throughput, and what it rules out
Measured on the Kaggle CPU shape (4 vCPU Xeon @ 2.2 GHz) with the gpt2 tokenizer over 6 MB of real
Wikipedia English, tokenizers 0.22.2 Tokenizer.encode_batch:
| cores | tokens/s | hours to tokenize 1B tokens |
|---|---|---|
| 1 | 245,660 | 1.13 |
| 2 | 492,191 | 0.56 |
| 4 | 660,943 | 0.42 |
chars_per_token 4.53 on this corpus. Scaling is near-linear to 2 cores and then sub-linear
(2.7x at 4 cores), which is what a GIL-bound parent feeding rayon workers looks like β so a Phase 2
pipeline that shards by process, not by thread, should recover the rest. datasets streaming ran
at 151.6 rows/s β 3.26 MB of text per second single-process, unauthenticated.
The conclusion is a negative result, which is the useful kind: tokenizing 1B tokens costs ~0.4 h of
the 4-core budget, and at 35β89 MB/s the raw bytes arrive faster than datasets can stream them.
Phase 2 is therefore not CPU-bound and not bandwidth-bound β it is bound by dedup memory and by the
30 GiB cgroup with no swap. Budget the design accordingly: the dedup/audit stage must be written to
stream and shard against fixed RAM, not to hold a billion sketches, and the only real wall-clock cost
in building the mix is the exact-duplicate/n-gram overlap pass. Note also that download throughput
varied by 2.5x between identical runs (36.9 β 72.8 β 63.2 MB/s), so it should be measured inside the
actual ingest job rather than treated as a constant.
6. Accelerator capability envelope (Turing, cc 7.5)
Established inside a GPU session, not assumed:
torch.compileworks (warm + 20 stepsof an MLP in 6.0 s, incl. first-call compile).- Flash-attention is not available.
SDPBackend.FLASH_ATTENTIONraisesNo available kernel. Aborting execution., preceded byFlash attention only supports gpu architectures in the range [sm80, sm121]. Attempting to run on a sm 7.5 gpu.β so this is architectural, not a missing dependency, and installingflash-attncannot fix it. Β§3's "never touch Hopper-only kernels" now has a concrete instance. EFFICIENT_ATTENTIONandMATHSDPA backends both run. Memory-efficient attention is the viable fast path; note torch warns "Memory efficient attention has been runtime disabled" when flash is force-requested, so backend selection must be explicit.torch.cuda.get_arch_list()includessm_75(also 70, 80, 86, 90, 100, 120) β this torch build does have kernels for the card.- CUDA 12.8, cuDNN 9.10.2, NCCL 2.27.5 present;
triton 3.6.0present. - Absent from the image:
flash-attn,xformers,trl,deepspeed,bitsandbytes,liger-kernel,evaluate. Present:transformers 5.0.0,datasets 5.0.0,accelerate 1.13.0,tokenizers 0.22.2,safetensors 0.7.0,numpy 2.0.2,pyarrow 24.0.0,peft 0.19.1. - bf16 is a hardware question, not a torch one:
torch.bfloat16tensors exist, and bf16 on cc 7.5 is at best emulated. Probe C records whether an autocast-bf16 forward even completes and what it costs, so the Phase 1 precision decision is evidence-backed rather than a recitation of "no bf16". nvidia-smihas nomultiprocessor_countquery field (rc=2); usetorch.cuda.β¦ get_device_properties(i).multi_processor_count(40/SM) instead.
7. fp16 training trap found by accident
Probe B ran a small Llama with weights cast to fp16 (model.cuda().to(torch.float16)) under
autocast with a plain AdamW and no GradScaler: every loss came back NaN in 65 steps. That is the
classic fp16 underflow/overflow signature, and it is exactly the kind of bug that would otherwise be
discovered 10 hours into the main run. Probe C re-runs the throughput measurements with the correct
recipe β fp32 master weights + torch.autocast(fp16) + torch.amp.GradScaler("cuda") + grad
clipping β and asserts finiteness, because on T4 there is no bf16 fallback: fp16 with loss scaling
is the only mixed-precision option, so getting it right is not optional.
8. What Phase 0 has not yet proven
Be honest about the boundary of this gate:
- Hub writes from inside a Kaggle job β never attempted. No credential path chosen (D-002 open).
- NCCL across these two T4s β still unproven; probe B's attempt was defeated by its own parser,
and one failed
all_reduceis not evidence either way. Probe C is fixing that now. - Checkpoint upload β independent verification β local delete β free-space-confirmed (Β§3.13) β untouched; that is a Phase 3 gate, and it needs the real 100M-model sizing arithmetic first.
- Cold resume on a fresh instance with an empty disk β Phase 3.
- Real session caps:
sessionTimeoutSecondswas accepted and no session ran long enough to hit any limit. The actual wall-clock ceiling, idle-timeout behaviour, and preemption policy are unknown, and the multi-week resume plan in Phase 4 depends on them. Probe E (ounce100m-p0e-session-cap) is a12.5 h CPU heartbeat running now specifically to find this; if Kaggle kills it, the last heartbeat is the answer. Update while writing: the session was still5β40 min) is the unit of lost work either way.RUNNINGat 108 minutes (submitted 15:38Z, checked 17:26Z), so no CPU cap exists below ~1.8 h. That already rules out the worst case for Phase 4 β a session that dies inside an hour β and the build is sized so one source (
9. GPU probe C results β the numbers the schedule is built on
Kernel dodosoomro/ounce100m-p0c-fp16-ddp v2, GPU image
gcr.io/kaggle-gpu-images/python@sha256:37c64f7dβ¦7d461, 2x Tesla T4, charged 157.704 s of quota.
Probe B's two void measurements are now replaced. Model shape used β deliberately near the likely
main-run shape so the arithmetic is about the real run and not a toy:
hidden 768, layers 12, heads 12, GQA kv_heads 3, FFN 2048, vocab 50257, tied embeddings, seq 1024
β 112,934,400 parameters, of which 38,597,376 are the (tied) embedding.
| measurement | result |
|---|---|
fp32 master weights + autocast(fp16) + GradScaler + clip |
loss 9.988, finite β the recipe works; probe B's NaN was the recipe, not the hardware |
bf16 autocast on cc 7.5 |
runs (loss 10.966, first fwd 0.36 s) β i.e. emulated, not tensor-cored |
| tokens/s, 1 GPU, sdpa, bs4Γ1024 | 6,478 |
| tokens/s, 1 GPU, eager, bs4Γ1024 | 7,239 |
| tokens/s per rank, DDP 2 ranks, sdpa, bs4Γ1024 | 5,531 β 11,062 aggregate |
| DDP scaling efficiency | 11,062 / 6,478 = 1.71x (85 % of ideal) |
| NCCL across these two T4s | works β rc 0, backend nccl, world 2, both ranks finite, 30 steps |
| peak CUDA memory, bs4Γ1024, 1 GPU | 11.51 GB (sdpa) / 13.79 GB (eager) of 14.56 GB usable |
| peak CUDA memory, bs4Γ1024, DDP | 8.68 GB per rank |
| micro-batch ceiling at this shape | bs 8 OOMs (Tried to allocate 1.54 GiB⦠520 MiB free) |
What each line forces
eagerbeatsdpahere (7,239 vs 6,478 tok/s). Atseq_len 1024the memory-efficient kernel's overhead is not repaid, and flash-attention is unavailable on sm 7.5 anyway. Do not assume "flash is the fast one" β Phase 3 must re-measure at the frozen sequence length, because the crossover is length-dependent.- Memory, not compute, is the binding constraint on batch size. bs4 fits in 11.5 GB of 14.56 GB; bs8 does not fit at all. So global batch must be built with gradient accumulation, not larger micro-batches β which also means the usual "increase batch until full" advice is unavailable here. The fp32 AdamW state for a 100M model (β1.2 GB) plus fp32 master weights is a fixed floor that does not shrink with batch size.
- 85 % DDP efficiency is good enough to plan on, and is the number the ETA below uses. It was measured on 30 steps with NCCL over PCIe between two T4s; treat it as optimistic-typical, not a guaranteed constant.
- bf16 running is a trap, not a permission. Β§2 and Β§8 both rule it out. On Turing it is emulated, so a config that silently selects bf16 will "work" and be slow. Precision must be asserted to be fp16 in the frozen config, not left to a library default.
First ETA, from measurement rather than hope
1e9 tokens Γ· 11,062 tok/s = 90,400 s β 25.1 wall-clock hours at the measured aggregate rate.
Against a 30 h weekly allowance accruing at 1x, the main run fits inside one week of quota on this shape β but with only ~5 h of headroom, which is not enough to survive both checkpoint upload time and the multi-session restarts that Β§4/Phase 4 assume. Three honest caveats on that number:
- It is a 30-step measurement on one shape that is over the parameter budget (112.9 M > 110 M), so the frozen config will differ and throughput with it.
- It excludes checkpoint write/upload/verify cycles,
DataLoaderstalls, and any tokenisation-time-vs-disk tradeoff in the input pipeline. - It assumes sessions can run long enough to be efficient; the cap is unknown until probe E reports.
Planning stance for docs/01-plan.md: treat ~9β11k tok/s aggregate as the working band, size the
token target from the low end (~28β31 h for 1 B tokens β exceeds a single week, so plan for a
two-week run), and revisit once Gate 3 measures the frozen configuration. Do not let the 25 h figure
become the plan's basis, because it leaves no room for the interruption budget that Β§3.1 explicitly
expects to be spent.
10. Credentials and the Kaggle-side secret store
dodosoomro/ounce100m-p1-credential-path (CPU, read-only) established:
- Internet-enabled Kaggle sessions arrive pre-authenticated against the Kaggle API. The
kaggleCLI 2.0.2 is preinstalled, andkaggle config viewreportsusername: dodosoomro,auth_method: ACCESS_TOKEN, config from/root/.config/kaggleβ with nokaggle.jsonwritten by us.kaggle kernels list --minereturns real rows, so the principal is genuinely usable, not merely present. Injected env vars:KAGGLE_API_V1_TOKEN(32 ch),KAGGLE_DATA_PROXY_TOKEN(473),KAGGLE_USER_SECRETS_TOKEN(245). - No Hugging Face credential exists in the container:
HF_TOKENandHUGGING_FACE_HUB_TOKENare both absent,HF_HOMEunset,~/.cache/huggingfacedoes not exist. An anonymous Hub write is correctly rejected βPOST /api/models/β¦/commit/mainβ 401 Unauthorized β while anonymous reads work. So jobs can pull data freely and cannot push anything until given a token. /kaggle/inputis empty and/kaggle/input/.secrets/does not exist, so Kaggle's UI-configured "user input / secrets" facility is unavailable: enabling it is a web-UI action, and there are no locally usable Kaggle credentials to do it with (~/.kaggleabsent on this machine; Kaggle auth lives in the MCP server and inside the container). That rules out option 1 of D-002 for a concrete reason.- Consequence: the in-job authenticated Kaggle API is the usable secret store. A job can create a
private Kaggle dataset holding the HF token; later jobs mount it via
datasetDataSourcesand read it from/kaggle/input/β¦, so the token appears in exactly one private kernel revision instead of in every job's source. Β§2 explicitly sanctions "a secret store", so this is the intended mechanism.
Ordering used, because the failure mode is unrecoverable. kaggle datasets create derives
visibility from a metadata field whose name has moved between CLI versions, and an unrecognised key is
silently ignored β which would create a public dataset containing the token. So the bootstrap
worker (a) creates the dataset with an inert placeholder only, (b) proves privacy by confirming an
anonymous download fails, and (c) uploads the credential as a later version only if (b) passes,
then re-proves. If anonymous read ever succeeds, the secret is never written and the job says so loudly.
Residual risk, stated: kernel p1-secret-bootstrap v1 necessarily contains the token inline, and
Kaggle retains private kernel revisions β overwrite-after-use does not erase history, and there is no
delete-kernel operation in CLI 2.0.2 or in the MCP toolset. This is a strictly smaller exposure than
putting the token in every training job and in anything that mirrors kernel source.
11. Pre-existing Kaggle kernels, not ours
kaggle kernels list --mine also returned five kernels predating this project
(quick-python-cpu-smoke-test, monte-carlo-cpu-smoke-test, cpu-test, exam-80-monte-carlo,
workbuddy-cpu-probe; last run 2026-09-17 β 2026-09-19). They are not touched (Β§3.10). This also
explains E-002: search_notebooks returning {} for this account was the search tool failing, not the
account being empty β so an empty search must never be read as "no prior work".
12. Measured notebook session cap: 12.0 h survived, the instance reported 45,000 s
dodosoomro/ounce100m-p0e-session-cap β a CPU kernel that printed one heartbeat every two minutes and
nothing else, costing zero GPU quota, left running since 2026-09-19T15:38Z purely as a measurement.
CAP_START {"epoch": 1789832305.41, "host": "096ef365b02e", "pid": 7, "limit": 45000}
HB 1 elapsed=0 free_GB=19.5 mem_avail_MB=31204 load=0.41 net=204
HB 361 elapsed=43204 free_GB=19.5 mem_avail_MB=31288 load=0.00 net=204
361 heartbeats, 43,204 s = 12.0 h of continuous running, then the platform took it
(CANCEL_ACKNOWLEDGED, not a crash β the last heartbeat is clean and disk/RAM are unchanged at 19.5 GB
free and 31.3 GB available). The instance itself reported limit: 45000 s, i.e. 12.5 h, so the kill
arrived at ~96 % of its own stated budget.
What this does and does not license:
- It is a CPU measurement. The GPU accelerator quota is tracked separately, and nothing here proves a
GPU session is allowed to run 12.5 h. So Phase 4 plans
SESSION_GPU_HOURS=6.9for session 1 β well inside every plausible cap β and can extend later sessions only on evidence from earlier ones. - It does settle the question that the plan had been hedging since Β§2 of
04-run-log.md("the session cap is unmeasured"): an interruption at 6 h is not the platform's limit, so a long session is not doomed, and a--stop-after-stepsschedule sized for 6-7 h is conservative rather than necessary. - It confirms the working volume is stable for the whole duration β
free_GBnever moved from 19.5, so an 11 GB drift over a session is not a thing to plan for, and the disk floor in the launcher is about the mix and checkpoints, not about slow leaks.