ounce100m-code / guide /07-platform-notes.md
Cion-lab's picture
publish guide/: the operating manual distilled from 54 failure records, for future agent sessions
6302710 verified
|
Raw History Blame Contribute Delete
7.13 kB

07 — Platform notes

Numbers measured on one account, one week, one hardware generation. Treat them as the shape of the fact to record, not as facts: re-measure anything you're about to rely on. Staleness here is the default state.

Stack: Kaggle for compute (2×Tesla T4), Hugging Face Hub for storage and publication.

Kaggle

Fact Measured
GPU billing model 1× container wall-clock, not accelerator-seconds. A job that idles, installs, or waits in a queue is billed. Start the clock deliberately.
Weekly allowance 108,000 s = 30 GPU-h. Reset Saturday 00:00Z (quota_refresh_time 2026-09-26T00:00:00Z, which is a Saturday — looked up, not assumed). Read it live; don't infer the weekday.
Small-model test cap (self-imposed) 6 GPU-h lifetime for all preflight; closed at 4.63 h
Session ceiling ~25,200 s (7 h) for a GPU session; treat the practical cap as lower and end on a checkpoint with margin
Kernel logs Not available until the session completes. Plan monitoring around remote artefacts instead
Env vars to a job Cannot be passed through create_notebook_session. Workaround: a thin pinned wrapper kernel that runpy.run_paths the real script fetched at a commit sha, and carries only non-secret values (chosen session size, planning rate, quota figure). Secrets never go in kernel source — see 01 §4 and D-006.
Quota read get_accelerator_quota; time_reserved was 0s in every read taken between jobs, so it is not a live-job indicator — use the session-status call for that
Throughput, 106M model, seq 1024, 2×T4, no checkpointing 12,221–12,312 tok/s sustained = 21.5 s/step at 262,144 tok/step. With gradient checkpointing: 9,358–9,696 tok/s
Memory, same config cross-rank max reserved 12.84 GiB / allocated 12.70 GiB, identical at step 10 and step 120 of the same process (i.e. stable over time; per-rank equality was never established — the number is an all_reduce(MAX))

MCP/API mechanics that bit:

  • save_notebook requires language and kernelType (and machineShape for some CPU shapes).
  • Passing newTitle rewrites the kernel slug. Address existing kernels by their current slug and re-read it after any title change; every stale reference breaks silently.
  • dockerImagePinningType accepts latest and will not pin a revision usefully, so reproducibility comes from your commit shas plus explicit pip version pins — not from the image.
  • pip --user installs land under /root/.local/..., which is not on the parent process's sys.path when the directory is created mid-process. Install, then start a fresh interpreter, or add the path explicitly.
  • Big tool outputs persist to a file rather than returning inline; decode the log field with stdlib JSON.
  • A job that exits within ~2 minutes almost always refused to start on its own assertions. That is a good outcome — make refusals print a named verdict line you can grep for.

Hugging Face Hub

Fact Measured
Commit-rate ceiling ~139 sequential commits in a burst starts failing. Batch multi-file pushes into one commit.
Checkpoint cycle, 1.27 GB 7–10 s on the live run, measured as the gap between the checkpoint ckpt/checkpoint-N commit and the latest -> N commit (verification happens in between, so this bounds push + ranged byte-verify + point). The earlier whole-object download design took 797 s off a cold CDN. Verification design sets your cadence.
LFS files lfs["oid"] is the content sha256 — and it's a dict entry, not an object attribute
Small (non-LFS) files expose only a sha1 blob id, not comparable to a sha256. Don't write an equality check that mixes them.
Range requests return 206 with content-range: bytes a-b/size through the LFS redirect. Compare the header to the window you asked for: a server that ignores Range answers 200 with the whole object, which is how a "verified" read-back can prove nothing.
Small-file fetches resolve/main/<small file> answers 307 → /api/resolve-cache/<commit sha>/<path>. A client that doesn't follow redirects reads the redirect stub as the file. The target embeds the head sha — a free consistency check.
Missing repo 404 / RepositoryNotFoundError. A timeout is not "missing", and neither is 403.
403 is ambiguous a private repo and a nonexistent repo both answer 403/404 to the wrong caller. Proving a secret store is private takes two different calls (01 §4).
Anonymous reads public repos are readable with no token, including a checkpoint's trainer_state.json — remote loss monitoring needs no credentials.

Evaluation harness (lm-eval)

Fact Detail
Metric keys shaped "<name>,<aggregation>": acc,none, exact_match,strict-match. Not settleable from config files; assert the key exists in the result dict at smoke time
MMLU grouped across subjects; the per-group value is a list, so len() of the result isn't a row count. Some tasks have two metrics (em + strict-match); pick one and record which
Splits several "zero-shot academic" tasks score on validation; some datasets have no test split at all
lm_eval --tasks list not a listing command in 0.4.13 — ValueError: Tasks not found: list. Use TaskManager().task_index
Optional deps --no-deps installs import fine and then die at run time on sacrebleu. Install with dependencies
Denominators n_input / n_logged live in per-task output JSONs; harvest by walking the output tree, and treat a missing denominator as a failed cell, not a zero
Not measured total eval wall-clock for the eight tasks on 2×T4. Unknown as of the freeze, which is exactly why D-019 puts training ahead of it

Hardware reality check

  • fp16 autocast + fp32 master weights + GradScaler works on Turing. bf16 is unavailable (no Ampere+), so any recipe assuming bf16 or Hopper-only kernels (flash-attn/FA2) is out. Eager attention is the price.
  • Optimizer state dominates the artefact: measured 12 bytes per parameter in a checkpoint (fp32 master 4 + two fp32 moments 8 + fp16 grads) ≈ 6× the fp16 weight file, 3× an fp32 one. So a 106M model ships 1,274,563,198 B. That number, not the weight size, decides the storage cycle.
  • 15.4 GB per card leaves real headroom at this scale; measure it before accepting a memory-saving technique that costs throughput (03 §3).
  • A 1B-token run here costs ~22.6 h of stepping against 24.96 h of remaining quota — a 1.48 h margin after D-019's arithmetic. Throughput decisions are quota decisions, and the margin is thin enough that a single lost session (4.55 h) spends the week's slack entirely.

The habit

Every external system gets a dated table like the above: the number, how it was measured, when, and what would invalidate it. When a plan breaks, the cause is rarely your model — it's a platform fact that was true in a doc and false on the account.