# 07 — Platform notes Numbers measured on one account, one week, one hardware generation. Treat them as **the shape of the fact to record**, not as facts: re-measure anything you're about to rely on. Staleness here is the default state. Stack: Kaggle for compute (2×Tesla T4), Hugging Face Hub for storage and publication. ## Kaggle | Fact | Measured | |---|---| | GPU billing model | **1× container wall-clock**, not accelerator-seconds. A job that idles, installs, or waits in a queue is billed. Start the clock deliberately. | | Weekly allowance | 108,000 s = 30 GPU-h. Reset **Saturday 00:00Z** (`quota_refresh_time` 2026-09-26T00:00:00Z, which is a Saturday — looked up, not assumed). Read it live; don't infer the weekday. | | Small-model test cap (self-imposed) | 6 GPU-h lifetime for all preflight; closed at 4.63 h | | Session ceiling | ~25,200 s (7 h) for a GPU session; treat the *practical* cap as lower and end on a checkpoint with margin | | Kernel logs | **Not available until the session completes.** Plan monitoring around remote artefacts instead | | Env vars to a job | Cannot be passed through `create_notebook_session`. Workaround: a thin **pinned wrapper kernel** that `runpy.run_path`s the real script fetched at a commit sha, and carries only **non-secret** values (chosen session size, planning rate, quota figure). Secrets never go in kernel source — see `01` §4 and D-006. | | Quota read | `get_accelerator_quota`; `time_reserved` was `0s` in every read taken between jobs, so it is **not** a live-job indicator — use the session-status call for that | | Throughput, 106M model, seq 1024, 2×T4, no checkpointing | **12,221–12,312 tok/s sustained** = 21.5 s/step at 262,144 tok/step. With gradient checkpointing: 9,358–9,696 tok/s | | Memory, same config | **cross-rank max** reserved 12.84 GiB / allocated 12.70 GiB, identical at step 10 and step 120 of the same process (i.e. stable *over time*; per-rank equality was never established — the number is an `all_reduce(MAX)`) | MCP/API mechanics that bit: - `save_notebook` requires `language` and `kernelType` (and `machineShape` for some CPU shapes). - Passing `newTitle` **rewrites the kernel slug**. Address existing kernels by their current slug and re-read it after any title change; every stale reference breaks silently. - `dockerImagePinningType` accepts `latest` and will not pin a revision usefully, so reproducibility comes from *your* commit shas plus explicit `pip` version pins — not from the image. - pip `--user` installs land under `/root/.local/...`, which is **not** on the parent process's `sys.path` when the directory is created mid-process. Install, then start a fresh interpreter, or add the path explicitly. - Big tool outputs persist to a file rather than returning inline; decode the `log` field with stdlib JSON. - A job that exits within ~2 minutes almost always *refused to start* on its own assertions. That is a good outcome — make refusals print a named verdict line you can grep for. ## Hugging Face Hub | Fact | Measured | |---|---| | Commit-rate ceiling | ~139 sequential commits in a burst starts failing. **Batch** multi-file pushes into one commit. | | Checkpoint cycle, 1.27 GB | **7–10 s** on the live run, measured as the gap between the `checkpoint ckpt/checkpoint-N` commit and the `latest -> N` commit (verification happens in between, so this bounds push + ranged byte-verify + point). The earlier whole-object download design took **797 s** off a cold CDN. Verification design sets your cadence. | | LFS files | `lfs["oid"]` **is** the content sha256 — and it's a dict entry, not an object attribute | | Small (non-LFS) files | expose only a sha1 *blob id*, **not** comparable to a sha256. Don't write an equality check that mixes them. | | Range requests | return 206 with `content-range: bytes a-b/size` through the LFS redirect. Compare the header to the window you asked for: a server that ignores `Range` answers 200 with the whole object, which is how a "verified" read-back can prove nothing. | | Small-file fetches | `resolve/main/` answers **307 → `/api/resolve-cache//`**. A client that doesn't follow redirects reads the redirect stub as the file. The target embeds the head sha — a free consistency check. | | Missing repo | 404 / `RepositoryNotFoundError`. A **timeout is not** "missing", and neither is 403. | | 403 is ambiguous | a *private* repo and a *nonexistent* repo both answer 403/404 to the wrong caller. Proving a secret store is private takes **two different calls** (`01` §4). | | Anonymous reads | public repos are readable with no token, including a checkpoint's `trainer_state.json` — remote loss monitoring needs no credentials. | ## Evaluation harness (lm-eval) | Fact | Detail | |---|---| | Metric keys | shaped `","`: `acc,none`, `exact_match,strict-match`. Not settleable from config files; assert the key exists in the *result dict* at smoke time | | MMLU | grouped across subjects; the per-group value is a list, so `len()` of the result isn't a row count. Some tasks have **two** metrics (em + strict-match); pick one and record which | | Splits | several "zero-shot academic" tasks score on `validation`; some datasets have no `test` split at all | | `lm_eval --tasks list` | **not** a listing command in 0.4.13 — `ValueError: Tasks not found: list`. Use `TaskManager().task_index` | | Optional deps | `--no-deps` installs import fine and then die at run time on `sacrebleu`. Install with dependencies | | Denominators | `n_input` / `n_logged` live in per-task output JSONs; harvest by walking the output tree, and treat a missing denominator as a failed cell, not a zero | | **Not measured** | total eval wall-clock for the eight tasks on 2×T4. Unknown as of the freeze, which is exactly why `D-019` puts training ahead of it | ## Hardware reality check - **fp16 autocast + fp32 master weights + GradScaler** works on Turing. bf16 is unavailable (no Ampere+), so any recipe assuming bf16 or Hopper-only kernels (flash-attn/FA2) is out. Eager attention is the price. - Optimizer state dominates the artefact: measured **12 bytes per parameter** in a checkpoint (fp32 master 4 + two fp32 moments 8 + fp16 grads) ≈ **6× the fp16 weight file**, 3× an fp32 one. So a 106M model ships 1,274,563,198 B. That number, not the weight size, decides the storage cycle. - 15.4 GB per card leaves real headroom at this scale; measure it before accepting a memory-saving technique that costs throughput (`03` §3). - A 1B-token run here costs ~22.6 h of stepping against 24.96 h of remaining quota — a **1.48 h margin** after D-019's arithmetic. Throughput decisions are quota decisions, and the margin is thin enough that a single lost session (4.55 h) spends the week's slack entirely. ## The habit Every external system gets a dated table like the above: the number, how it was measured, when, and what would invalidate it. When a plan breaks, the cause is rarely your model — it's a platform fact that was true in a doc and false on the account.