ounce100m-code / guide /07-platform-notes.md
Cion-lab's picture
publish guide/: the operating manual distilled from 54 failure records, for future agent sessions
6302710 verified
|
Raw History Blame Contribute Delete
7.13 kB
# 07 β€” Platform notes
Numbers measured on one account, one week, one hardware generation. Treat them as **the shape of the fact to
record**, not as facts: re-measure anything you're about to rely on. Staleness here is the default state.
Stack: Kaggle for compute (2Γ—Tesla T4), Hugging Face Hub for storage and publication.
## Kaggle
| Fact | Measured |
|---|---|
| GPU billing model | **1Γ— container wall-clock**, not accelerator-seconds. A job that idles, installs, or waits in a queue is billed. Start the clock deliberately. |
| Weekly allowance | 108,000 s = 30 GPU-h. Reset **Saturday 00:00Z** (`quota_refresh_time` 2026-09-26T00:00:00Z, which is a Saturday β€” looked up, not assumed). Read it live; don't infer the weekday. |
| Small-model test cap (self-imposed) | 6 GPU-h lifetime for all preflight; closed at 4.63 h |
| Session ceiling | ~25,200 s (7 h) for a GPU session; treat the *practical* cap as lower and end on a checkpoint with margin |
| Kernel logs | **Not available until the session completes.** Plan monitoring around remote artefacts instead |
| Env vars to a job | Cannot be passed through `create_notebook_session`. Workaround: a thin **pinned wrapper kernel** that `runpy.run_path`s the real script fetched at a commit sha, and carries only **non-secret** values (chosen session size, planning rate, quota figure). Secrets never go in kernel source β€” see `01` Β§4 and D-006. |
| Quota read | `get_accelerator_quota`; `time_reserved` was `0s` in every read taken between jobs, so it is **not** a live-job indicator β€” use the session-status call for that |
| Throughput, 106M model, seq 1024, 2Γ—T4, no checkpointing | **12,221–12,312 tok/s sustained** = 21.5 s/step at 262,144 tok/step. With gradient checkpointing: 9,358–9,696 tok/s |
| Memory, same config | **cross-rank max** reserved 12.84 GiB / allocated 12.70 GiB, identical at step 10 and step 120 of the same process (i.e. stable *over time*; per-rank equality was never established β€” the number is an `all_reduce(MAX)`) |
MCP/API mechanics that bit:
- `save_notebook` requires `language` and `kernelType` (and `machineShape` for some CPU shapes).
- Passing `newTitle` **rewrites the kernel slug**. Address existing kernels by their current slug and re-read it
after any title change; every stale reference breaks silently.
- `dockerImagePinningType` accepts `latest` and will not pin a revision usefully, so reproducibility comes from
*your* commit shas plus explicit `pip` version pins β€” not from the image.
- pip `--user` installs land under `/root/.local/...`, which is **not** on the parent process's `sys.path` when
the directory is created mid-process. Install, then start a fresh interpreter, or add the path explicitly.
- Big tool outputs persist to a file rather than returning inline; decode the `log` field with stdlib JSON.
- A job that exits within ~2 minutes almost always *refused to start* on its own assertions. That is a good
outcome β€” make refusals print a named verdict line you can grep for.
## Hugging Face Hub
| Fact | Measured |
|---|---|
| Commit-rate ceiling | ~139 sequential commits in a burst starts failing. **Batch** multi-file pushes into one commit. |
| Checkpoint cycle, 1.27 GB | **7–10 s** on the live run, measured as the gap between the `checkpoint ckpt/checkpoint-N` commit and the `latest -> N` commit (verification happens in between, so this bounds push + ranged byte-verify + point). The earlier whole-object download design took **797 s** off a cold CDN. Verification design sets your cadence. |
| LFS files | `lfs["oid"]` **is** the content sha256 β€” and it's a dict entry, not an object attribute |
| Small (non-LFS) files | expose only a sha1 *blob id*, **not** comparable to a sha256. Don't write an equality check that mixes them. |
| Range requests | return 206 with `content-range: bytes a-b/size` through the LFS redirect. Compare the header to the window you asked for: a server that ignores `Range` answers 200 with the whole object, which is how a "verified" read-back can prove nothing. |
| Small-file fetches | `resolve/main/<small file>` answers **307 β†’ `/api/resolve-cache/<commit sha>/<path>`**. A client that doesn't follow redirects reads the redirect stub as the file. The target embeds the head sha β€” a free consistency check. |
| Missing repo | 404 / `RepositoryNotFoundError`. A **timeout is not** "missing", and neither is 403. |
| 403 is ambiguous | a *private* repo and a *nonexistent* repo both answer 403/404 to the wrong caller. Proving a secret store is private takes **two different calls** (`01` Β§4). |
| Anonymous reads | public repos are readable with no token, including a checkpoint's `trainer_state.json` β€” remote loss monitoring needs no credentials. |
## Evaluation harness (lm-eval)
| Fact | Detail |
|---|---|
| Metric keys | shaped `"<name>,<aggregation>"`: `acc,none`, `exact_match,strict-match`. Not settleable from config files; assert the key exists in the *result dict* at smoke time |
| MMLU | grouped across subjects; the per-group value is a list, so `len()` of the result isn't a row count. Some tasks have **two** metrics (em + strict-match); pick one and record which |
| Splits | several "zero-shot academic" tasks score on `validation`; some datasets have no `test` split at all |
| `lm_eval --tasks list` | **not** a listing command in 0.4.13 β€” `ValueError: Tasks not found: list`. Use `TaskManager().task_index` |
| Optional deps | `--no-deps` installs import fine and then die at run time on `sacrebleu`. Install with dependencies |
| Denominators | `n_input` / `n_logged` live in per-task output JSONs; harvest by walking the output tree, and treat a missing denominator as a failed cell, not a zero |
| **Not measured** | total eval wall-clock for the eight tasks on 2Γ—T4. Unknown as of the freeze, which is exactly why `D-019` puts training ahead of it |
## Hardware reality check
- **fp16 autocast + fp32 master weights + GradScaler** works on Turing. bf16 is unavailable (no Ampere+), so any
recipe assuming bf16 or Hopper-only kernels (flash-attn/FA2) is out. Eager attention is the price.
- Optimizer state dominates the artefact: measured **12 bytes per parameter** in a checkpoint
(fp32 master 4 + two fp32 moments 8 + fp16 grads) β‰ˆ **6Γ— the fp16 weight file**, 3Γ— an fp32 one. So a 106M
model ships 1,274,563,198 B. That number, not the weight size, decides the storage cycle.
- 15.4 GB per card leaves real headroom at this scale; measure it before accepting a memory-saving technique
that costs throughput (`03` Β§3).
- A 1B-token run here costs ~22.6 h of stepping against 24.96 h of remaining quota β€” a **1.48 h margin** after
D-019's arithmetic. Throughput decisions are quota decisions, and the margin is thin enough that a single
lost session (4.55 h) spends the week's slack entirely.
## The habit
Every external system gets a dated table like the above: the number, how it was measured, when, and what would
invalidate it. When a plan breaks, the cause is rarely your model β€” it's a platform fact that was true in a doc
and false on the account.