|
Download guide/07-platform-notes.md from Cion-lab/ounce100m-code: direct link, hf CLI and curl.
- Browser
- Download file 7.13 kB
-
https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/guide/07-platform-notes.md
- Command line
-
hf download hf://Cion-lab/ounce100m-code/guide/07-platform-notes.md
-
curl -L -o 07-platform-notes.md https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/guide/07-platform-notes.md
7.13 kB
07 — Platform notes
Numbers measured on one account, one week, one hardware generation. Treat them as the shape of the fact to record, not as facts: re-measure anything you're about to rely on. Staleness here is the default state.
Stack: Kaggle for compute (2×Tesla T4), Hugging Face Hub for storage and publication.
Kaggle
| Fact | Measured |
|---|---|
| GPU billing model | 1× container wall-clock, not accelerator-seconds. A job that idles, installs, or waits in a queue is billed. Start the clock deliberately. |
| Weekly allowance | 108,000 s = 30 GPU-h. Reset Saturday 00:00Z (quota_refresh_time 2026-09-26T00:00:00Z, which is a Saturday — looked up, not assumed). Read it live; don't infer the weekday. |
| Small-model test cap (self-imposed) | 6 GPU-h lifetime for all preflight; closed at 4.63 h |
| Session ceiling | ~25,200 s (7 h) for a GPU session; treat the practical cap as lower and end on a checkpoint with margin |
| Kernel logs | Not available until the session completes. Plan monitoring around remote artefacts instead |
| Env vars to a job | Cannot be passed through create_notebook_session. Workaround: a thin pinned wrapper kernel that runpy.run_paths the real script fetched at a commit sha, and carries only non-secret values (chosen session size, planning rate, quota figure). Secrets never go in kernel source — see 01 §4 and D-006. |
| Quota read | get_accelerator_quota; time_reserved was 0s in every read taken between jobs, so it is not a live-job indicator — use the session-status call for that |
| Throughput, 106M model, seq 1024, 2×T4, no checkpointing | 12,221–12,312 tok/s sustained = 21.5 s/step at 262,144 tok/step. With gradient checkpointing: 9,358–9,696 tok/s |
| Memory, same config | cross-rank max reserved 12.84 GiB / allocated 12.70 GiB, identical at step 10 and step 120 of the same process (i.e. stable over time; per-rank equality was never established — the number is an all_reduce(MAX)) |
MCP/API mechanics that bit:
save_notebookrequireslanguageandkernelType(andmachineShapefor some CPU shapes).- Passing
newTitlerewrites the kernel slug. Address existing kernels by their current slug and re-read it after any title change; every stale reference breaks silently. dockerImagePinningTypeacceptslatestand will not pin a revision usefully, so reproducibility comes from your commit shas plus explicitpipversion pins — not from the image.- pip
--userinstalls land under/root/.local/..., which is not on the parent process'ssys.pathwhen the directory is created mid-process. Install, then start a fresh interpreter, or add the path explicitly. - Big tool outputs persist to a file rather than returning inline; decode the
logfield with stdlib JSON. - A job that exits within ~2 minutes almost always refused to start on its own assertions. That is a good outcome — make refusals print a named verdict line you can grep for.
Hugging Face Hub
| Fact | Measured |
|---|---|
| Commit-rate ceiling | ~139 sequential commits in a burst starts failing. Batch multi-file pushes into one commit. |
| Checkpoint cycle, 1.27 GB | 7–10 s on the live run, measured as the gap between the checkpoint ckpt/checkpoint-N commit and the latest -> N commit (verification happens in between, so this bounds push + ranged byte-verify + point). The earlier whole-object download design took 797 s off a cold CDN. Verification design sets your cadence. |
| LFS files | lfs["oid"] is the content sha256 — and it's a dict entry, not an object attribute |
| Small (non-LFS) files | expose only a sha1 blob id, not comparable to a sha256. Don't write an equality check that mixes them. |
| Range requests | return 206 with content-range: bytes a-b/size through the LFS redirect. Compare the header to the window you asked for: a server that ignores Range answers 200 with the whole object, which is how a "verified" read-back can prove nothing. |
| Small-file fetches | resolve/main/<small file> answers 307 → /api/resolve-cache/<commit sha>/<path>. A client that doesn't follow redirects reads the redirect stub as the file. The target embeds the head sha — a free consistency check. |
| Missing repo | 404 / RepositoryNotFoundError. A timeout is not "missing", and neither is 403. |
| 403 is ambiguous | a private repo and a nonexistent repo both answer 403/404 to the wrong caller. Proving a secret store is private takes two different calls (01 §4). |
| Anonymous reads | public repos are readable with no token, including a checkpoint's trainer_state.json — remote loss monitoring needs no credentials. |
Evaluation harness (lm-eval)
| Fact | Detail |
|---|---|
| Metric keys | shaped "<name>,<aggregation>": acc,none, exact_match,strict-match. Not settleable from config files; assert the key exists in the result dict at smoke time |
| MMLU | grouped across subjects; the per-group value is a list, so len() of the result isn't a row count. Some tasks have two metrics (em + strict-match); pick one and record which |
| Splits | several "zero-shot academic" tasks score on validation; some datasets have no test split at all |
lm_eval --tasks list |
not a listing command in 0.4.13 — ValueError: Tasks not found: list. Use TaskManager().task_index |
| Optional deps | --no-deps installs import fine and then die at run time on sacrebleu. Install with dependencies |
| Denominators | n_input / n_logged live in per-task output JSONs; harvest by walking the output tree, and treat a missing denominator as a failed cell, not a zero |
| Not measured | total eval wall-clock for the eight tasks on 2×T4. Unknown as of the freeze, which is exactly why D-019 puts training ahead of it |
Hardware reality check
- fp16 autocast + fp32 master weights + GradScaler works on Turing. bf16 is unavailable (no Ampere+), so any recipe assuming bf16 or Hopper-only kernels (flash-attn/FA2) is out. Eager attention is the price.
- Optimizer state dominates the artefact: measured 12 bytes per parameter in a checkpoint (fp32 master 4 + two fp32 moments 8 + fp16 grads) ≈ 6× the fp16 weight file, 3× an fp32 one. So a 106M model ships 1,274,563,198 B. That number, not the weight size, decides the storage cycle.
- 15.4 GB per card leaves real headroom at this scale; measure it before accepting a memory-saving technique
that costs throughput (
03§3). - A 1B-token run here costs ~22.6 h of stepping against 24.96 h of remaining quota — a 1.48 h margin after D-019's arithmetic. Throughput decisions are quota decisions, and the margin is thin enough that a single lost session (4.55 h) spends the week's slack entirely.
The habit
Every external system gets a dated table like the above: the number, how it was measured, when, and what would invalidate it. When a plan breaks, the cause is rarely your model — it's a platform fact that was true in a doc and false on the account.