|
Download guide/07-platform-notes.md from Cion-lab/ounce100m-code: direct link, hf CLI and curl.
- Browser
- Download file 7.13 kB
-
https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/guide/07-platform-notes.md
- Command line
-
hf download hf://Cion-lab/ounce100m-code/guide/07-platform-notes.md
-
curl -L -o 07-platform-notes.md https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/guide/07-platform-notes.md
7.13 kB
| # 07 β Platform notes | |
| Numbers measured on one account, one week, one hardware generation. Treat them as **the shape of the fact to | |
| record**, not as facts: re-measure anything you're about to rely on. Staleness here is the default state. | |
| Stack: Kaggle for compute (2ΓTesla T4), Hugging Face Hub for storage and publication. | |
| ## Kaggle | |
| | Fact | Measured | | |
| |---|---| | |
| | GPU billing model | **1Γ container wall-clock**, not accelerator-seconds. A job that idles, installs, or waits in a queue is billed. Start the clock deliberately. | | |
| | Weekly allowance | 108,000 s = 30 GPU-h. Reset **Saturday 00:00Z** (`quota_refresh_time` 2026-09-26T00:00:00Z, which is a Saturday β looked up, not assumed). Read it live; don't infer the weekday. | | |
| | Small-model test cap (self-imposed) | 6 GPU-h lifetime for all preflight; closed at 4.63 h | | |
| | Session ceiling | ~25,200 s (7 h) for a GPU session; treat the *practical* cap as lower and end on a checkpoint with margin | | |
| | Kernel logs | **Not available until the session completes.** Plan monitoring around remote artefacts instead | | |
| | Env vars to a job | Cannot be passed through `create_notebook_session`. Workaround: a thin **pinned wrapper kernel** that `runpy.run_path`s the real script fetched at a commit sha, and carries only **non-secret** values (chosen session size, planning rate, quota figure). Secrets never go in kernel source β see `01` Β§4 and D-006. | | |
| | Quota read | `get_accelerator_quota`; `time_reserved` was `0s` in every read taken between jobs, so it is **not** a live-job indicator β use the session-status call for that | | |
| | Throughput, 106M model, seq 1024, 2ΓT4, no checkpointing | **12,221β12,312 tok/s sustained** = 21.5 s/step at 262,144 tok/step. With gradient checkpointing: 9,358β9,696 tok/s | | |
| | Memory, same config | **cross-rank max** reserved 12.84 GiB / allocated 12.70 GiB, identical at step 10 and step 120 of the same process (i.e. stable *over time*; per-rank equality was never established β the number is an `all_reduce(MAX)`) | | |
| MCP/API mechanics that bit: | |
| - `save_notebook` requires `language` and `kernelType` (and `machineShape` for some CPU shapes). | |
| - Passing `newTitle` **rewrites the kernel slug**. Address existing kernels by their current slug and re-read it | |
| after any title change; every stale reference breaks silently. | |
| - `dockerImagePinningType` accepts `latest` and will not pin a revision usefully, so reproducibility comes from | |
| *your* commit shas plus explicit `pip` version pins β not from the image. | |
| - pip `--user` installs land under `/root/.local/...`, which is **not** on the parent process's `sys.path` when | |
| the directory is created mid-process. Install, then start a fresh interpreter, or add the path explicitly. | |
| - Big tool outputs persist to a file rather than returning inline; decode the `log` field with stdlib JSON. | |
| - A job that exits within ~2 minutes almost always *refused to start* on its own assertions. That is a good | |
| outcome β make refusals print a named verdict line you can grep for. | |
| ## Hugging Face Hub | |
| | Fact | Measured | | |
| |---|---| | |
| | Commit-rate ceiling | ~139 sequential commits in a burst starts failing. **Batch** multi-file pushes into one commit. | | |
| | Checkpoint cycle, 1.27 GB | **7β10 s** on the live run, measured as the gap between the `checkpoint ckpt/checkpoint-N` commit and the `latest -> N` commit (verification happens in between, so this bounds push + ranged byte-verify + point). The earlier whole-object download design took **797 s** off a cold CDN. Verification design sets your cadence. | | |
| | LFS files | `lfs["oid"]` **is** the content sha256 β and it's a dict entry, not an object attribute | | |
| | Small (non-LFS) files | expose only a sha1 *blob id*, **not** comparable to a sha256. Don't write an equality check that mixes them. | | |
| | Range requests | return 206 with `content-range: bytes a-b/size` through the LFS redirect. Compare the header to the window you asked for: a server that ignores `Range` answers 200 with the whole object, which is how a "verified" read-back can prove nothing. | | |
| | Small-file fetches | `resolve/main/<small file>` answers **307 β `/api/resolve-cache/<commit sha>/<path>`**. A client that doesn't follow redirects reads the redirect stub as the file. The target embeds the head sha β a free consistency check. | | |
| | Missing repo | 404 / `RepositoryNotFoundError`. A **timeout is not** "missing", and neither is 403. | | |
| | 403 is ambiguous | a *private* repo and a *nonexistent* repo both answer 403/404 to the wrong caller. Proving a secret store is private takes **two different calls** (`01` Β§4). | | |
| | Anonymous reads | public repos are readable with no token, including a checkpoint's `trainer_state.json` β remote loss monitoring needs no credentials. | | |
| ## Evaluation harness (lm-eval) | |
| | Fact | Detail | | |
| |---|---| | |
| | Metric keys | shaped `"<name>,<aggregation>"`: `acc,none`, `exact_match,strict-match`. Not settleable from config files; assert the key exists in the *result dict* at smoke time | | |
| | MMLU | grouped across subjects; the per-group value is a list, so `len()` of the result isn't a row count. Some tasks have **two** metrics (em + strict-match); pick one and record which | | |
| | Splits | several "zero-shot academic" tasks score on `validation`; some datasets have no `test` split at all | | |
| | `lm_eval --tasks list` | **not** a listing command in 0.4.13 β `ValueError: Tasks not found: list`. Use `TaskManager().task_index` | | |
| | Optional deps | `--no-deps` installs import fine and then die at run time on `sacrebleu`. Install with dependencies | | |
| | Denominators | `n_input` / `n_logged` live in per-task output JSONs; harvest by walking the output tree, and treat a missing denominator as a failed cell, not a zero | | |
| | **Not measured** | total eval wall-clock for the eight tasks on 2ΓT4. Unknown as of the freeze, which is exactly why `D-019` puts training ahead of it | | |
| ## Hardware reality check | |
| - **fp16 autocast + fp32 master weights + GradScaler** works on Turing. bf16 is unavailable (no Ampere+), so any | |
| recipe assuming bf16 or Hopper-only kernels (flash-attn/FA2) is out. Eager attention is the price. | |
| - Optimizer state dominates the artefact: measured **12 bytes per parameter** in a checkpoint | |
| (fp32 master 4 + two fp32 moments 8 + fp16 grads) β **6Γ the fp16 weight file**, 3Γ an fp32 one. So a 106M | |
| model ships 1,274,563,198 B. That number, not the weight size, decides the storage cycle. | |
| - 15.4 GB per card leaves real headroom at this scale; measure it before accepting a memory-saving technique | |
| that costs throughput (`03` Β§3). | |
| - A 1B-token run here costs ~22.6 h of stepping against 24.96 h of remaining quota β a **1.48 h margin** after | |
| D-019's arithmetic. Throughput decisions are quota decisions, and the margin is thin enough that a single | |
| lost session (4.55 h) spends the week's slack entirely. | |
| ## The habit | |
| Every external system gets a dated table like the above: the number, how it was measured, when, and what would | |
| invalidate it. When a plan breaks, the cause is rarely your model β it's a platform fact that was true in a doc | |
| and false on the account. | |