|
Download docs/00-platform-notes.md from Cion-lab/ounce100m-code: direct link, hf CLI and curl.
- Browser
- Download file 23.7 kB
-
https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/docs/00-platform-notes.md
- Command line
-
hf download hf://Cion-lab/ounce100m-code/docs/00-platform-notes.md
-
curl -L -o 00-platform-notes.md https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/docs/00-platform-notes.md
23.7 kB
| # 00 β Platform notes (Phase 0) | |
| Observed mechanics of the two execution environments this project spans: this workspace and the | |
| Kaggle image. Everything here was **measured on 2026-09-19**, not recalled. Where a claim contradicts | |
| prior documentation (including the Kaggle skill's own notes), the contradiction is called out, because | |
| a future session will otherwise trust the older statement. | |
| Raw probe JSON lives in the run logs of the kernels listed in `memory/ASSETS.md`. | |
| --- | |
| ## 1. The two environments are not the same machine | |
| | | this workspace | Kaggle CPU session | Kaggle GPU session (2xT4) | | |
| |---|---|---|---| | |
| | OS | Windows 11, git-bash | Linux 6.12.90+, glibc 2.35 | Linux 6.12.90+, glibc 2.35 | | |
| | Python | 3.11 (and 3.14) | **3.12.13** | **3.12.13** | | |
| | Docker image | n/a | `gcr.io/kaggle-images/python@sha256:dafd4ce5β¦c40b9` | `gcr.io/kaggle-gpu-images/python@sha256:37c64f7dβ¦7d461` | | |
| | torch | β | **2.10.0+cpu** | **2.10.0+cu128** | | |
| | internet | β | yes (needs `enableInternet: true`) | yes (needs `enableInternet: true`) | | |
| | GPU | none | `device_count()==0` | **2x Tesla T4, cc 7.5, 40 SM, 14.56 GiB usable each** | | |
| **Consequence:** CPU and GPU sessions are *different images*, not one image with a GPU attached. The | |
| CPU image's `torch` is a `+cpu` build, so a script that imports `torch.cuda` will work on the GPU | |
| shape and silently do nothing useful on the CPU shape. Pin these digests in any job that must be | |
| reproducible; a "latest" Kaggle image can move under the run. | |
| `huggingface_hub` is **1.32.0 locally** and reported as `huggingface-hub 1.11.0` inside the Kaggle | |
| image. Local code and in-job code are therefore on different minor versions β an API that exists | |
| locally may not exist in the job (`create_repo(tags=β¦)` is one such example; it raised | |
| `TypeError` locally). Check the target environment's version before using a new API. | |
| ## 2. Job mechanics that the main run will depend on | |
| - **Submission is one call.** `save_notebook` with `kernelExecutionType: "SaveAndRunAll"` creates | |
| *and* runs. `kernelType: "script"` and `machineShape` must both be present or the call fails with a | |
| message-less error. `machineShape: "GPU"` + `enableGpu: true` is accepted and yields 2xT4 β no | |
| T4-specific shape string is needed, and none should be invented. | |
| - **`newTitle` rewrites the slug.** Kept `newTitle` equal to the slug suffix in every probe so the | |
| returned `ref` matched what was requested. Still: **poll the returned slug, never the requested one.** | |
| - **Latency.** A ~30 s CPU script went submit β `COMPLETE` in ~85 s; a ~58 s GPU script in ~145 s. | |
| So budget **~60β90 s of queue+boot per job**, and don't poll sooner than ~60 s. Both figures are | |
| single samples from an uncongested account at ~15:20 UTC. | |
| - **Where output lands.** `list_notebook_session_output` returns `files[]` (signed URLs, expiry | |
| unknown β download promptly) and `log`, a JSON array of `{stream_name, time, data}` where `time` is | |
| **seconds since session start**, one array element per line. The log carries the whole stdout, so a | |
| probe that prints its result as JSON needs no file download. **Only files written under | |
| `/kaggle/working` are returned** β the probes deliberately wrote nothing to `/tmp`, because `/tmp` | |
| looks enormous (Β§4) and is therefore the easiest place to silently lose an artifact. | |
| - **Kernel logs are not secret-safe.** Anything printed is retained in the kernel revision. No token | |
| may be printed, and by extension no token may be *in the kernel source*, since source and logs are | |
| both Kaggle-held. | |
| - **`CUDA_VISIBLE_DEVICES` is unset in both shapes** β it is absent from the environment entirely, not | |
| set to the string `"None"`. The Kaggle skill documents `CUDA_VISIBLE_DEVICES=None` for CPU runs; | |
| that is wrong for this image (E-002 in `memory/ERRORS.md`), and code that branches on | |
| `== "None"` will take the GPU branch on a CPU box. Gate on `torch.cuda.is_available()` or | |
| `torch.cuda.device_count()` instead. | |
| ## 3. Quota accounting | |
| - Readout is **seconds**: `total_time_allowed: "108000s"` = 30 h, `quota_refresh_time: | |
| 2026-09-26T00:00:00Z`. Weekly reset is **Saturday 00:00 UTC**. | |
| - CPU sessions cost **zero** GPU quota, confirmed: two CPU probes ran and `time_used` stayed `0s`. | |
| - A 2xT4 session charged **66.411 s** for **58.4 s** of script time β **~1x wall-clock, not 2x**. | |
| Full reasoning, the caveat, and the second sample pending in `memory/QUOTA.md`. This is the single | |
| most schedule-relevant number measured in Phase 0: at 1x, 30 h/week is 30 wall-clock hours with | |
| both cards, not the 15 that Β§2's "shared across both cards" phrasing suggests. | |
| ## 4. Disk and memory β the assumption in Β§2 is only half right | |
| | Mount | Total | Free | Notes | | |
| |---|---|---|---| | |
| | `/kaggle/working`, `/kaggle/input` | **19.5 GB** | 19.5 GB | the small Kaggle disk; **only this is retrieved as output** | | |
| | `/dev/loop1` β `/kaggle/src` | 20 GB | 20 GB | separate loop device | | |
| | `/`, `/tmp`, `/usr`, `/root/.cache` (overlay) | 7.9 TB | **1.1 TB** | **shared host filesystem, 88 % used by other tenants** | | |
| | `/dev/shm` | 14 GB | 14 GB | useful for DataLoader worker traffic | | |
| | RAM | 32 GB (`cgroup memory.max` = 30 GiB) | 31.2 GB available | `SwapTotal: 0` | | |
| | CPU | 4 vCPU (`cpu.max` = `400000 100000`), Intel Xeon @ 2.20 GHz | | | | |
| **Read this carefully before planning the checkpoint cycle.** Β§2 says "Kaggle disk is small", and for | |
| `/kaggle/working` that is exactly true: 19.5 GB. But the root overlay reports ~1 TB free, and a naive | |
| optimisation would write caches or staged shards there. Two reasons not to: | |
| 1. The overlay is a **shared host volume already at 88 %**, so "1 TB free" is other tenants' headroom | |
| and can disappear mid-run. Β§3.13's "never let it fill the disk" must be enforced against a number | |
| we do not control. | |
| 2. Nothing outside `/kaggle/working` is retrievable when the session ends, so it is not storage in | |
| any sense that matters. | |
| Default `HF_HOME`/datasets cache resolves under `/root/.cache`, i.e. the overlay, not the 19.5 GB | |
| volume. So **a training job's apparent free space can be large while `/kaggle/working` fills up** β | |
| the failure mode Β§3.13 warns about is invisible to a naive `disk_usage("/")` check. **Every job must | |
| measure the specific filesystem it writes to, and log free space for that path.** On the CPU shape the | |
| effective budget is ~19.5 GB; treat that as the design number for shards resident at once. | |
| No swap: a tokenizer build or a large in-memory dedup on a 30 GiB cgroup must be sized, not streamed | |
| optimistically. | |
| ## 5. Network behaviour, including one thing that does not resolve | |
| Working, measured from inside a GPU and a CPU session: | |
| `huggingface.co/api/models` β 200 in 149 ms; `pypi.org/simple/` β 200 in 73 ms; | |
| `datasets/β¦/resolve/main/README.md` β 200 via `api/resolve-cache/β¦`; | |
| `datasets.load_dataset("wikimedia/wikipedia", "20231101.en", streaming=True)` β first row in **7.4 s**, | |
| unauthenticated. `pip install` from PyPI works: `xxhash` in **5.8 s**, import verified. | |
| **Does not resolve:** `cdn-lfs.huggingface.co`, `cdn-lfs-us-1.huggingface.co`, `xethub.huggingface.co` | |
| β `gaierror`. That looked alarming, so probe D tested it directly by pulling 300 MB from two large | |
| public parquet files. | |
| **Sustained Hub download throughput from inside a session: 72.8 MB/s** (`wikimedia/wikipedia`, 745 MB | |
| file) **and 88.9 MB/s** (`HuggingFaceFW/fineweb-edu`, 2334 MB file), with 0.8β0.94 s time-to-first-byte. | |
| So the non-resolving LFS hostnames do not block bulk transfer β `resolve/` redirects land on hosts that | |
| do resolve. **Ingest of the mix is not network-bound:** at a conservative 70 MB/s, one billion tokens | |
| (β3.6 GB of raw text at ~3.6 chars/token, or ~4β6 GB as parquet) downloads in single-digit minutes. | |
| Phase 2 should plan shard staging around CPU and disk limits, not bandwidth. | |
| Also measured in the same probe: `datasets` streaming at **139.5 rows/s β 3.0 MB of text per second** | |
| single-process, unauthenticated (83 rows/s / 1.8 MB/s cold, ~3 MB/s warm). Unauthenticated Hub access | |
| works but the API warns about rate limits (`Please set a HF_TOKEN`); at multi-hundred-shard scale that | |
| is a plausible 429 source, which is an argument for the token-in-job path (D-002) for the *data plane* | |
| even if the *checkpoint plane* stays local. | |
| Enumerating repo files without a client library: `GET https://huggingface.co/api/datasets/<id>/tree/main?recursive=true` | |
| returns `path`/`size`/`type` and works with plain `urllib`. It appeared to cap at 1000 entries | |
| (`fineweb-edu` returned exactly 1000), so **assume pagination is required** for a large repo. | |
| ### CPU-side ingest throughput, and what it rules out | |
| Measured on the Kaggle CPU shape (4 vCPU Xeon @ 2.2 GHz) with the gpt2 tokenizer over 6 MB of real | |
| Wikipedia English, `tokenizers 0.22.2` `Tokenizer.encode_batch`: | |
| | cores | tokens/s | hours to tokenize 1B tokens | | |
| |---|---|---| | |
| | 1 | 245,660 | 1.13 | | |
| | 2 | 492,191 | 0.56 | | |
| | 4 | **660,943** | **0.42** | | |
| `chars_per_token` 4.53 on this corpus. Scaling is near-linear to 2 cores and then sub-linear | |
| (2.7x at 4 cores), which is what a GIL-bound parent feeding rayon workers looks like β so a Phase 2 | |
| pipeline that shards by **process**, not by thread, should recover the rest. `datasets` streaming ran | |
| at 151.6 rows/s β **3.26 MB of text per second** single-process, unauthenticated. | |
| **The conclusion is a negative result, which is the useful kind:** tokenizing 1B tokens costs ~0.4 h of | |
| the 4-core budget, and at 35β89 MB/s the raw bytes arrive faster than `datasets` can stream them. | |
| **Phase 2 is therefore not CPU-bound and not bandwidth-bound β it is bound by dedup memory and by the | |
| 30 GiB cgroup with no swap.** Budget the design accordingly: the dedup/audit stage must be written to | |
| stream and shard against fixed RAM, not to hold a billion sketches, and the only real wall-clock cost | |
| in building the mix is the exact-duplicate/n-gram overlap pass. Note also that download throughput | |
| varied by 2.5x between identical runs (36.9 β 72.8 β 63.2 MB/s), so it should be measured inside the | |
| actual ingest job rather than treated as a constant. | |
| ## 6. Accelerator capability envelope (Turing, cc 7.5) | |
| Established inside a GPU session, not assumed: | |
| - **`torch.compile` works** (`warm + 20 steps` of an MLP in 6.0 s, incl. first-call compile). | |
| - **Flash-attention is not available.** `SDPBackend.FLASH_ATTENTION` raises | |
| `No available kernel. Aborting execution.`, preceded by | |
| `Flash attention only supports gpu architectures in the range [sm80, sm121]. Attempting to run on | |
| a sm 7.5 gpu.` β so this is architectural, not a missing dependency, and installing `flash-attn` | |
| cannot fix it. **Β§3's "never touch Hopper-only kernels" now has a concrete instance.** | |
| - **`EFFICIENT_ATTENTION` and `MATH` SDPA backends both run.** Memory-efficient attention is the | |
| viable fast path; note torch warns "Memory efficient attention has been runtime disabled" *when | |
| flash is force-requested*, so backend selection must be explicit. | |
| - `torch.cuda.get_arch_list()` includes `sm_75` (also 70, 80, 86, 90, 100, 120) β this torch build | |
| does have kernels for the card. | |
| - CUDA 12.8, cuDNN 9.10.2, **NCCL 2.27.5** present; `triton 3.6.0` present. | |
| - Absent from the image: `flash-attn`, `xformers`, `trl`, `deepspeed`, `bitsandbytes`, `liger-kernel`, | |
| `evaluate`. Present: `transformers 5.0.0`, `datasets 5.0.0`, `accelerate 1.13.0`, | |
| `tokenizers 0.22.2`, `safetensors 0.7.0`, `numpy 2.0.2`, `pyarrow 24.0.0`, `peft 0.19.1`. | |
| - bf16 is a hardware question, not a torch one: `torch.bfloat16` *tensors* exist, and bf16 on cc 7.5 | |
| is at best emulated. Probe C records whether an autocast-bf16 forward even completes and what it | |
| costs, so the Phase 1 precision decision is evidence-backed rather than a recitation of "no bf16". | |
| - `nvidia-smi` has no `multiprocessor_count` query field (rc=2); use `torch.cuda.β¦ | |
| get_device_properties(i).multi_processor_count` (40/SM) instead. | |
| ## 7. fp16 training trap found by accident | |
| Probe B ran a small Llama with **weights cast to fp16** (`model.cuda().to(torch.float16)`) under | |
| autocast with a plain AdamW and **no GradScaler**: every loss came back `NaN` in 65 steps. That is the | |
| classic fp16 underflow/overflow signature, and it is exactly the kind of bug that would otherwise be | |
| discovered 10 hours into the main run. Probe C re-runs the throughput measurements with the correct | |
| recipe β **fp32 master weights + `torch.autocast(fp16)` + `torch.amp.GradScaler("cuda")` + grad | |
| clipping** β and asserts finiteness, because on T4 there is no bf16 fallback: fp16 with loss scaling | |
| is the only mixed-precision option, so getting it right is not optional. | |
| ## 8. What Phase 0 has *not* yet proven | |
| Be honest about the boundary of this gate: | |
| - **Hub writes from inside a Kaggle job** β never attempted. No credential path chosen (D-002 open). | |
| - **NCCL across these two T4s** β still unproven; probe B's attempt was defeated by its own parser, | |
| and one failed `all_reduce` is not evidence either way. Probe C is fixing that now. | |
| - **Checkpoint upload β independent verification β local delete β free-space-confirmed (Β§3.13)** β | |
| untouched; that is a Phase 3 gate, and it needs the real 100M-model sizing arithmetic first. | |
| - **Cold resume on a fresh instance with an empty disk** β Phase 3. | |
| - **Real session caps**: `sessionTimeoutSeconds` was accepted and no session ran long enough to hit | |
| any limit. The actual wall-clock ceiling, idle-timeout behaviour, and preemption policy are | |
| **unknown**, and the multi-week resume plan in Phase 4 depends on them. Probe E | |
| (`ounce100m-p0e-session-cap`) is a ~12.5 h CPU heartbeat running now specifically to find this; | |
| if Kaggle kills it, the last heartbeat is the answer. **Update while writing: the session was still | |
| `RUNNING` at 108 minutes** (submitted 15:38Z, checked 17:26Z), so no CPU cap exists below ~1.8 h. That | |
| already rules out the worst case for Phase 4 β a session that dies inside an hour β and the build is | |
| sized so one source (~5β40 min) is the unit of lost work either way. | |
| ## 9. GPU probe C results β the numbers the schedule is built on | |
| Kernel `dodosoomro/ounce100m-p0c-fp16-ddp` v2, GPU image | |
| `gcr.io/kaggle-gpu-images/python@sha256:37c64f7dβ¦7d461`, 2x Tesla T4, charged 157.704 s of quota. | |
| Probe B's two void measurements are now replaced. Model shape used β deliberately near the likely | |
| main-run shape so the arithmetic is about the real run and not a toy: | |
| `hidden 768, layers 12, heads 12, GQA kv_heads 3, FFN 2048, vocab 50257, tied embeddings, seq 1024` | |
| β **112,934,400 parameters**, of which 38,597,376 are the (tied) embedding. | |
| | measurement | result | | |
| |---|---| | |
| | fp32 master weights + `autocast(fp16)` + `GradScaler` + clip | **loss 9.988, finite** β the recipe works; probe B's NaN was the recipe, not the hardware | | |
| | bf16 `autocast` on cc 7.5 | **runs** (loss 10.966, first fwd 0.36 s) β i.e. *emulated*, not tensor-cored | | |
| | tokens/s, 1 GPU, sdpa, bs4Γ1024 | 6,478 | | |
| | tokens/s, 1 GPU, **eager**, bs4Γ1024 | **7,239** | | |
| | tokens/s per rank, DDP 2 ranks, sdpa, bs4Γ1024 | 5,531 β **11,062 aggregate** | | |
| | DDP scaling efficiency | 11,062 / 6,478 = **1.71x** (85 % of ideal) | | |
| | NCCL across these two T4s | **works** β `rc 0`, `backend nccl`, `world 2`, both ranks finite, 30 steps | | |
| | peak CUDA memory, bs4Γ1024, 1 GPU | **11.51 GB** (sdpa) / 13.79 GB (eager) of 14.56 GB usable | | |
| | peak CUDA memory, bs4Γ1024, DDP | 8.68 GB per rank | | |
| | micro-batch ceiling at this shape | **bs 8 OOMs** (`Tried to allocate 1.54 GiB⦠520 MiB free`) | | |
| ### What each line forces | |
| - **`eager` beat `sdpa` here (7,239 vs 6,478 tok/s).** At `seq_len 1024` the memory-efficient kernel's | |
| overhead is not repaid, and flash-attention is unavailable on sm 7.5 anyway. Do not assume "flash is | |
| the fast one" β Phase 3 must re-measure at the frozen sequence length, because the crossover is | |
| length-dependent. | |
| - **Memory, not compute, is the binding constraint on batch size.** bs4 fits in 11.5 GB of 14.56 GB; | |
| bs8 does not fit at all. So global batch must be built with **gradient accumulation**, not larger | |
| micro-batches β which also means the usual "increase batch until full" advice is unavailable here. | |
| The fp32 AdamW state for a 100M model (β1.2 GB) plus fp32 master weights is a fixed floor that | |
| does not shrink with batch size. | |
| - **85 % DDP efficiency is good enough to plan on**, and is the number the ETA below uses. It was | |
| measured on 30 steps with NCCL over PCIe between two T4s; treat it as optimistic-typical, not a | |
| guaranteed constant. | |
| - **bf16 running is a trap, not a permission.** Β§2 and Β§8 both rule it out. On Turing it is emulated, | |
| so a config that silently selects bf16 will "work" and be slow. Precision must be asserted to be | |
| fp16 in the frozen config, not left to a library default. | |
| ### First ETA, from measurement rather than hope | |
| `1e9 tokens Γ· 11,062 tok/s = 90,400 s β 25.1 wall-clock hours` at the measured aggregate rate. | |
| Against a 30 h weekly allowance accruing at 1x, **the main run fits inside one week of quota on this | |
| shape β but with only ~5 h of headroom**, which is not enough to survive both checkpoint upload time | |
| and the multi-session restarts that Β§4/Phase 4 assume. Three honest caveats on that number: | |
| 1. It is a 30-step measurement on one shape that is **over the parameter budget** (112.9 M > 110 M), | |
| so the frozen config will differ and throughput with it. | |
| 2. It excludes checkpoint write/upload/verify cycles, `DataLoader` stalls, and any | |
| tokenisation-time-vs-disk tradeoff in the input pipeline. | |
| 3. It assumes sessions can run long enough to be efficient; the cap is unknown until probe E reports. | |
| Planning stance for `docs/01-plan.md`: treat **~9β11k tok/s aggregate** as the working band, size the | |
| token target from the *low* end (~28β31 h for 1 B tokens β exceeds a single week, so plan for a | |
| two-week run), and revisit once Gate 3 measures the frozen configuration. Do not let the 25 h figure | |
| become the plan's basis, because it leaves no room for the interruption budget that Β§3.1 explicitly | |
| expects to be spent. | |
| ## 10. Credentials and the Kaggle-side secret store | |
| `dodosoomro/ounce100m-p1-credential-path` (CPU, read-only) established: | |
| - Internet-enabled Kaggle sessions **arrive pre-authenticated against the Kaggle API**. The `kaggle` | |
| CLI **2.0.2 is preinstalled**, and `kaggle config view` reports `username: dodosoomro`, | |
| `auth_method: ACCESS_TOKEN`, config from `/root/.config/kaggle` β with no `kaggle.json` written by | |
| us. `kaggle kernels list --mine` returns real rows, so the principal is genuinely usable, not merely | |
| present. Injected env vars: `KAGGLE_API_V1_TOKEN` (32 ch), `KAGGLE_DATA_PROXY_TOKEN` (473), | |
| `KAGGLE_USER_SECRETS_TOKEN` (245). | |
| - **No Hugging Face credential exists in the container**: `HF_TOKEN` and `HUGGING_FACE_HUB_TOKEN` are | |
| both absent, `HF_HOME` unset, `~/.cache/huggingface` does not exist. An anonymous Hub *write* is | |
| correctly rejected β `POST /api/models/β¦/commit/main` β **401 Unauthorized** β while anonymous reads | |
| work. So jobs can pull data freely and cannot push anything until given a token. | |
| - `/kaggle/input` is empty and `/kaggle/input/.secrets/` **does not exist**, so Kaggle's UI-configured | |
| "user input / secrets" facility is unavailable: enabling it is a web-UI action, and there are no | |
| locally usable Kaggle credentials to do it with (`~/.kaggle` absent on this machine; Kaggle auth lives | |
| in the MCP server and inside the container). That rules out option 1 of D-002 for a concrete reason. | |
| - **Consequence: the in-job authenticated Kaggle API is the usable secret store.** A job can create a | |
| **private** Kaggle dataset holding the HF token; later jobs mount it via `datasetDataSources` and read | |
| it from `/kaggle/input/β¦`, so the token appears in exactly one private kernel revision instead of in | |
| every job's source. Β§2 explicitly sanctions "a secret store", so this is the intended mechanism. | |
| **Ordering used, because the failure mode is unrecoverable.** `kaggle datasets create` derives | |
| visibility from a metadata field whose name has moved between CLI versions, and an unrecognised key is | |
| **silently ignored** β which would create a *public* dataset containing the token. So the bootstrap | |
| worker (a) creates the dataset with an inert placeholder only, (b) proves privacy by confirming an | |
| **anonymous** download fails, and (c) uploads the credential as a later version only if (b) passes, | |
| then re-proves. If anonymous read ever succeeds, the secret is never written and the job says so loudly. | |
| **Residual risk, stated:** kernel `p1-secret-bootstrap` v1 necessarily contains the token inline, and | |
| Kaggle retains private kernel revisions β overwrite-after-use does not erase history, and there is no | |
| delete-kernel operation in CLI 2.0.2 or in the MCP toolset. This is a strictly smaller exposure than | |
| putting the token in every training job and in anything that mirrors kernel source. | |
| ## 11. Pre-existing Kaggle kernels, not ours | |
| `kaggle kernels list --mine` also returned five kernels predating this project | |
| (`quick-python-cpu-smoke-test`, `monte-carlo-cpu-smoke-test`, `cpu-test`, `exam-80-monte-carlo`, | |
| `workbuddy-cpu-probe`; last run 2026-09-17 β 2026-09-19). They are **not touched** (Β§3.10). This also | |
| explains E-002: `search_notebooks` returning `{}` for this account was the search tool failing, not the | |
| account being empty β so an empty search must never be read as "no prior work". | |
| ## 12. Measured notebook session cap: 12.0 h survived, the instance reported 45,000 s | |
| `dodosoomro/ounce100m-p0e-session-cap` β a CPU kernel that printed one heartbeat every two minutes and | |
| nothing else, costing zero GPU quota, left running since 2026-09-19T15:38Z purely as a measurement. | |
| ``` | |
| CAP_START {"epoch": 1789832305.41, "host": "096ef365b02e", "pid": 7, "limit": 45000} | |
| HB 1 elapsed=0 free_GB=19.5 mem_avail_MB=31204 load=0.41 net=204 | |
| HB 361 elapsed=43204 free_GB=19.5 mem_avail_MB=31288 load=0.00 net=204 | |
| ``` | |
| **361 heartbeats, 43,204 s = 12.0 h of continuous running, then the platform took it** | |
| (`CANCEL_ACKNOWLEDGED`, not a crash β the last heartbeat is clean and disk/RAM are unchanged at 19.5 GB | |
| free and 31.3 GB available). The instance itself reported `limit: 45000` s, i.e. **12.5 h**, so the kill | |
| arrived at ~96 % of its own stated budget. | |
| What this does and does not license: | |
| - It is a **CPU** measurement. The GPU accelerator quota is tracked separately, and nothing here proves a | |
| GPU session is allowed to run 12.5 h. So Phase 4 plans `SESSION_GPU_HOURS=6.9` for session 1 β well | |
| inside every plausible cap β and can extend later sessions only on evidence from earlier ones. | |
| - It does settle the question that the plan had been hedging since Β§2 of `04-run-log.md` ("the session cap | |
| is unmeasured"): an interruption at 6 h is not the platform's limit, so a long session is not doomed, and | |
| a `--stop-after-steps` schedule sized for 6-7 h is conservative rather than necessary. | |
| - It confirms the working volume is stable for the whole duration β `free_GB` never moved from 19.5, so an | |
| 11 GB drift over a session is not a thing to plan for, and the disk floor in the launcher is about the | |
| mix and checkpoints, not about slow leaks. | |