File size: 23,664 Bytes
f345921 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 | # 00 β Platform notes (Phase 0)
Observed mechanics of the two execution environments this project spans: this workspace and the
Kaggle image. Everything here was **measured on 2026-09-19**, not recalled. Where a claim contradicts
prior documentation (including the Kaggle skill's own notes), the contradiction is called out, because
a future session will otherwise trust the older statement.
Raw probe JSON lives in the run logs of the kernels listed in `memory/ASSETS.md`.
---
## 1. The two environments are not the same machine
| | this workspace | Kaggle CPU session | Kaggle GPU session (2xT4) |
|---|---|---|---|
| OS | Windows 11, git-bash | Linux 6.12.90+, glibc 2.35 | Linux 6.12.90+, glibc 2.35 |
| Python | 3.11 (and 3.14) | **3.12.13** | **3.12.13** |
| Docker image | n/a | `gcr.io/kaggle-images/python@sha256:dafd4ce5β¦c40b9` | `gcr.io/kaggle-gpu-images/python@sha256:37c64f7dβ¦7d461` |
| torch | β | **2.10.0+cpu** | **2.10.0+cu128** |
| internet | β | yes (needs `enableInternet: true`) | yes (needs `enableInternet: true`) |
| GPU | none | `device_count()==0` | **2x Tesla T4, cc 7.5, 40 SM, 14.56 GiB usable each** |
**Consequence:** CPU and GPU sessions are *different images*, not one image with a GPU attached. The
CPU image's `torch` is a `+cpu` build, so a script that imports `torch.cuda` will work on the GPU
shape and silently do nothing useful on the CPU shape. Pin these digests in any job that must be
reproducible; a "latest" Kaggle image can move under the run.
`huggingface_hub` is **1.32.0 locally** and reported as `huggingface-hub 1.11.0` inside the Kaggle
image. Local code and in-job code are therefore on different minor versions β an API that exists
locally may not exist in the job (`create_repo(tags=β¦)` is one such example; it raised
`TypeError` locally). Check the target environment's version before using a new API.
## 2. Job mechanics that the main run will depend on
- **Submission is one call.** `save_notebook` with `kernelExecutionType: "SaveAndRunAll"` creates
*and* runs. `kernelType: "script"` and `machineShape` must both be present or the call fails with a
message-less error. `machineShape: "GPU"` + `enableGpu: true` is accepted and yields 2xT4 β no
T4-specific shape string is needed, and none should be invented.
- **`newTitle` rewrites the slug.** Kept `newTitle` equal to the slug suffix in every probe so the
returned `ref` matched what was requested. Still: **poll the returned slug, never the requested one.**
- **Latency.** A ~30 s CPU script went submit β `COMPLETE` in ~85 s; a ~58 s GPU script in ~145 s.
So budget **~60β90 s of queue+boot per job**, and don't poll sooner than ~60 s. Both figures are
single samples from an uncongested account at ~15:20 UTC.
- **Where output lands.** `list_notebook_session_output` returns `files[]` (signed URLs, expiry
unknown β download promptly) and `log`, a JSON array of `{stream_name, time, data}` where `time` is
**seconds since session start**, one array element per line. The log carries the whole stdout, so a
probe that prints its result as JSON needs no file download. **Only files written under
`/kaggle/working` are returned** β the probes deliberately wrote nothing to `/tmp`, because `/tmp`
looks enormous (Β§4) and is therefore the easiest place to silently lose an artifact.
- **Kernel logs are not secret-safe.** Anything printed is retained in the kernel revision. No token
may be printed, and by extension no token may be *in the kernel source*, since source and logs are
both Kaggle-held.
- **`CUDA_VISIBLE_DEVICES` is unset in both shapes** β it is absent from the environment entirely, not
set to the string `"None"`. The Kaggle skill documents `CUDA_VISIBLE_DEVICES=None` for CPU runs;
that is wrong for this image (E-002 in `memory/ERRORS.md`), and code that branches on
`== "None"` will take the GPU branch on a CPU box. Gate on `torch.cuda.is_available()` or
`torch.cuda.device_count()` instead.
## 3. Quota accounting
- Readout is **seconds**: `total_time_allowed: "108000s"` = 30 h, `quota_refresh_time:
2026-09-26T00:00:00Z`. Weekly reset is **Saturday 00:00 UTC**.
- CPU sessions cost **zero** GPU quota, confirmed: two CPU probes ran and `time_used` stayed `0s`.
- A 2xT4 session charged **66.411 s** for **58.4 s** of script time β **~1x wall-clock, not 2x**.
Full reasoning, the caveat, and the second sample pending in `memory/QUOTA.md`. This is the single
most schedule-relevant number measured in Phase 0: at 1x, 30 h/week is 30 wall-clock hours with
both cards, not the 15 that Β§2's "shared across both cards" phrasing suggests.
## 4. Disk and memory β the assumption in Β§2 is only half right
| Mount | Total | Free | Notes |
|---|---|---|---|
| `/kaggle/working`, `/kaggle/input` | **19.5 GB** | 19.5 GB | the small Kaggle disk; **only this is retrieved as output** |
| `/dev/loop1` β `/kaggle/src` | 20 GB | 20 GB | separate loop device |
| `/`, `/tmp`, `/usr`, `/root/.cache` (overlay) | 7.9 TB | **1.1 TB** | **shared host filesystem, 88 % used by other tenants** |
| `/dev/shm` | 14 GB | 14 GB | useful for DataLoader worker traffic |
| RAM | 32 GB (`cgroup memory.max` = 30 GiB) | 31.2 GB available | `SwapTotal: 0` |
| CPU | 4 vCPU (`cpu.max` = `400000 100000`), Intel Xeon @ 2.20 GHz | | |
**Read this carefully before planning the checkpoint cycle.** Β§2 says "Kaggle disk is small", and for
`/kaggle/working` that is exactly true: 19.5 GB. But the root overlay reports ~1 TB free, and a naive
optimisation would write caches or staged shards there. Two reasons not to:
1. The overlay is a **shared host volume already at 88 %**, so "1 TB free" is other tenants' headroom
and can disappear mid-run. Β§3.13's "never let it fill the disk" must be enforced against a number
we do not control.
2. Nothing outside `/kaggle/working` is retrievable when the session ends, so it is not storage in
any sense that matters.
Default `HF_HOME`/datasets cache resolves under `/root/.cache`, i.e. the overlay, not the 19.5 GB
volume. So **a training job's apparent free space can be large while `/kaggle/working` fills up** β
the failure mode Β§3.13 warns about is invisible to a naive `disk_usage("/")` check. **Every job must
measure the specific filesystem it writes to, and log free space for that path.** On the CPU shape the
effective budget is ~19.5 GB; treat that as the design number for shards resident at once.
No swap: a tokenizer build or a large in-memory dedup on a 30 GiB cgroup must be sized, not streamed
optimistically.
## 5. Network behaviour, including one thing that does not resolve
Working, measured from inside a GPU and a CPU session:
`huggingface.co/api/models` β 200 in 149 ms; `pypi.org/simple/` β 200 in 73 ms;
`datasets/β¦/resolve/main/README.md` β 200 via `api/resolve-cache/β¦`;
`datasets.load_dataset("wikimedia/wikipedia", "20231101.en", streaming=True)` β first row in **7.4 s**,
unauthenticated. `pip install` from PyPI works: `xxhash` in **5.8 s**, import verified.
**Does not resolve:** `cdn-lfs.huggingface.co`, `cdn-lfs-us-1.huggingface.co`, `xethub.huggingface.co`
β `gaierror`. That looked alarming, so probe D tested it directly by pulling 300 MB from two large
public parquet files.
**Sustained Hub download throughput from inside a session: 72.8 MB/s** (`wikimedia/wikipedia`, 745 MB
file) **and 88.9 MB/s** (`HuggingFaceFW/fineweb-edu`, 2334 MB file), with 0.8β0.94 s time-to-first-byte.
So the non-resolving LFS hostnames do not block bulk transfer β `resolve/` redirects land on hosts that
do resolve. **Ingest of the mix is not network-bound:** at a conservative 70 MB/s, one billion tokens
(β3.6 GB of raw text at ~3.6 chars/token, or ~4β6 GB as parquet) downloads in single-digit minutes.
Phase 2 should plan shard staging around CPU and disk limits, not bandwidth.
Also measured in the same probe: `datasets` streaming at **139.5 rows/s β 3.0 MB of text per second**
single-process, unauthenticated (83 rows/s / 1.8 MB/s cold, ~3 MB/s warm). Unauthenticated Hub access
works but the API warns about rate limits (`Please set a HF_TOKEN`); at multi-hundred-shard scale that
is a plausible 429 source, which is an argument for the token-in-job path (D-002) for the *data plane*
even if the *checkpoint plane* stays local.
Enumerating repo files without a client library: `GET https://huggingface.co/api/datasets/<id>/tree/main?recursive=true`
returns `path`/`size`/`type` and works with plain `urllib`. It appeared to cap at 1000 entries
(`fineweb-edu` returned exactly 1000), so **assume pagination is required** for a large repo.
### CPU-side ingest throughput, and what it rules out
Measured on the Kaggle CPU shape (4 vCPU Xeon @ 2.2 GHz) with the gpt2 tokenizer over 6 MB of real
Wikipedia English, `tokenizers 0.22.2` `Tokenizer.encode_batch`:
| cores | tokens/s | hours to tokenize 1B tokens |
|---|---|---|
| 1 | 245,660 | 1.13 |
| 2 | 492,191 | 0.56 |
| 4 | **660,943** | **0.42** |
`chars_per_token` 4.53 on this corpus. Scaling is near-linear to 2 cores and then sub-linear
(2.7x at 4 cores), which is what a GIL-bound parent feeding rayon workers looks like β so a Phase 2
pipeline that shards by **process**, not by thread, should recover the rest. `datasets` streaming ran
at 151.6 rows/s β **3.26 MB of text per second** single-process, unauthenticated.
**The conclusion is a negative result, which is the useful kind:** tokenizing 1B tokens costs ~0.4 h of
the 4-core budget, and at 35β89 MB/s the raw bytes arrive faster than `datasets` can stream them.
**Phase 2 is therefore not CPU-bound and not bandwidth-bound β it is bound by dedup memory and by the
30 GiB cgroup with no swap.** Budget the design accordingly: the dedup/audit stage must be written to
stream and shard against fixed RAM, not to hold a billion sketches, and the only real wall-clock cost
in building the mix is the exact-duplicate/n-gram overlap pass. Note also that download throughput
varied by 2.5x between identical runs (36.9 β 72.8 β 63.2 MB/s), so it should be measured inside the
actual ingest job rather than treated as a constant.
## 6. Accelerator capability envelope (Turing, cc 7.5)
Established inside a GPU session, not assumed:
- **`torch.compile` works** (`warm + 20 steps` of an MLP in 6.0 s, incl. first-call compile).
- **Flash-attention is not available.** `SDPBackend.FLASH_ATTENTION` raises
`No available kernel. Aborting execution.`, preceded by
`Flash attention only supports gpu architectures in the range [sm80, sm121]. Attempting to run on
a sm 7.5 gpu.` β so this is architectural, not a missing dependency, and installing `flash-attn`
cannot fix it. **Β§3's "never touch Hopper-only kernels" now has a concrete instance.**
- **`EFFICIENT_ATTENTION` and `MATH` SDPA backends both run.** Memory-efficient attention is the
viable fast path; note torch warns "Memory efficient attention has been runtime disabled" *when
flash is force-requested*, so backend selection must be explicit.
- `torch.cuda.get_arch_list()` includes `sm_75` (also 70, 80, 86, 90, 100, 120) β this torch build
does have kernels for the card.
- CUDA 12.8, cuDNN 9.10.2, **NCCL 2.27.5** present; `triton 3.6.0` present.
- Absent from the image: `flash-attn`, `xformers`, `trl`, `deepspeed`, `bitsandbytes`, `liger-kernel`,
`evaluate`. Present: `transformers 5.0.0`, `datasets 5.0.0`, `accelerate 1.13.0`,
`tokenizers 0.22.2`, `safetensors 0.7.0`, `numpy 2.0.2`, `pyarrow 24.0.0`, `peft 0.19.1`.
- bf16 is a hardware question, not a torch one: `torch.bfloat16` *tensors* exist, and bf16 on cc 7.5
is at best emulated. Probe C records whether an autocast-bf16 forward even completes and what it
costs, so the Phase 1 precision decision is evidence-backed rather than a recitation of "no bf16".
- `nvidia-smi` has no `multiprocessor_count` query field (rc=2); use `torch.cuda.β¦
get_device_properties(i).multi_processor_count` (40/SM) instead.
## 7. fp16 training trap found by accident
Probe B ran a small Llama with **weights cast to fp16** (`model.cuda().to(torch.float16)`) under
autocast with a plain AdamW and **no GradScaler**: every loss came back `NaN` in 65 steps. That is the
classic fp16 underflow/overflow signature, and it is exactly the kind of bug that would otherwise be
discovered 10 hours into the main run. Probe C re-runs the throughput measurements with the correct
recipe β **fp32 master weights + `torch.autocast(fp16)` + `torch.amp.GradScaler("cuda")` + grad
clipping** β and asserts finiteness, because on T4 there is no bf16 fallback: fp16 with loss scaling
is the only mixed-precision option, so getting it right is not optional.
## 8. What Phase 0 has *not* yet proven
Be honest about the boundary of this gate:
- **Hub writes from inside a Kaggle job** β never attempted. No credential path chosen (D-002 open).
- **NCCL across these two T4s** β still unproven; probe B's attempt was defeated by its own parser,
and one failed `all_reduce` is not evidence either way. Probe C is fixing that now.
- **Checkpoint upload β independent verification β local delete β free-space-confirmed (Β§3.13)** β
untouched; that is a Phase 3 gate, and it needs the real 100M-model sizing arithmetic first.
- **Cold resume on a fresh instance with an empty disk** β Phase 3.
- **Real session caps**: `sessionTimeoutSeconds` was accepted and no session ran long enough to hit
any limit. The actual wall-clock ceiling, idle-timeout behaviour, and preemption policy are
**unknown**, and the multi-week resume plan in Phase 4 depends on them. Probe E
(`ounce100m-p0e-session-cap`) is a ~12.5 h CPU heartbeat running now specifically to find this;
if Kaggle kills it, the last heartbeat is the answer. **Update while writing: the session was still
`RUNNING` at 108 minutes** (submitted 15:38Z, checked 17:26Z), so no CPU cap exists below ~1.8 h. That
already rules out the worst case for Phase 4 β a session that dies inside an hour β and the build is
sized so one source (~5β40 min) is the unit of lost work either way.
## 9. GPU probe C results β the numbers the schedule is built on
Kernel `dodosoomro/ounce100m-p0c-fp16-ddp` v2, GPU image
`gcr.io/kaggle-gpu-images/python@sha256:37c64f7dβ¦7d461`, 2x Tesla T4, charged 157.704 s of quota.
Probe B's two void measurements are now replaced. Model shape used β deliberately near the likely
main-run shape so the arithmetic is about the real run and not a toy:
`hidden 768, layers 12, heads 12, GQA kv_heads 3, FFN 2048, vocab 50257, tied embeddings, seq 1024`
β **112,934,400 parameters**, of which 38,597,376 are the (tied) embedding.
| measurement | result |
|---|---|
| fp32 master weights + `autocast(fp16)` + `GradScaler` + clip | **loss 9.988, finite** β the recipe works; probe B's NaN was the recipe, not the hardware |
| bf16 `autocast` on cc 7.5 | **runs** (loss 10.966, first fwd 0.36 s) β i.e. *emulated*, not tensor-cored |
| tokens/s, 1 GPU, sdpa, bs4Γ1024 | 6,478 |
| tokens/s, 1 GPU, **eager**, bs4Γ1024 | **7,239** |
| tokens/s per rank, DDP 2 ranks, sdpa, bs4Γ1024 | 5,531 β **11,062 aggregate** |
| DDP scaling efficiency | 11,062 / 6,478 = **1.71x** (85 % of ideal) |
| NCCL across these two T4s | **works** β `rc 0`, `backend nccl`, `world 2`, both ranks finite, 30 steps |
| peak CUDA memory, bs4Γ1024, 1 GPU | **11.51 GB** (sdpa) / 13.79 GB (eager) of 14.56 GB usable |
| peak CUDA memory, bs4Γ1024, DDP | 8.68 GB per rank |
| micro-batch ceiling at this shape | **bs 8 OOMs** (`Tried to allocate 1.54 GiB⦠520 MiB free`) |
### What each line forces
- **`eager` beat `sdpa` here (7,239 vs 6,478 tok/s).** At `seq_len 1024` the memory-efficient kernel's
overhead is not repaid, and flash-attention is unavailable on sm 7.5 anyway. Do not assume "flash is
the fast one" β Phase 3 must re-measure at the frozen sequence length, because the crossover is
length-dependent.
- **Memory, not compute, is the binding constraint on batch size.** bs4 fits in 11.5 GB of 14.56 GB;
bs8 does not fit at all. So global batch must be built with **gradient accumulation**, not larger
micro-batches β which also means the usual "increase batch until full" advice is unavailable here.
The fp32 AdamW state for a 100M model (β1.2 GB) plus fp32 master weights is a fixed floor that
does not shrink with batch size.
- **85 % DDP efficiency is good enough to plan on**, and is the number the ETA below uses. It was
measured on 30 steps with NCCL over PCIe between two T4s; treat it as optimistic-typical, not a
guaranteed constant.
- **bf16 running is a trap, not a permission.** Β§2 and Β§8 both rule it out. On Turing it is emulated,
so a config that silently selects bf16 will "work" and be slow. Precision must be asserted to be
fp16 in the frozen config, not left to a library default.
### First ETA, from measurement rather than hope
`1e9 tokens Γ· 11,062 tok/s = 90,400 s β 25.1 wall-clock hours` at the measured aggregate rate.
Against a 30 h weekly allowance accruing at 1x, **the main run fits inside one week of quota on this
shape β but with only ~5 h of headroom**, which is not enough to survive both checkpoint upload time
and the multi-session restarts that Β§4/Phase 4 assume. Three honest caveats on that number:
1. It is a 30-step measurement on one shape that is **over the parameter budget** (112.9 M > 110 M),
so the frozen config will differ and throughput with it.
2. It excludes checkpoint write/upload/verify cycles, `DataLoader` stalls, and any
tokenisation-time-vs-disk tradeoff in the input pipeline.
3. It assumes sessions can run long enough to be efficient; the cap is unknown until probe E reports.
Planning stance for `docs/01-plan.md`: treat **~9β11k tok/s aggregate** as the working band, size the
token target from the *low* end (~28β31 h for 1 B tokens β exceeds a single week, so plan for a
two-week run), and revisit once Gate 3 measures the frozen configuration. Do not let the 25 h figure
become the plan's basis, because it leaves no room for the interruption budget that Β§3.1 explicitly
expects to be spent.
## 10. Credentials and the Kaggle-side secret store
`dodosoomro/ounce100m-p1-credential-path` (CPU, read-only) established:
- Internet-enabled Kaggle sessions **arrive pre-authenticated against the Kaggle API**. The `kaggle`
CLI **2.0.2 is preinstalled**, and `kaggle config view` reports `username: dodosoomro`,
`auth_method: ACCESS_TOKEN`, config from `/root/.config/kaggle` β with no `kaggle.json` written by
us. `kaggle kernels list --mine` returns real rows, so the principal is genuinely usable, not merely
present. Injected env vars: `KAGGLE_API_V1_TOKEN` (32 ch), `KAGGLE_DATA_PROXY_TOKEN` (473),
`KAGGLE_USER_SECRETS_TOKEN` (245).
- **No Hugging Face credential exists in the container**: `HF_TOKEN` and `HUGGING_FACE_HUB_TOKEN` are
both absent, `HF_HOME` unset, `~/.cache/huggingface` does not exist. An anonymous Hub *write* is
correctly rejected β `POST /api/models/β¦/commit/main` β **401 Unauthorized** β while anonymous reads
work. So jobs can pull data freely and cannot push anything until given a token.
- `/kaggle/input` is empty and `/kaggle/input/.secrets/` **does not exist**, so Kaggle's UI-configured
"user input / secrets" facility is unavailable: enabling it is a web-UI action, and there are no
locally usable Kaggle credentials to do it with (`~/.kaggle` absent on this machine; Kaggle auth lives
in the MCP server and inside the container). That rules out option 1 of D-002 for a concrete reason.
- **Consequence: the in-job authenticated Kaggle API is the usable secret store.** A job can create a
**private** Kaggle dataset holding the HF token; later jobs mount it via `datasetDataSources` and read
it from `/kaggle/input/β¦`, so the token appears in exactly one private kernel revision instead of in
every job's source. Β§2 explicitly sanctions "a secret store", so this is the intended mechanism.
**Ordering used, because the failure mode is unrecoverable.** `kaggle datasets create` derives
visibility from a metadata field whose name has moved between CLI versions, and an unrecognised key is
**silently ignored** β which would create a *public* dataset containing the token. So the bootstrap
worker (a) creates the dataset with an inert placeholder only, (b) proves privacy by confirming an
**anonymous** download fails, and (c) uploads the credential as a later version only if (b) passes,
then re-proves. If anonymous read ever succeeds, the secret is never written and the job says so loudly.
**Residual risk, stated:** kernel `p1-secret-bootstrap` v1 necessarily contains the token inline, and
Kaggle retains private kernel revisions β overwrite-after-use does not erase history, and there is no
delete-kernel operation in CLI 2.0.2 or in the MCP toolset. This is a strictly smaller exposure than
putting the token in every training job and in anything that mirrors kernel source.
## 11. Pre-existing Kaggle kernels, not ours
`kaggle kernels list --mine` also returned five kernels predating this project
(`quick-python-cpu-smoke-test`, `monte-carlo-cpu-smoke-test`, `cpu-test`, `exam-80-monte-carlo`,
`workbuddy-cpu-probe`; last run 2026-09-17 β 2026-09-19). They are **not touched** (Β§3.10). This also
explains E-002: `search_notebooks` returning `{}` for this account was the search tool failing, not the
account being empty β so an empty search must never be read as "no prior work".
## 12. Measured notebook session cap: 12.0 h survived, the instance reported 45,000 s
`dodosoomro/ounce100m-p0e-session-cap` β a CPU kernel that printed one heartbeat every two minutes and
nothing else, costing zero GPU quota, left running since 2026-09-19T15:38Z purely as a measurement.
```
CAP_START {"epoch": 1789832305.41, "host": "096ef365b02e", "pid": 7, "limit": 45000}
HB 1 elapsed=0 free_GB=19.5 mem_avail_MB=31204 load=0.41 net=204
HB 361 elapsed=43204 free_GB=19.5 mem_avail_MB=31288 load=0.00 net=204
```
**361 heartbeats, 43,204 s = 12.0 h of continuous running, then the platform took it**
(`CANCEL_ACKNOWLEDGED`, not a crash β the last heartbeat is clean and disk/RAM are unchanged at 19.5 GB
free and 31.3 GB available). The instance itself reported `limit: 45000` s, i.e. **12.5 h**, so the kill
arrived at ~96 % of its own stated budget.
What this does and does not license:
- It is a **CPU** measurement. The GPU accelerator quota is tracked separately, and nothing here proves a
GPU session is allowed to run 12.5 h. So Phase 4 plans `SESSION_GPU_HOURS=6.9` for session 1 β well
inside every plausible cap β and can extend later sessions only on evidence from earlier ones.
- It does settle the question that the plan had been hedging since Β§2 of `04-run-log.md` ("the session cap
is unmeasured"): an interruption at 6 h is not the platform's limit, so a long session is not doomed, and
a `--stop-after-steps` schedule sized for 6-7 h is conservative rather than necessary.
- It confirms the working volume is stable for the whole duration β `free_GB` never moved from 19.5, so an
11 GB drift over a session is not a thing to plan for, and the disk floor in the launcher is about the
mix and checkpoints, not about slow leaks.
|