|
Download code/docs/tutorials/04-gpu-feasibility.md from nima1/stackcraft-clef-flash-lora: direct link, hf CLI and curl.
- Browser
- Download file 16.4 kB
-
https://huggingface.co/nima1/stackcraft-clef-flash-lora/resolve/main/code/docs/tutorials/04-gpu-feasibility.md
- Command line
-
hf download hf://nima1/stackcraft-clef-flash-lora/code/docs/tutorials/04-gpu-feasibility.md
-
curl -L -o 04-gpu-feasibility.md https://huggingface.co/nima1/stackcraft-clef-flash-lora/resolve/main/code/docs/tutorials/04-gpu-feasibility.md
16.4 kB
| # Tutorial 04 — Prove training and reload before running an experiment | |
| This milestone asks a narrow engineering question: can the real pinned Clef-flash | |
| model take gradient steps on this GPU, save all learned parameters, and reproduce | |
| its probabilities in a fresh process? It does not ask whether the model plays | |
| better. The latter requires complete games on the held-out sequences in Milestone 5. | |
| **The bounded real-weight training and fresh-process reload gate passed.** Both | |
| head-only and LoRA-plus-head artifacts reproduced their reference probabilities | |
| exactly. The first LoRA reload exposed a saved-configuration compatibility issue; | |
| the corrected loader and successful retry are documented below. This establishes | |
| local training feasibility, not improved game performance. | |
| ## Start from the native decision model | |
| The base is `Cloudflare/clef-flash`, revision | |
| `17f0b0ad64efb65d273590632833508766b2aae6`. The local loader verifies the pinned | |
| `joint_schema_model.py` source hash and loads an explicitly resolved snapshot. | |
| It requires `trust_pinned_code=True`: pinning and inspecting Python source controls | |
| which code executes; it does not turn that code into a sandbox. | |
| An observation becomes one native choice question containing every legal | |
| placement. The answer is a distribution over action IDs, not a generated command | |
| or natural-language explanation. `encode_observation` checks the complete token | |
| budget before native encoding, because silently truncating a board would change | |
| the task. Training uses only the visible board, current piece, one preview and | |
| legal placements. Dataset seed and episode metadata remain outside the input. | |
| The native encoder sorts option IDs lexicographically. `decision_loss` locates the | |
| teacher's action in `encoded.questions[0].option_ids`; it must not reuse the | |
| engine's rotation/column index as a target index. Both orders contain the same | |
| legal moves, but their order can differ. | |
| Do not train through `systemone()` or `ClefPlayer.choose()`: those are inference | |
| paths. The probe calls `native.collate_records` and then `model(batch)` with | |
| normal autograd enabled. `choose()` is appropriate for the before/after reference | |
| probabilities, where gradients are unnecessary. | |
| ## Keep the backbone small in memory and the head stable | |
| `src/stackcraft/training.py` implements two bounded configurations: | |
| | Configuration | Updated parameters | Purpose | | |
| | --- | --- | --- | | |
| | Head only | Native joint decision head | Isolate head training and checkpoint handling before a backbone backward pass. | | |
| | Rank-4 LoRA plus head | Small added text-layer matrices and the native head | Check whether the task can adapt both text representations and decision scoring. | | |
| The backbone remains BF16. The native decision head is converted to FP32, and | |
| `FP32DecisionHead` disables autocast inside the head. Its floating inputs are | |
| converted to FP32 while token IDs and masks retain their original types. This | |
| keeps the trainable scoring operations in full precision without converting the | |
| whole multi-billion-parameter backbone. | |
| The head reads selected rows of the output vocabulary embedding. Converting the | |
| entire vocabulary matrix to FP32 first creates a large temporary allocation | |
| (roughly 3.79 GiB for this release). `GatheredFloat32Embedding` instead implements | |
| `weight[indices].float()`: select the needed rows first, then cast. This preserves | |
| the actual native head computation and autograd connections. A tiny actual-native- | |
| head test compares outputs and gradients with the full-cast reference exactly. | |
| That test checks the optimization, not full-model training feasibility. | |
| LoRA targets are full module paths under the text transformer layers. They include | |
| both ordinary attention projections and Qwen3.5's linear-attention projections, | |
| as well as the MLP projections. Matching only suffixes such as `q_proj` could | |
| accidentally select the vision encoder; matching only ordinary attention would | |
| miss the hybrid backbone. The vision tower, vocabulary output layer and unrelated | |
| modules are excluded. Original backbone parameters are frozen. The probe uses | |
| rank 4, alpha 8, no LoRA dropout and no trained backbone biases. | |
| Non-reentrant gradient checkpointing recomputes intermediate activations during | |
| backward instead of retaining all of them. This trades time for memory while | |
| preserving gradients through frozen parts of the backbone into LoRA matrices. | |
| A tiny actual Qwen3.5 hybrid-backbone test checks nonzero LoRA gradients in ordinary | |
| attention, linear attention and MLP layers. These are stronger plumbing checks | |
| than testing a stand-in linear model alone, but still do not establish the real | |
| 9B GPU memory requirement. | |
| We start with BF16 rather than immediately adding 4-bit quantized training. The | |
| native model has a custom head and loader; another quantization layer would add | |
| an unverified compatibility dependency. If BF16 fails, record the failure and | |
| make a bounded next decision rather than silently changing the experiment. | |
| ## Prepare before borrowing the GPU | |
| Run these commands from the repository root or the public bundle's `code/` | |
| directory. Install the ML tools and cache the pinned release before testing: | |
| ```bash | |
| uv sync --locked --extra ml --group dev | |
| uv run --locked --extra ml hf download Cloudflare/clef-flash \ | |
| --revision 17f0b0ad64efb65d273590632833508766b2aae6 | |
| ``` | |
| Downloading only populates the cache; it does not load GPU weights. It can run | |
| while other GPU services remain available. The actual native-head tests need | |
| this cached source. Explicitly enable the native tokenizer/encoder check too: | |
| ```bash | |
| STACKCRAFT_TEST_NATIVE_ENCODING=1 uv run --locked --extra ml \ | |
| pytest tests/test_training.py tests/test_clef.py | |
| ``` | |
| These checks use CPU fixtures and the real tokenizer/source; they do not load | |
| the full backbone or establish GPU feasibility. Inspect the test summary: a | |
| skipped test is not a passed native-model test. | |
| The training probe expects the audited `data/study-v1` bundle established by | |
| Tutorial03's exact-study reconstruction or a verified dataset download. If you | |
| used a different directory, pass that same path with `--dataset` below. | |
| It validates the train/validation bundle, then uses only the first four training | |
| positions. It does not sample held-out test seeds. Record the current source | |
| commit and any uncommitted probe changes in `record.md` before running. | |
| This workstation normally runs a local GLM Docker service. The user explicitly | |
| authorized temporarily stopping that service and restoring it after these | |
| Stackcraft experiments. `scripts/gpu_session.py` is specific to that approved | |
| service; it is not a generic instruction to stop another machine's workloads. | |
| It first requires the configured container to be running, then stops it, runs a | |
| bounded child job, and attempts restoration during cleanup. The training script | |
| itself never stops services. It refuses to load unless CUDA is available and at | |
| least **25 GiB is free**. Free-memory admission is a precondition, not a guarantee | |
| that backward will fit. | |
| ## Run the bounded real-weight probe | |
| The initial budget is two configurations and at most 30 minutes of child-job | |
| time. The commands below allocate 20 minutes to the training probe and 10 minutes | |
| to the separate reload check. Choose fresh output and session-record paths on a | |
| retry; neither existing evidence nor checkpoints should be overwritten. | |
| ```bash | |
| uv run --locked --extra ml python scripts/gpu_session.py \ | |
| --timeout 1200 --record runs/gpu-training-probe-v1.json -- \ | |
| uv run --locked --extra ml python scripts/probe_clef_training.py \ | |
| --dataset data/study-v1 --output runs/clef-training-probe-v1 | |
| ``` | |
| The probe first records unchanged native probabilities. It measures the initial | |
| probability drift introduced by moving the head to FP32, before training; a | |
| precision change must not be mistaken for a learned improvement. It then runs | |
| three head-only steps. After saving that checkpoint, it releases the model and | |
| loads a fresh unchanged base for five LoRA-plus-head steps. The LoRA experiment | |
| therefore does not inherit the head-only warm-up. | |
| Both configurations use batch size one, AdamW with learning rate `1e-5`, norm | |
| clipping at 1.0, and this loss: | |
| smoothed_cross_entropy(label_smoothing=0.05) | |
| + 0.1 * sum_over_options((probability - one_hot_teacher_label)**2) | |
| The cross-entropy term learns the expert action. Smoothing avoids assigning the | |
| entire training target to one action. The Brier term also penalizes the predicted | |
| probability distribution's distance from the teacher target. These fixed choices | |
| are a probe configuration, not evidence of calibrated confidence or optimal | |
| hyperparameters. | |
| For every step, the probe checks finite loss and gradients and records nonzero | |
| head/LoRA gradient sums, tokens, elapsed time, and CUDA peak allocated/reserved | |
| bytes. Before/after parameter hashes check that intended parameters changed and | |
| frozen parameters did not. Hashing the entire frozen backbone costs time but | |
| checks something that `requires_grad=False` alone cannot prove: what bytes actually | |
| changed during this run. | |
| Inspect `runs/clef-training-probe-v1/report.json` and the session record. A timeout, | |
| CUDA error, absent gradient or failed hash check is a failed gate. A report left at | |
| `running` after forced termination is incomplete evidence, even if some checkpoint | |
| files exist. Do not start long training merely because loss fell on four examples. | |
| ## Save the head as well as the adapter | |
| A LoRA checkpoint has this structure: | |
| ```text | |
| checkpoint/ | |
| joint_head.safetensors | |
| adapter/ | |
| adapter_config.json | |
| adapter_model.safetensors | |
| training_config.json | |
| reference.json | |
| ``` | |
| The head is an independently trained part of Clef; saving only the PEFT adapter | |
| would omit learned parameters. `joint_head.safetensors` stores FP32 native head | |
| keys without wrapper-specific prefixes. The adapter stores only the added LoRA | |
| weights, not another copy of the frozen base. Metadata identifies the base | |
| revision, native source hash, input encoding, head format, LoRA settings and | |
| teacher rows. The loader checks this contract, head shapes and dtypes, and actual | |
| saved adapter configuration before restoring weights. | |
| The head-only probe writes `head-checkpoint/` without an adapter. It is diagnostic | |
| output, not a selected release model. | |
| ## Require a separate fresh-process reload | |
| Only after the first probe succeeds, run: | |
| ```bash | |
| uv run --locked --extra ml python scripts/gpu_session.py \ | |
| --timeout 600 --record runs/gpu-training-reload-v1.json -- \ | |
| uv run --locked --extra ml python scripts/probe_clef_training.py \ | |
| --dataset data/study-v1 --output runs/clef-training-reload-v1 \ | |
| --reload runs/clef-training-probe-v1/checkpoint | |
| ``` | |
| This starts another Python process, loads the unchanged pinned base, restores both | |
| adapter and head, and evaluates the same four training observations. Require the | |
| same option IDs and a **maximum absolute probability difference of at most | |
| `1e-4`** across all recorded options. The tolerance was declared before the run. | |
| Compare unrounded probabilities; rounded display values can conceal a mismatch. | |
| A same-process test could accidentally reuse trained parameters that were never | |
| saved. The fresh process tests the actual inference artifact. A passing reload | |
| establishes serialization fidelity on these references, not general game skill | |
| or correctness on every possible input. | |
| ## Restore service and record the limits | |
| Inspect both `restored_running` and `restored_healthy` after success and failure. | |
| The wrapper checks `/health` inside the restored container and allows up to 120 | |
| seconds for readiness, retrying transient timeouts and malformed startup responses. | |
| It does not query the unrelated service on host port 8080. A running process alone | |
| would not establish readiness. Nested cleanup ensures a child-termination error | |
| still reaches restoration; repeated ordinary interrupts are ignored during that | |
| cleanup. Preserve the wrapper record, probe reports and console logs, including | |
| errors. Mocked tests exercise these recovery paths without touching Docker. | |
| Process cleanup cannot guarantee recovery after `SIGKILL`, host power loss or a | |
| Docker daemon failure. If the wrapper was forcibly terminated, inspect the exact | |
| approved container and any remaining job before acting. Restore the existing | |
| service only after the experiment has released GPU memory; do not start another | |
| probe simply because the last terminal stopped printing output. | |
| The earlier M2 real-weight development pilot is useful context: on this RTX 5090, | |
| it recorded 55 unchanged-Clef decisions at a mean of about **0.144 seconds per | |
| decision**, with **19,686,539,776 bytes** peak CUDA allocation (about 19.7 GB, | |
| 18.3 GiB). The benchmark script took 12.60 seconds; the stop/job/restore wrapper | |
| interval was about 14.20 seconds and recorded successful restoration. These are | |
| inference measurements on two short development games, not training memory or | |
| held-out quality results. Sources are `runs/clef-development-v1.json` and | |
| `runs/gpu-baseline-session.json`. | |
| ## Measured training probe and the first reload failure | |
| The real-weight probe passed on 2026-10-05 using the RTX 5090 and Torch | |
| `2.14.1+cu130`. Evidence is `runs/probe-v1/report.json`; the actual run directory | |
| differs from the fresh reproduction paths shown above. It used training positions | |
| `seed-10000-turn-0` through `seed-10000-turn-3`, with 1,302–2,252 encoded tokens. | |
| | Measurement | Head-only probe | LoRA-plus-head probe | | |
| | --- | ---: | ---: | | |
| | Optimization steps | 3 | 5 | | |
| | Observed step times | 0.154–0.328 s | 0.670–1.322 s | | |
| | Peak allocated CUDA bytes | 21,323,760,640 | 23,180,389,888 | | |
| | Peak reserved CUDA bytes | 21,541,945,344 | 23,595,057,152 | | |
| | Intended parameter changes | Verified | Verified | | |
| | Frozen parameter hashes | Unchanged | Unchanged | | |
| All reported losses and gradients were finite, with nonzero gradients in the | |
| intended groups. Total probe time was **51.81 seconds**, including model reload, | |
| parameter hashing and checkpoint work. This is not an estimate of a full epoch: | |
| only four short training positions were exercised. Peak reserved memory is the | |
| PyTorch allocator's reservation, not the same quantity as all memory reported by | |
| `nvidia-smi`. | |
| The FP32 head changed initial probabilities by at most **0.00074412** before any | |
| training. That is a precision effect relative to the original BF16 head, not a | |
| failed serialization comparison. The `1e-4` reload threshold compares the saved | |
| trained FP32-head model with its own fresh reload. | |
| The first fresh reload, `runs/probe-reload-v1/report.json`, failed after 4.15 | |
| seconds because the saved adapter configuration differed from our metadata. PEFT | |
| automatically shortened **248 full target paths to 12 module suffixes**. The strict | |
| loader correctly refused a mismatch; the failure does not by itself mean the | |
| weight tensors are corrupt. Accepting suffixes without checking what they match | |
| could attach an adapter to unintended modules. The correction must expand saved | |
| target selectors against the actual pinned backbone and require exactly the | |
| intended full target set, rejecting any extra destinations. This is why fresh | |
| reload is a milestone gate rather than a final packaging chore. | |
| The corrected loader resolves saved selectors against the pinned backbone and | |
| requires the exact intended destinations. The retry preserved the original | |
| checkpoint and failed report; it changed validation of equivalent target selectors, | |
| not the trained weights. A regression test covers target shortening and rejects | |
| selectors that also match unintended modules. | |
| Both independent fresh-process checks then passed: | |
| | Artifact | Report | Maximum probability difference | Process time | | |
| | --- | --- | ---: | ---: | | |
| | Head only | `runs/probe-head-reload-v1/report.json` | 0.0 | 5.44 s | | |
| | LoRA plus head | `runs/probe-reload-v2/report.json` | 0.0 | 5.65 s | | |
| The LoRA retry used source commit | |
| `f7be7ceee9c38566dfebad0ebf5f1c04958faee3`. The dataset manifest SHA-256 was | |
| `aeddcfc1ea5390f12122d2e786c1e2030dcc97d3790b6d7523432548198b1f94` throughout. | |
| The approved GLM service was restored and passed its health check after these | |
| probes. **M4 is complete.** Milestone 5 can now train a selected checkpoint and | |
| measure complete games; the four-position probe does not demonstrate that the | |
| trained policy wins or generalizes. | |