stackcraft-clef-flash-lora / code /docs /tutorials /04-gpu-feasibility.md
nima1's picture
Publish verified checkpoint and losslessly compressed study evidence
4be6a52 verified
|
Raw History Blame Contribute Delete
16.4 kB

Tutorial 04 — Prove training and reload before running an experiment

This milestone asks a narrow engineering question: can the real pinned Clef-flash model take gradient steps on this GPU, save all learned parameters, and reproduce its probabilities in a fresh process? It does not ask whether the model plays better. The latter requires complete games on the held-out sequences in Milestone 5.

The bounded real-weight training and fresh-process reload gate passed. Both head-only and LoRA-plus-head artifacts reproduced their reference probabilities exactly. The first LoRA reload exposed a saved-configuration compatibility issue; the corrected loader and successful retry are documented below. This establishes local training feasibility, not improved game performance.

Start from the native decision model

The base is Cloudflare/clef-flash, revision 17f0b0ad64efb65d273590632833508766b2aae6. The local loader verifies the pinned joint_schema_model.py source hash and loads an explicitly resolved snapshot. It requires trust_pinned_code=True: pinning and inspecting Python source controls which code executes; it does not turn that code into a sandbox.

An observation becomes one native choice question containing every legal placement. The answer is a distribution over action IDs, not a generated command or natural-language explanation. encode_observation checks the complete token budget before native encoding, because silently truncating a board would change the task. Training uses only the visible board, current piece, one preview and legal placements. Dataset seed and episode metadata remain outside the input.

The native encoder sorts option IDs lexicographically. decision_loss locates the teacher's action in encoded.questions[0].option_ids; it must not reuse the engine's rotation/column index as a target index. Both orders contain the same legal moves, but their order can differ.

Do not train through systemone() or ClefPlayer.choose(): those are inference paths. The probe calls native.collate_records and then model(batch) with normal autograd enabled. choose() is appropriate for the before/after reference probabilities, where gradients are unnecessary.

Keep the backbone small in memory and the head stable

src/stackcraft/training.py implements two bounded configurations:

Configuration Updated parameters Purpose
Head only Native joint decision head Isolate head training and checkpoint handling before a backbone backward pass.
Rank-4 LoRA plus head Small added text-layer matrices and the native head Check whether the task can adapt both text representations and decision scoring.

The backbone remains BF16. The native decision head is converted to FP32, and FP32DecisionHead disables autocast inside the head. Its floating inputs are converted to FP32 while token IDs and masks retain their original types. This keeps the trainable scoring operations in full precision without converting the whole multi-billion-parameter backbone.

The head reads selected rows of the output vocabulary embedding. Converting the entire vocabulary matrix to FP32 first creates a large temporary allocation (roughly 3.79 GiB for this release). GatheredFloat32Embedding instead implements weight[indices].float(): select the needed rows first, then cast. This preserves the actual native head computation and autograd connections. A tiny actual-native- head test compares outputs and gradients with the full-cast reference exactly. That test checks the optimization, not full-model training feasibility.

LoRA targets are full module paths under the text transformer layers. They include both ordinary attention projections and Qwen3.5's linear-attention projections, as well as the MLP projections. Matching only suffixes such as q_proj could accidentally select the vision encoder; matching only ordinary attention would miss the hybrid backbone. The vision tower, vocabulary output layer and unrelated modules are excluded. Original backbone parameters are frozen. The probe uses rank 4, alpha 8, no LoRA dropout and no trained backbone biases.

Non-reentrant gradient checkpointing recomputes intermediate activations during backward instead of retaining all of them. This trades time for memory while preserving gradients through frozen parts of the backbone into LoRA matrices. A tiny actual Qwen3.5 hybrid-backbone test checks nonzero LoRA gradients in ordinary attention, linear attention and MLP layers. These are stronger plumbing checks than testing a stand-in linear model alone, but still do not establish the real 9B GPU memory requirement.

We start with BF16 rather than immediately adding 4-bit quantized training. The native model has a custom head and loader; another quantization layer would add an unverified compatibility dependency. If BF16 fails, record the failure and make a bounded next decision rather than silently changing the experiment.

Prepare before borrowing the GPU

Run these commands from the repository root or the public bundle's code/ directory. Install the ML tools and cache the pinned release before testing:

uv sync --locked --extra ml --group dev
uv run --locked --extra ml hf download Cloudflare/clef-flash \
  --revision 17f0b0ad64efb65d273590632833508766b2aae6

Downloading only populates the cache; it does not load GPU weights. It can run while other GPU services remain available. The actual native-head tests need this cached source. Explicitly enable the native tokenizer/encoder check too:

STACKCRAFT_TEST_NATIVE_ENCODING=1 uv run --locked --extra ml \
  pytest tests/test_training.py tests/test_clef.py

These checks use CPU fixtures and the real tokenizer/source; they do not load the full backbone or establish GPU feasibility. Inspect the test summary: a skipped test is not a passed native-model test.

The training probe expects the audited data/study-v1 bundle established by Tutorial03's exact-study reconstruction or a verified dataset download. If you used a different directory, pass that same path with --dataset below. It validates the train/validation bundle, then uses only the first four training positions. It does not sample held-out test seeds. Record the current source commit and any uncommitted probe changes in record.md before running.

This workstation normally runs a local GLM Docker service. The user explicitly authorized temporarily stopping that service and restoring it after these Stackcraft experiments. scripts/gpu_session.py is specific to that approved service; it is not a generic instruction to stop another machine's workloads. It first requires the configured container to be running, then stops it, runs a bounded child job, and attempts restoration during cleanup. The training script itself never stops services. It refuses to load unless CUDA is available and at least 25 GiB is free. Free-memory admission is a precondition, not a guarantee that backward will fit.

Run the bounded real-weight probe

The initial budget is two configurations and at most 30 minutes of child-job time. The commands below allocate 20 minutes to the training probe and 10 minutes to the separate reload check. Choose fresh output and session-record paths on a retry; neither existing evidence nor checkpoints should be overwritten.

uv run --locked --extra ml python scripts/gpu_session.py \
  --timeout 1200 --record runs/gpu-training-probe-v1.json -- \
  uv run --locked --extra ml python scripts/probe_clef_training.py \
  --dataset data/study-v1 --output runs/clef-training-probe-v1

The probe first records unchanged native probabilities. It measures the initial probability drift introduced by moving the head to FP32, before training; a precision change must not be mistaken for a learned improvement. It then runs three head-only steps. After saving that checkpoint, it releases the model and loads a fresh unchanged base for five LoRA-plus-head steps. The LoRA experiment therefore does not inherit the head-only warm-up.

Both configurations use batch size one, AdamW with learning rate 1e-5, norm clipping at 1.0, and this loss:

smoothed_cross_entropy(label_smoothing=0.05)
  + 0.1 * sum_over_options((probability - one_hot_teacher_label)**2)

The cross-entropy term learns the expert action. Smoothing avoids assigning the entire training target to one action. The Brier term also penalizes the predicted probability distribution's distance from the teacher target. These fixed choices are a probe configuration, not evidence of calibrated confidence or optimal hyperparameters.

For every step, the probe checks finite loss and gradients and records nonzero head/LoRA gradient sums, tokens, elapsed time, and CUDA peak allocated/reserved bytes. Before/after parameter hashes check that intended parameters changed and frozen parameters did not. Hashing the entire frozen backbone costs time but checks something that requires_grad=False alone cannot prove: what bytes actually changed during this run.

Inspect runs/clef-training-probe-v1/report.json and the session record. A timeout, CUDA error, absent gradient or failed hash check is a failed gate. A report left at running after forced termination is incomplete evidence, even if some checkpoint files exist. Do not start long training merely because loss fell on four examples.

Save the head as well as the adapter

A LoRA checkpoint has this structure:

checkpoint/
  joint_head.safetensors
  adapter/
    adapter_config.json
    adapter_model.safetensors
  training_config.json
  reference.json

The head is an independently trained part of Clef; saving only the PEFT adapter would omit learned parameters. joint_head.safetensors stores FP32 native head keys without wrapper-specific prefixes. The adapter stores only the added LoRA weights, not another copy of the frozen base. Metadata identifies the base revision, native source hash, input encoding, head format, LoRA settings and teacher rows. The loader checks this contract, head shapes and dtypes, and actual saved adapter configuration before restoring weights.

The head-only probe writes head-checkpoint/ without an adapter. It is diagnostic output, not a selected release model.

Require a separate fresh-process reload

Only after the first probe succeeds, run:

uv run --locked --extra ml python scripts/gpu_session.py \
  --timeout 600 --record runs/gpu-training-reload-v1.json -- \
  uv run --locked --extra ml python scripts/probe_clef_training.py \
  --dataset data/study-v1 --output runs/clef-training-reload-v1 \
  --reload runs/clef-training-probe-v1/checkpoint

This starts another Python process, loads the unchanged pinned base, restores both adapter and head, and evaluates the same four training observations. Require the same option IDs and a maximum absolute probability difference of at most 1e-4 across all recorded options. The tolerance was declared before the run. Compare unrounded probabilities; rounded display values can conceal a mismatch.

A same-process test could accidentally reuse trained parameters that were never saved. The fresh process tests the actual inference artifact. A passing reload establishes serialization fidelity on these references, not general game skill or correctness on every possible input.

Restore service and record the limits

Inspect both restored_running and restored_healthy after success and failure. The wrapper checks /health inside the restored container and allows up to 120 seconds for readiness, retrying transient timeouts and malformed startup responses. It does not query the unrelated service on host port 8080. A running process alone would not establish readiness. Nested cleanup ensures a child-termination error still reaches restoration; repeated ordinary interrupts are ignored during that cleanup. Preserve the wrapper record, probe reports and console logs, including errors. Mocked tests exercise these recovery paths without touching Docker.

Process cleanup cannot guarantee recovery after SIGKILL, host power loss or a Docker daemon failure. If the wrapper was forcibly terminated, inspect the exact approved container and any remaining job before acting. Restore the existing service only after the experiment has released GPU memory; do not start another probe simply because the last terminal stopped printing output.

The earlier M2 real-weight development pilot is useful context: on this RTX 5090, it recorded 55 unchanged-Clef decisions at a mean of about 0.144 seconds per decision, with 19,686,539,776 bytes peak CUDA allocation (about 19.7 GB, 18.3 GiB). The benchmark script took 12.60 seconds; the stop/job/restore wrapper interval was about 14.20 seconds and recorded successful restoration. These are inference measurements on two short development games, not training memory or held-out quality results. Sources are runs/clef-development-v1.json and runs/gpu-baseline-session.json.

Measured training probe and the first reload failure

The real-weight probe passed on 2026-10-05 using the RTX 5090 and Torch 2.14.1+cu130. Evidence is runs/probe-v1/report.json; the actual run directory differs from the fresh reproduction paths shown above. It used training positions seed-10000-turn-0 through seed-10000-turn-3, with 1,302–2,252 encoded tokens.

Measurement Head-only probe LoRA-plus-head probe
Optimization steps 3 5
Observed step times 0.154–0.328 s 0.670–1.322 s
Peak allocated CUDA bytes 21,323,760,640 23,180,389,888
Peak reserved CUDA bytes 21,541,945,344 23,595,057,152
Intended parameter changes Verified Verified
Frozen parameter hashes Unchanged Unchanged

All reported losses and gradients were finite, with nonzero gradients in the intended groups. Total probe time was 51.81 seconds, including model reload, parameter hashing and checkpoint work. This is not an estimate of a full epoch: only four short training positions were exercised. Peak reserved memory is the PyTorch allocator's reservation, not the same quantity as all memory reported by nvidia-smi.

The FP32 head changed initial probabilities by at most 0.00074412 before any training. That is a precision effect relative to the original BF16 head, not a failed serialization comparison. The 1e-4 reload threshold compares the saved trained FP32-head model with its own fresh reload.

The first fresh reload, runs/probe-reload-v1/report.json, failed after 4.15 seconds because the saved adapter configuration differed from our metadata. PEFT automatically shortened 248 full target paths to 12 module suffixes. The strict loader correctly refused a mismatch; the failure does not by itself mean the weight tensors are corrupt. Accepting suffixes without checking what they match could attach an adapter to unintended modules. The correction must expand saved target selectors against the actual pinned backbone and require exactly the intended full target set, rejecting any extra destinations. This is why fresh reload is a milestone gate rather than a final packaging chore.

The corrected loader resolves saved selectors against the pinned backbone and requires the exact intended destinations. The retry preserved the original checkpoint and failed report; it changed validation of equivalent target selectors, not the trained weights. A regression test covers target shortening and rejects selectors that also match unintended modules.

Both independent fresh-process checks then passed:

Artifact Report Maximum probability difference Process time
Head only runs/probe-head-reload-v1/report.json 0.0 5.44 s
LoRA plus head runs/probe-reload-v2/report.json 0.0 5.65 s

The LoRA retry used source commit f7be7ceee9c38566dfebad0ebf5f1c04958faee3. The dataset manifest SHA-256 was aeddcfc1ea5390f12122d2e786c1e2030dcc97d3790b6d7523432548198b1f94 throughout. The approved GLM service was restored and passed its health check after these probes. M4 is complete. Milestone 5 can now train a selected checkpoint and measure complete games; the four-position probe does not demonstrate that the trained policy wins or generalizes.