|
Download code/docs/tutorials/05-evaluation.md from nima1/stackcraft-clef-flash-lora: direct link, hf CLI and curl.
- Browser
- Download file 15.4 kB
-
https://huggingface.co/nima1/stackcraft-clef-flash-lora/resolve/main/code/docs/tutorials/05-evaluation.md
- Command line
-
hf download hf://nima1/stackcraft-clef-flash-lora/code/docs/tutorials/05-evaluation.md
-
curl -L -o 05-evaluation.md https://huggingface.co/nima1/stackcraft-clef-flash-lora/resolve/main/code/docs/tutorials/05-evaluation.md
15.4 kB
| # Tutorial 05 — Select on validation, then measure complete games | |
| Training, validation selection, selected-checkpoint reload and the held-out game | |
| evaluation are complete. Epoch 02 cleared **16.81 lines per game**, versus **0.07** | |
| for native Clef and **76.53** for the cheap heuristic on all 200 reserved seeds. | |
| This tutorial gives the frozen procedure and explains the measured outcome. See | |
| [the final study](../../reports/final-study.md) for full results and limitations, | |
| and [the compact JSON](../../reports/final-summary.json) for exact metrics and hashes. | |
| The experiment results are complete; release verification is tracked separately. | |
| ## Separate learning, selection and final testing | |
| The two candidate checkpoints are epochs 01 and 02 of the same seed-42 LoRA run. | |
| Their training configuration and dataset are fixed before comparison. Evaluate | |
| both on **all 215 validation positions** and select the lowest finite mean target | |
| negative log likelihood (NLL). An exact tie chooses the earlier epoch. A candidate | |
| with inference errors, missing probability distributions, or nonfinite NLL is | |
| ineligible. If both are ineligible, stop before opening the test set. | |
| NLL measures how much probability the model assigns to the teacher's recommended | |
| placement. Selecting on a smooth probability score can distinguish checkpoints | |
| even when their top choices tie. Teacher agreement and Brier score remain useful | |
| diagnostics, but cannot override the preregistered rule after results are visible. | |
| Copying the teacher is not the same as playing well. Each small mistake changes | |
| the next board; the model eventually encounters states different from its training | |
| examples. The teacher itself uses a limited search horizon and a heuristic value | |
| function. Final evaluation must therefore play complete games on unseen sequences. | |
| ## Train exactly the preregistered candidates | |
| Pass the real GPU feasibility and fresh-process reload gates in Tutorial 04 first. | |
| Use the exact 827/215-row dataset from Tutorial 03. The training runner rejects a | |
| substituted dataset even when its labels look similar: hashes bind the experiment | |
| to specific bytes and provenance. Use a fresh run directory when reproducing. | |
| ```bash | |
| uv sync --locked --extra ml --group dev | |
| uv run --locked --extra ml python scripts/train_clef.py \ | |
| --dataset data/study-v1 --output runs/study-v1 \ | |
| --mode lora --epochs 2 --learning-rate 1e-5 \ | |
| --accumulation 8 --max-length 4096 | |
| ``` | |
| This command needs a free GPU. On the original workstation only, run it through | |
| the authorized stop/restore wrapper from Tutorial 04 with a 3,600-second timeout. | |
| The script uses seed 42, batch size one, rank-4/alpha-8 LoRA with zero dropout, | |
| AdamW weight decay 0.01, gradient norm clipping at 1.0, and the fixed smoothed | |
| cross-entropy plus Brier loss from Tutorial 04. The backbone stays BF16 and frozen; | |
| LoRA matrices and the native FP32 head learn. Each epoch visits all training rows | |
| in a recorded deterministic shuffled order. Validation rows never train the model. | |
| Accumulating eight single-example gradients fits the memory budget while averaging | |
| over several boards per optimizer step. The final group has only three rows and | |
| is divided by three, not eight. Two complete epochs give two declared candidates | |
| without an open-ended search. Neither falling loss nor a promising demo authorizes | |
| extra epochs after the test set has been opened. The script defaults to one epoch, | |
| so the explicit `--epochs 2` above is essential for reproducing this study. | |
| The completed run processed 827 examples and 104 optimizer steps in each epoch: | |
| | Epoch | Mean training loss | Time | | |
| | --- | ---: | ---: | | |
| | 01 | 2.577967 | 742.48 s | | |
| | 02 | 1.966718 | 744.66 s | | |
| Total training time including hashes and checkpoint work was 1,511.25 seconds | |
| (about 25.2 minutes). Peak CUDA allocation was 23.23 GB, or about 21.64 GiB. | |
| The head and LoRA parameters changed, while all frozen parameter hashes remained | |
| unchanged. These are training measurements, not held-out performance. See | |
| [the training report](../../reports/study-training.json), | |
| [the original launch source](../../reports/training-launch-source.json), and | |
| [the archived preregistration](../../reports/study-preregistration.md). | |
| ## Prepare checkpoint metadata before freezing it | |
| PEFT can write a generic adapter README with placeholders and a local cache path. | |
| Keep raw training outputs immutable. Prepare candidate copies before validation: | |
| ```bash | |
| uv run --locked --extra ml python scripts/prepare_checkpoint.py \ | |
| --source runs/study-v1/epoch-01 --output checkpoints/candidates/epoch-01 | |
| uv run --locked --extra ml python scripts/prepare_checkpoint.py \ | |
| --source runs/study-v1/epoch-02 --output checkpoints/candidates/epoch-02 | |
| ``` | |
| The script replaces only `adapter/README.md`, compares every other file byte for | |
| byte, and saves a sibling preparation record. It does not alter learned weights | |
| or training settings. Validate and select these exact copies. Once selected, | |
| even a README edit would break the frozen checkpoint file mapping. See Tutorial 06 | |
| for how release documentation stays outside the checkpoint directory. | |
| ## Generate validation evidence | |
| These commands load real GPU models. Use the memory admission and service-restoration | |
| procedure from Tutorial 04. The evaluator never downloads weights implicitly or | |
| stops another workload. Replace paths if your run directory differs. | |
| ```sh | |
| uv run --locked --extra ml python scripts/evaluate_clef.py positions \ | |
| --dataset data/study-v1 --checkpoint checkpoints/candidates/epoch-01 \ | |
| --players base base-fp32 trained --output runs/validation-epoch-01 | |
| uv run --locked --extra ml python scripts/evaluate_clef.py positions \ | |
| --dataset data/study-v1 --checkpoint checkpoints/candidates/epoch-02 \ | |
| --players trained --output runs/validation-epoch-02 | |
| ``` | |
| Every validation row gets an unrounded probability record or an explicit error. | |
| The position report binds those records to checkpoint file hashes and the dataset | |
| manifest. Aggregate NLL and Brier are unavailable if any probability distribution | |
| is missing; a successful subset cannot masquerade as a complete evaluation. | |
| The first command also measures two unchanged-model references. `base` retains | |
| the original native BF16 head. `base-fp32` keeps those same learned weights but | |
| uses our FP32 head wrapper. Training uses that wrapper, so the extra comparison | |
| shows whether a change comes from numerical precision rather than learning. | |
| The primary comparison remains trained versus native base. All three share the | |
| same complete observation encoding and BF16 backbone. | |
| ## Freeze the choice before running test episodes | |
| ```sh | |
| uv run --locked --extra ml python scripts/select_checkpoint.py \ | |
| --dataset data/study-v1 \ | |
| --candidate01 checkpoints/candidates/epoch-01 runs/validation-epoch-01/report.json \ | |
| --candidate02 checkpoints/candidates/epoch-02 runs/validation-epoch-02/report.json \ | |
| --output runs/selection-v1 | |
| ``` | |
| The selector independently recomputes metrics from every raw prediction against | |
| the audited validation rows. It verifies actual checkpoint bytes, training seed | |
| and configuration, epoch identity, common dataset and evaluation provenance. It | |
| accepts exactly the two preregistered candidates. It then copies both validation | |
| reports, prediction files and checkpoint metadata into a self-contained evidence | |
| directory. The large model weights remain in their checkpoint directories and | |
| are referenced by hashes. | |
| `selection.json` records the selected epoch, decision rule, both candidates, | |
| their eligibility, validation metrics, and the frozen test protocol. The runner | |
| rechecks that the selected epoch actually wins the declared rule. Manual editing | |
| of a label such as "best model" is not sufficient to unlock final evaluation. | |
| Use a new output directory for a new selection attempt; existing evidence is not | |
| silently overwritten. | |
| ## Inspect selection and reload the selected artifact | |
| The measured validation results used all 215 positions for every policy, with | |
| zero errors and complete probability distributions: | |
| | Policy | Teacher agreement | Mean target NLL | Mean Brier score | | |
| | --- | ---: | ---: | ---: | | |
| | Native base | 21/215 (9.77%) | 2.952702 | 0.923173 | | |
| | Unchanged FP32 head | 22/215 (10.23%) | 2.952393 | 0.923102 | | |
| | Epoch 01 | 89/215 (41.40%) | 1.860870 | 0.722190 | | |
| | Epoch 02 | 92/215 (42.79%) | 1.775923 | 0.708980 | | |
| The declared NLL rule selected epoch 02. The saved | |
| [validation summary](../../reports/validation-summary.json) identifies its complete | |
| file hashes and the selection hash. Teacher agreement improved on these positions; | |
| that still does not establish better complete-game play. | |
| Before consuming test seeds, reload the selected artifact in another process: | |
| ```bash | |
| uv run --locked --extra ml python scripts/probe_clef_training.py \ | |
| --dataset data/study-v1 --reload checkpoints/candidates/epoch-02 \ | |
| --output runs/selected-reload-v1 | |
| ``` | |
| This command also needs the GPU admission/restoration procedure. In this study it | |
| reproduced all four training-position reference distributions with **0.0 maximum | |
| absolute difference**, below the predeclared `1e-4` tolerance. See | |
| [the reload report](../../reports/selected-checkpoint-reload.json). | |
| The references use training positions, so this check does not consume test games. | |
| ## Run the paired test once | |
| Use the selected checkpoint path from the evidence, not whichever epoch you hoped | |
| would win. The reserved pool is all **200 seeds, 30000–30199**, with a **200-piece | |
| cap**. Each player gets the same piece sequence for a given seed. There are five | |
| players: native base, FP32-head base ablation, selected trained model, random and | |
| the fixed cheap heuristic. | |
| ```sh | |
| uv run --locked --extra ml python scripts/evaluate_clef.py tournament \ | |
| --checkpoint SELECTED_CHECKPOINT \ | |
| --seeds 30000:30200 --max-pieces 200 --final-test \ | |
| --selection-file runs/selection-v1/selection.json \ | |
| --output runs/final-evaluation | |
| ``` | |
| `--final-test` explicitly acknowledges that these held-out seeds are being consumed. | |
| Without it, any overlap with the reserved pool is rejected before model loading. | |
| The cap and full seed pool cannot silently shrink to finish sooner. The initial | |
| development timing estimate permits hours in the worst case; short early deaths | |
| can make actual evaluation much faster. | |
| Each completed episode is saved atomically, with actions, probabilities, runtime | |
| configuration, errors and a replay. To continue an interrupted run, repeat the same | |
| command with `--resume`. The runner checks source/dependency hashes, checkpoint | |
| identity, selection, and saved replay outcomes. It skips completed players without | |
| loading their models again. Recorded errors remain part of the experiment, rather | |
| than being retried until a nicer result appears. | |
| ## Read the report without hiding failures | |
| The primary outcome is lines cleared. Score and pieces survived are secondary. | |
| The report includes means, medians, cap-hit rate, failure counts, latency and input | |
| lengths. The [hardware snapshot](../../reports/evaluation-hardware.json) records | |
| the RTX 5090, driver, CPU, Python and inference settings observed during this run. | |
| Compare latency only with that environment in mind; the snapshot does not measure | |
| energy consumption or continuous GPU clock behavior. A game reaching 200 pieces survived **at least** 200; its eventual death | |
| was not observed. Always report cap-hit rates with survival statistics. | |
| For the primary reliability-adjusted analysis, any errored episode receives zero | |
| lines, score and pieces. Its measured partial outcome remains visible separately. | |
| Dropping failures would bias the result toward easier successful games. Both base | |
| and trained must have zero episode errors for the default positive-signal flag. | |
| For each seed, subtract base lines from trained lines, then bootstrap these paired | |
| differences 10,000 times with RNG seed 2026. Report the mean and percentile 95% | |
| interval. Resampling complete episode pairs respects the fact that moves within | |
| one game are correlated. Treating every move as independent would exaggerate the | |
| amount of evidence. | |
| A lower interval bound above zero supports improved mean lines for this checkpoint | |
| and seed distribution. It does not establish competitive Tetris ability, training | |
| seed stability, or superiority over the heuristic. Report the trained-versus- | |
| heuristic comparison even if it is embarrassing. A reproducible negative result | |
| with clear limitations is a stronger portfolio artifact than a selective win. | |
| After test results are opened, do not use them to choose between epochs or change | |
| the observation format. A subsequent experiment needs a new preregistered study | |
| and fresh held-out sequences. The current report should preserve failures and | |
| limitations alongside the playable recorded comparison. | |
| ## Completed outcome and what it teaches | |
| All five players completed the 200 paired seeds at the declared 200-piece cap: | |
| | Player | Mean lines | Median lines | Mean pieces | Cap hits | | |
| | --- | ---: | ---: | ---: | ---: | | |
| | Native base | 0.070 | 0 | 26.020 | 0/200 | | |
| | Unchanged FP32 head | 0.105 | 0 | 26.000 | 0/200 | | |
| | Selected epoch 02 | 16.810 | 16 | 83.565 | 0/200 | | |
| | Random | 0.140 | 0 | 26.645 | 0/200 | | |
| | Fixed heuristic | 76.530 | 77 | 200.000 | 200/200 | | |
| There were zero inference errors and zero invalid decisions. The paired mean line | |
| gain over native base was **16.74 [15.74, 17.77]** (95% percentile bootstrap, | |
| 10,000 replicates, seed 2026). Against the unchanged FP32 head, the gain was | |
| **16.705 [15.705, 17.74]**, so the observed improvement cannot be attributed to head | |
| precision alone. Against the heuristic the difference was | |
| **−59.72 [−60.805, −58.619875]**. Both the improvement and this large shortfall belong | |
| in the result. The heuristic's eventual survival remains unknown because every | |
| one of its games reached the cap. | |
| The selected model's mean decision latency was **206.8 ms**, compared with the | |
| heuristic's **0.1985 ms**, about 1,042 times slower on these measured workloads. | |
| Policies visited different states, and neural inference used the GPU while the | |
| heuristic used the CPU. This is an operational comparison, not an isolated | |
| same-input speed benchmark. The final wrapper took 85.1 minutes, including model | |
| loading, report writing and service restoration. Energy and monetary costs were | |
| not measured. The simple heuristic remains the practical choice for this game. | |
| An independent CPU audit replayed all 1,000 saved episodes and independently | |
| recomputed every paired interval. See [its results](../../reports/final-audit.json). | |
| The learned model passes the narrow improvement test over its unchanged base; | |
| that does not turn its 42.79% teacher agreement into a general measure of game | |
| quality. The final report includes score/survival intervals, all latency summaries, | |
| hardware, source identities and portable release evidence paths. | |
| To replicate the analysis, first verify the report/checkpoint/dataset hashes in | |
| [final-summary.json](../../reports/final-summary.json), then use the frozen source | |
| and commands above in fresh output directories. Reproducing inference needs the | |
| GPU runtime; inspecting metrics and replaying saved actions do not. Reproduction | |
| uses the already disclosed sequences, so it is not a new blind test. Any follow-up | |
| training or design change needs a separate declared experiment and fresh test | |
| seeds. No post-test tuning was used to produce the results reported here. | |