nima1's picture
Finalize release documentation and native loader verification
c427231 verified
|
Raw History Blame Contribute Delete
13.8 kB

Stackcraft: completed Clef fine-tuning study

Fine-tuning improved Clef's play in this small, turn-based block game, but a fixed arithmetic heuristic remained much stronger and cheaper. The selected model cleared 16.81 lines per game, compared with 0.07 for the unchanged native model and 76.53 for the heuristic. The paired trained-minus-native improvement was 16.74 lines, 95% bootstrap interval [15.74, 17.77].

These are results from all 200 held-out piece sequences, seeds 30000–30199, with a 200-piece cap for each of five players: 1,000 completed episodes. There were zero inference errors and zero invalid decisions. Training, selection, reload verification, final evaluation and an independent computational audit are complete. Publication and fresh-download verification are tracked separately in the source repository.

The compact machine-readable report contains full-precision summaries, all paired intervals, runtime configuration and evidence hashes.

What was tested

Stackcraft uses a 10×20 board, seven tetrominoes and deterministic seven-bag piece sequences. A turn selects a legal rotation-and-column placement that drops vertically. There is no hold, wall kick, tuck, T-spin scoring or real-time movement deadline. Scores are 100/300/500/800 for clearing one/two/three/four lines in a move. Each player receives the board, current piece, one next-piece preview and the same legal options. Seed, sequence index and future random state are excluded from the policy observation. This tests decisions over structured state, not visual play.

The neural players use Cloudflare/clef-flash, pinned to revision 17f0b0ad64efb65d273590632833508766b2aae6, with the same full observation encoding, option order, BF16 backbone and 4,096-token limit. The three neural conditions are:

  • Native base: unchanged weights and original BF16 decision head.
  • FP32 base: unchanged weights, using the FP32 head wrapper also used in training.
  • Trained: selected rank-4 LoRA adapters plus the learned FP32 native decision head.

Random chooses uniformly among legal options using its own recorded RNG. Heuristic greedily scores each resulting board with weights frozen before final evaluation: −aggregate height − 4 × holes − bumpiness + 8 × cleared lines. It does not use the available preview. The lookahead teacher used for labeling is distinct from this greedy heuristic and was not one of the five final tournament players.

Held-out game outcomes

Entries show mean / median across all 200 episodes for each player. Survival means pieces placed. A cap hit means survival of at least 200 pieces, not death at piece 200.

Player Lines Score Pieces survived Cap hits Errors / invalid decisions
Native base 0.070 / 0 7.0 / 0 26.020 / 26 0/200 0 / 0
FP32 base 0.105 / 0 10.5 / 0 26.000 / 26 0/200 0 / 0
Trained 16.810 / 16 1745.5 / 1650 83.565 / 82 0/200 0 / 0
Random 0.140 / 0 14.0 / 0 26.645 / 27 0/200 0 / 0
Heuristic 76.530 / 77 8306.5 / 8300 200.000 / 200 200/200 0 / 0

No failures were dropped. The frozen analysis assigns zero lines, score and pieces to any errored episode and preserves the observed partial outcome separately. Because none failed, failure-adjusted and observed summaries coincide here.

For each comparison, subtract the second player's outcome from the first on the same seed. Resample the 200 paired differences with replacement 10,000 times, using bootstrap seed 2026, and take the percentile 95% interval. Resampling whole episodes preserves dependence among moves within a game.

Comparison Mean line difference Paired 95% interval
Trained − native base (primary) +16.740 [15.740, 17.770]
Trained − FP32 base +16.705 [15.705, 17.740]
Trained − heuristic −59.720 [−60.805, −58.619875]
FP32 base − native base +0.035 [−0.020, 0.090]

The trained gain persists against the unchanged FP32 head. Head precision alone does not explain that gain in this experiment. The FP32-versus-native interval includes zero. This is not an ablation separating the effects of learned LoRA from learned head weights: both were trained together.

The primary trained-minus-native score difference was +1738.5 [1633.0, 1848.5], and survival difference was +57.545 pieces [54.930, 60.210]. All twelve intervals (lines, score and pieces for all four comparisons) are preserved in final-summary.json. Lines are the primary endpoint; score and survival intervals are descriptive secondary results, without a multiplicity correction. These intervals describe episode variability for one selected checkpoint, not variability across training runs.

Data, training and checkpoint selection

The disclosed synthetic dataset contains 827 training positions and 215 validation positions. Separate episode seed pools supplied mixed random, heuristic and lookahead-expert trajectories, sampled for at most 40 moves per episode. Occupancy-normalized deduplication removed six within-training duplicates and two cross-split duplicates. The teacher enumerates placements of the current piece and the single visible preview, scoring line clears and the resulting board; it cannot inspect unseen pieces. Its labels are bounded-search recommendations, not proven optimal moves. See the dataset manifest and Tutorial 03.

One seed-42 training run produced exactly two preregistered epoch candidates. Training used BF16 frozen backbone weights without quantization, rank-4/alpha-8 LoRA with zero dropout, and the full FP32 native joint head: 132,582,404 trainable parameters. The fixed loss was cross-entropy with 0.05 label smoothing plus Brier loss weighted 0.1. AdamW used learning rate 1e-5 and weight decay 0.01, with batch size one, accumulation eight and gradient clipping at 1.0. The final accumulation group was normalized by its actual three examples. Each epoch processed all 827 training positions in 104 optimizer steps. Frozen parameter hashes stayed unchanged.

Selection used the lowest finite mean target negative log likelihood (NLL) on all 215 validation positions, with exact ties assigned to the earlier epoch. Candidates with errors or incomplete probabilities were ineligible. Both were eligible; epoch 02 won. Test games did not choose the checkpoint.

Validation condition Teacher agreement Mean target NLL Mean Brier
Native base 21/215 (9.77%) 2.952702 0.923173
FP32 base 22/215 (10.23%) 2.952393 0.923102
Epoch 01 89/215 (41.40%) 1.860870 0.722190
Epoch 02, selected 92/215 (42.79%) 1.775923 0.708980

42.79% teacher agreement is not 42.79% game quality. It measures agreement with one chosen teacher action, including its deterministic tie rule. Small errors change subsequent states, and imitation can fail on boards absent from training. Complete games provide a separate behavioral measurement.

The selected checkpoint reloaded in a fresh process with 0.0 maximum absolute probability difference on four fixed training-position references, below the 1e-4 tolerance. This checks those reference distributions, not all possible boards or downloaded publication artifacts. Full details are in the training, validation and reload reports. The training report's status reflects its creation before external validation; later evidence completes that stage. The archived preregistration preserves the pre-test choices; it is a local source archive, not a signed external registration.

Runtime and practical cost

The workstation used an NVIDIA RTX 5090 (reported 32,607 MiB), AMD Ryzen 9 9950X3D, eight Torch CPU threads and about 123 GiB host RAM. Runtime was Python 3.13.16, Torch 2.14.1+cu130, CUDA 13.0, Transformers 5.18.0 and PEFT 0.21.2. TF32 was disabled. The hardware snapshot records the remaining details.

Player Mean decision ms Median ms p95 ms Decisions Local model initialization s
Native base 149.110 139.381 244.916 5,204 3.866
FP32 base 150.188 140.968 247.880 5,200 2.176
Trained 206.814 171.844 309.000 16,713 2.528
Random 0.001497 0.001250 0.002879 5,329 0.000035
Heuristic 0.198510 0.139489 0.280528 40,000 0.000002

Latencies cover all Player.choose calls, including encoding and probability transfer, with no excluded warmup. They exclude rendering, replay serialization and model loading. Initialization was measured separately with weights already cached locally; these are not download or reliably disk-cold timings. The neural players' mean input lengths were 1,354.30, 1,349.73 and 1,532.11 tokens respectively (native, FP32, trained); the trained p95 was 2,259 tokens.

The trained model took about 1,042 times the heuristic's mean decision time while clearing far fewer lines. This is an observed workload comparison: neural players use the GPU, the heuristic uses the CPU, and policies visit different boards. It is not a controlled same-state kernel benchmark. For a practical game bot under these rules, the heuristic is the clear choice. The fine-tuning result is useful as a reproducible learning experiment, not a reason to replace the heuristic in production.

Training took 1,511.25 seconds (25.2 minutes) including checkpoint and hash work; peak CUDA allocation was 23.23 GB (21.64 GiB) and reservation 24.18 GB. The final evaluation wrapper took 5,103.11 seconds (85.1 minutes) including loading, reporting and the local service stop/restore health checks. This was within the pre-test multi-hour budget. These durations exclude earlier development, validation and dataset generation. No energy consumption or monetary cost was measured; the GPU's configured 575 W ceiling is not an energy-use measurement.

A fixed replay illustration

Seed 30000, the first reserved seed and the declared demo sequence, provides an exploratory illustration. These values come from saved replays; no additional games were generated for this description.

Player Lines Pieces Final maximum height Final holes
Native base 0 23 20 86
FP32 base 0 26 20 79
Trained 16 82 20 28
Random 0 24 20 88
Heuristic 78 200 (cap) 3 1

A hole is an empty cell below an occupied cell in its column. These are different endpoints: the heuristic is still alive at the cap, whereas the other players have topped out. The trained player clears lines but eventually reaches the top. The table describes this replay; it does not establish why the model fails or prove a causal relationship between final holes and the aggregate result.

Reproducibility and remaining limits

The independently implemented CPU audit checked all 1,000 persisted episodes: replayed states, observations, choices, outcomes, source/checkpoint/selection bindings and aggregate statistics. It independently recomputed all twelve paired intervals and matched the report. The audit results and archived audit source preserve that check; its script hash and compact outcome are also included in the summary. This is an independent code-path check by another agent, not an external scientific replication.

The full evaluation JSON is identified by SHA-256 5585574b77e81c8d7dd44464c4dfb09f9511ccec0d9150c3ea44793760416857. The frozen selection SHA-256 is 025b62587a26bd8cc78c80f362524b8c178190df5188b317299121afa6b1203b. Dataset manifest, checkpoint files, evaluation source files and supporting report hashes are all included in final-summary.json.

The prepared model-release layout contains code/ with these reports, source, locked environment, tests and tutorials; evidence.zip losslessly preserves evidence/evaluation/ with the full report, request and per-player episode records; and evidence/selection/ with the frozen choice and both candidates' validation evidence. Extract the ZIP from the model repository root to restore these exact paths; evidence-files.json records their unchanged uncompressed hashes and sizes. These portable paths are relative to the model repository root and do not require access to private Git history. See the model release, dataset, and local demo instructions. The owner deferred hosted deployment; no live Space is claimed. Tutorial 05 gives the reproduction procedure.

The principal limitations are one training seed, a small synthetic dataset, an approximate teacher, a single structured encoding and simplified rules. Training positions cover shorter trajectories than the 200-piece test horizon. All 200 heuristic games reached the cap, so their uncapped longevity is unknown. The weak unchanged-model results characterize this pinned model and encoding, not every possible way to use Clef. No model, prompt, dataset, policy or checkpoint was tuned after opening this final test. Future improvements require a new declared study and fresh held-out sequences; these results remain the record of this experiment.