|
Download code/docs/tutorials/02-baselines.md from nima1/stackcraft-clef-flash-lora: direct link, hf CLI and curl.
- Browser
- Download file 7.39 kB
-
https://huggingface.co/nima1/stackcraft-clef-flash-lora/resolve/main/code/docs/tutorials/02-baselines.md
- Command line
-
hf download hf://nima1/stackcraft-clef-flash-lora/code/docs/tutorials/02-baselines.md
-
curl -L -o 02-baselines.md https://huggingface.co/nima1/stackcraft-clef-flash-lora/resolve/main/code/docs/tutorials/02-baselines.md
7.39 kB
| # Tutorial 02 — Build a comparison before training | |
| This milestone establishes offline opponents and a replayable tournament. **Unchanged Clef now has a small real-GPU development benchmark.** These results cannot establish an advantage from fine-tuning. | |
| ## Give every player the same information | |
| `players.observe(state)` creates a frozen `Observation`: board, current piece, exactly | |
| one next piece, legal placements, and the rules version. It omits the episode seed, | |
| piece index, and random generator. A seed would reveal the entire future sequence; | |
| giving it to a search player would invalidate the intended experiment. The tournament | |
| retains the seed separately so another person can replay an episode. | |
| Players implement `choose(observation) -> Decision`. A decision identifies one legal | |
| placement and can provide a complete probability distribution. The runner rejects | |
| illegal moves, missing probabilities, nonfinite probabilities, and incorrect totals. | |
| An exception ends that player's episode as an explicit error. A random fallback would | |
| hide deployment failures and turn the measured player into a different policy. | |
| ## Start with two cheap opponents | |
| `RandomPlayer(seed=0)` samples uniformly from legal placements using its own random | |
| generator. Its seed is unrelated to the game's sequence seed. The tournament creates | |
| a fresh instance for each episode, so earlier game length cannot alter later choices. | |
| `HeuristicPlayer` scores every legal afterboard with these fixed, untuned weights: | |
| value = -aggregate_height - 4*holes - bumpiness + 8*cleared_lines | |
| Aggregate height sums the ten column heights. A hole is an empty cell below a filled | |
| cell in its column. Bumpiness sums differences between adjacent column heights. The | |
| line bonus uses lines cleared by this move. All features describe the board after | |
| simultaneous row removal. Ties follow the engine's rotation-then-column ordering. | |
| These intentionally simple integer weights are a declared baseline, not an optimized | |
| or competitive Tetris claim. The heuristic receives the preview but does not use it; | |
| the later search expert will be able to search one additional placement. | |
| We call the authoritative engine's `place` helper to obtain afterboards. Calling | |
| `step` on the full episode state would expose the hidden future sequence, and a | |
| separate simulator could disagree with the game's actual rules. | |
| ## Preregister the development pilot | |
| Before running any tournament, the pilot is fixed as: development seeds **0–19**, | |
| 100-piece cap, random RNG seed 0, the weights above, and two players (random and | |
| heuristic). Question: does the runner produce repeatable paired games, and does the | |
| simple board heuristic behave plausibly? This is engineering development, not a | |
| held-out model study. Do not put these seeds into the final test pool or use this | |
| pilot to claim model improvement. Stop after 40 episodes or a runner failure. | |
| Expected artifact: `runs/baselines-development.json`. Record mean/median lines, | |
| score, pieces placed, cap-hit and error rates, and decision latency. Acceptance is | |
| valid replays and no unexplained errors; poor heuristic play is reportable, not a | |
| reason to secretly select nicer seeds. Starting commit is | |
| `53db6bd0e7f42723bc63db7ad56c4ac139a04031`, plus the uncommitted baseline | |
| implementation recorded in Git at integration. | |
| ## Reproduce the pilot | |
| From the repository, after `uv sync --locked --group dev`: | |
| ```bash | |
| uv run --locked python - <<'PY' | |
| import json | |
| from pathlib import Path | |
| from stackcraft.players import HeuristicPlayer, RandomPlayer | |
| from stackcraft.tournament import run_tournament | |
| result = run_tournament({'random': RandomPlayer, 'heuristic': HeuristicPlayer}, range(20), 100) | |
| output = Path('runs/baselines-development.json') | |
| output.parent.mkdir(parents=True, exist_ok=True) | |
| output.write_text(json.dumps(result, indent=2) + '\n') | |
| print(json.dumps(result['summary'], indent=2)) | |
| PY | |
| ``` | |
| Each episode records visible observations, chosen action IDs, optional probabilities, | |
| the complete action replay, final outcomes, explicit errors, and decision times. | |
| `replay_states(episode['replay'])` reconstructs its final board independently. All | |
| players receive identical piece sequences for each seed, even when one dies earlier. | |
| The cap censors survival: 100 pieces means "survived at least 100", not "would die | |
| at 100". Report cap-hit rate alongside survival. Latency times only `choose`, not | |
| rendering, observation construction, engine transition, or replay validation. It | |
| includes Python overhead and is machine dependent. Game content is deterministic; | |
| latency is deliberately excluded when checking repeatability. | |
| ## Development pilot result | |
| The preregistered pilot ran on 2026-10-05. All 40 episode replays matched their | |
| recorded final outcomes; neither player produced an error. | |
| | Player | Mean lines | Median lines | Mean score | Mean pieces | Cap-hit rate | Mean decision time | | |
| | --- | ---: | ---: | ---: | ---: | ---: | ---: | | |
| | Random | 0.10 | 0 | 10 | 26.9 | 0% | 0.0016 ms | | |
| | Heuristic | 35.95 | 36 | 3,900 | 100 | 100% | 0.179 ms | | |
| Every heuristic episode reached the cap. This supports using a longer development | |
| cap to estimate the cost of final evaluation; it does not measure unlimited survival. | |
| The pilot ran in approximately 3.37 seconds including independent replay checks | |
| and writing JSON. Timing is illustrative local CPU execution, with no GPU model. | |
| Artifact SHA-256: | |
| `ed4973c69496cdc0ebb5720e3beaae621406988a6d298233c10ab962e1c499af`. | |
| Regeneration preserves game outcomes but changes timing fields and therefore the | |
| full artifact hash. | |
| ## Actual unchanged Clef development run | |
| The pinned native Clef-flash backbone and joint head ran on RTX 5090 in BF16. | |
| First cache the exact upstream snapshot from the repository root: | |
| ```bash | |
| uv run --locked --extra ml hf download Cloudflare/clef-flash \ | |
| --revision 17f0b0ad64efb65d273590632833508766b2aae6 | |
| ``` | |
| This installs the optional ML dependencies and downloads files; it does not load | |
| the model onto the GPU. The benchmark loader uses local files only and will fail | |
| on a fresh machine without this step. Once at least25GiB GPU memory is free, run: | |
| ```bash | |
| uv run --locked --extra ml python scripts/benchmark_clef.py \ | |
| --output runs/clef-development-repeat.json --episodes 2 --max-pieces 30 | |
| ``` | |
| Choose a new output path if it already exists. Do not stop other workloads without | |
| authorization. The loader explicitly trusts the pinned, hash-checked upstream | |
| Python implementation; see Tutorial04 for that trust boundary and memory checks. | |
| On development seeds0–1 (30-piece cap), Clef averaged0lines and27.5pieces; | |
| the heuristic averaged9lines and reached30pieces both times. No player errored. | |
| Clef's55decisions averaged144ms (p95 241ms), after a separate warmup. | |
| These two episodes establish feasibility, not statistical performance. Raw probabilities, | |
| input-token counts, pinned revisions and timing are retained in | |
| `runs/clef-development-v1.json`. Reference PyTorch attention kernels were used; | |
| faster optional kernels were not installed, so latency is implementation-specific. | |
| The GLM service was temporarily stopped under user authorization and restored; | |
| its in-container health endpoint returned `{"status":"ok"}`. | |
| Use fresh held-out episode seeds for the final study. A large neural model can lose | |
| to a few lines of board arithmetic; the report must make that outcome visible. | |