|
Download README.md from nima1/stackcraft-clef-flash-lora: direct link, hf CLI and curl.
- Browser
- Download file 9.68 kB
-
https://huggingface.co/nima1/stackcraft-clef-flash-lora/resolve/main/README.md
- Command line
-
hf download hf://nima1/stackcraft-clef-flash-lora/README.md
-
curl -L -o README.md https://huggingface.co/nima1/stackcraft-clef-flash-lora/resolve/main/README.md
9.68 kB
| license: apache-2.0 | |
| base_model: Cloudflare/clef-flash | |
| base_model_relation: adapter | |
| library_name: stackcraft | |
| language: | |
| - en | |
| tags: | |
| - clef | |
| - decision-model | |
| - lora | |
| - imitation-learning | |
| - stackcraft | |
| # Stackcraft Clef-Flash | |
| ## Technical report | |
| **[Read the technical report (PDF)](paper/stackcraft-technical-report.pdf)** · | |
| [Read online](paper/stackcraft-technical-report.md) · | |
| [Source and reproducible build](https://github.com/kkarimi/stackcraft/tree/4f09178ce6b3dfca0844e240fc8e42daee33b5da/paper) | |
| *Stackcraft: Adapting Clef-flash to a Falling-Block Decision Task* — Nima Karimi, | |
| 6 October 2026. A seven-page technical report with methods, measured results, | |
| figures, limitations and AI-assistance disclosure. **Not peer reviewed.** | |
| The paper reports 16.81 mean lines for the trained policy versus 0.07 for native | |
| base and 76.53 for the heuristic. It is an additive document release; the model | |
| weights, dataset and original experimental evidence are unchanged. | |
| Selected checkpoint: **epoch-02**, chosen on validation before final tests. | |
| This is a small, synthetic imitation-learning study of a falling-block placement | |
| policy. It is not a generally capable game agent, reinforcement-learning result, | |
| or competitive Tetris implementation. | |
| The artifact contains a rank-4 LoRA adapter **and a separately trained FP32 native | |
| joint decision head** for `Cloudflare/clef-flash`, pinned revision | |
| `17f0b0ad64efb65d273590632833508766b2aae6`. Load both. The frozen BF16 backbone | |
| is downloaded separately; these files are not a standalone model. Clef-flash is | |
| based on Qwen3.5-9B. The native schema head scores every supplied legal placement | |
| in one forward pass, with no generated text or command parsing. | |
| Related releases: [source and tutorials](https://github.com/kkarimi/stackcraft), | |
| [model](https://huggingface.co/nima1/stackcraft-clef-flash-lora), | |
| [dataset](https://huggingface.co/datasets/nima1/stackcraft-data), and | |
| [local playable demo](https://github.com/kkarimi/stackcraft/blob/main/docs/demo-hosting.md). | |
| The owner deferred hosted deployment. The game and Docker image need no GPU | |
| or Hugging Face subscription when run locally. | |
| ## Intended use | |
| Learn how to build deterministic environments, generate search-teacher labels, | |
| train a custom decision model, preserve its native head, and evaluate complete | |
| games with paired uncertainty. The playable demo uses recorded model replays; | |
| it does not provide public live GPU inference. | |
| The observation is a 10×20 occupied-cell grid, current tetromino, exactly one next | |
| piece, rules and every legal placement with its coordinates. The action selects | |
| an orientation and column, then drops vertically. There is no hold, wall kick, | |
| tuck, gravity timer, T-spin bonus or changing speed. Hidden future pieces and | |
| sequence seeds are excluded from neural input. Probabilities reflect the model's | |
| choice distribution; they are not calibrated chances of clearing a line or winning. | |
| ## Training and data | |
| The study uses 827 training and 215 validation positions generated from fixed | |
| seven-bag streams. Random, heuristic and search policies collect states. Every | |
| label comes from a fixed two-placement search using the current piece and one | |
| preview, with a board-height/holes/bumpiness/line-clear heuristic. These are | |
| synthetic teacher preferences, not human demonstrations or globally optimal moves. | |
| Exact duplicate observations are removed across splits. Six source hashes and | |
| exact file hashes accompany the dataset manifest. | |
| Preregistered training: seed 42, two epochs, batch 1, accumulation 8, rank 4/alpha 8, | |
| no LoRA dropout, AdamW lr 1e-5/weight_decay 0.01, gradient norm cap 1.0, label smoothing | |
| 0.05 and Brier sum weight0.1. BF16 backbone, FP32 head, text-layer LoRA including | |
| both hybrid attention types and MLPs. Vision, output and other original parameters | |
| remain frozen. There are 132,582,404 trainable parameters: 121,762,820 in the | |
| joint head and 10,819,584 in LoRA. This is why the head is a substantial part of | |
| the saved artifact, even though most backbone weights stay frozen. Select only between epoch 01/02 using lowest finite target NLL on | |
| all 215 validation positions; exact ties choose the earlier epoch. | |
| ## Evaluation and limitations | |
| Final preregistered comparison: 200 untouched seeds 30000–30199, 200-piece cap, | |
| selected model versus unchanged native BF16-head Clef, unchanged FP32-head | |
| ablation, random legal placement and a cheap heuristic. Primary metric is lines | |
| cleared. Report score, survival, cap hits, failures and latency. Paired 95% bootstrap | |
| intervals use 10000 episode resamples with seed 2026. Failed episodes remain in the | |
| analysis, assigned zero primary outcomes with partial outcomes reported separately. | |
| Verified final outcomes on all 200 paired seeds, with a 200-piece cap. Lines, score and placed-piece summaries use the preregistered zero-on-error policy. Cap hits indicate censored survival. | |
| | Player | Lines mean (median) | Score mean (median) | Placed mean (median) | Cap hits | Errors | | |
| | --- | ---: | ---: | ---: | ---: | ---: | | |
| | base | 0.070 (0.0) | 7.000 (0.0) | 26.020 (26.0) | 0.0% | 0.0% | | |
| | base-fp32 | 0.105 (0.0) | 10.500 (0.0) | 26.000 (26.0) | 0.0% | 0.0% | | |
| | trained | 16.810 (16.0) | 1745.500 (1650.0) | 83.565 (82.0) | 0.0% | 0.0% | | |
| | random | 0.140 (0.0) | 14.000 (0.0) | 26.645 (27.0) | 0.0% | 0.0% | | |
| | heuristic | 76.530 (77.0) | 8306.500 (8300.0) | 200.000 (200.0) | 100.0% | 0.0% | | |
| Each paired 95% bootstrap interval resamples complete episode differences. Trained versus native base is the primary comparison; the heuristic and FP32-head comparisons are separate checks. Intervals are not adjusted for multiple comparisons. | |
| | Comparison (first minus second) | Mean lines difference | Paired 95% interval | | |
| | --- | ---: | ---: | | |
| | trained − base | 16.740 | [15.740, 17.770] | | |
| | trained − heuristic | -59.720 | [-60.805, -58.620] | | |
| | trained − base-fp32 | 16.705 | [15.705, 17.740] | | |
| | base-fp32 − base | 0.035 | [-0.020, 0.090] | | |
| Decision latency includes recorded first-call effects, tokenization and policy overhead, but excludes game rendering and replay playback. | |
| | Player | Decision mean ms | Median ms | p95 ms | Invalid decisions | | |
| | --- | ---: | ---: | ---: | ---: | | |
| | base | 149.110 | 139.381 | 244.916 | 0 | | |
| | base-fp32 | 150.188 | 140.968 | 247.880 | 0 | | |
| | trained | 206.814 | 171.844 | 309.000 | 0 | | |
| | random | 0.001 | 0.001 | 0.003 | 0 | | |
| | heuristic | 0.199 | 0.139 | 0.281 | 0 | | |
| The full reports retain raw outcomes, all failed seeds and probability logs. | |
| Fine-tuning improved complete-game play over both unchanged Clef references in | |
| this study. The fixed heuristic still cleared substantially more lines and used | |
| about one thousandth of the trained model's decision time on this workstation. | |
| All five policies had zero errors and invalid decisions. Every heuristic game hit | |
| the 200-piece cap; no neural or random game did. This is evidence of task adaptation, | |
| not a reason to prefer this 9B model over the heuristic for the simplified game. | |
| See [the readable study report](code/reports/final-study.md) and | |
| [the hardware/runtime snapshot](code/reports/evaluation-hardware.json). | |
| Only one training seed, one small synthetic dataset, one simplified game ruleset | |
| and one hardware/software configuration are studied. The finite cap censors | |
| survival. Validation teacher agreement does not establish game quality. Test | |
| uncertainty does not capture variation across training seeds. | |
| ## Loading | |
| Use the released Stackcraft source and locked ML dependencies. Stackcraft uses | |
| PEFT internally but requires its custom native-head loader. Generic | |
| `AutoModel` or adapter-only loading omits the custom decision head and is incorrect. | |
| From the downloaded model repository's `code/` directory, install the locked ML | |
| environment with `uv sync --locked --extra ml` and cache the pinned upstream | |
| snapshot as described in [Tutorial 02](code/docs/tutorials/02-baselines.md). | |
| Run this example with `uv run --locked --extra ml python`; the sibling | |
| `../checkpoint` directory contains the released adapter and head: | |
| ```python | |
| from stackcraft.clef import ClefPlayer | |
| from stackcraft.training import load_checkpoint | |
| from stackcraft.engine import new_game | |
| from stackcraft.players import observe | |
| player = ClefPlayer.from_pretrained(trust_pinned_code=True, local_files_only=True) | |
| load_checkpoint(player.model, "../checkpoint") | |
| print(player.choose(observe(new_game(42)))) | |
| ``` | |
| `trust_pinned_code=True` executes reviewed, hash-checked upstream Python. It is | |
| not a sandbox. The measured RTX 5090 inference baseline allocated 19.7 GB; bounded | |
| LoRA training allocated 23.2 GB. Load only with sufficient free VRAM. Context is | |
| limited to 4096 tokens, with explicit rejection before any state truncation. | |
| See the released tutorials, training configuration, artifact hashes and evaluation | |
| report for exact reproduction. Apache-2.0; see LICENSE and NOTICE for upstream | |
| attribution. This project is independent of Cloudflare, Qwen and Tetris. | |
| ## Bundle layout | |
| `checkpoint/` preserves the selected checkpoint bytes. `code/` contains the source, locked environment, tests, scripts and milestone tutorials. `evidence.zip` losslessly compresses the frozen selection, both validation candidates and complete final evaluation. From the model repository root, run `python -m zipfile -e evidence.zip .` after download verification to restore `evidence/` and all original relative paths. `evidence-files.json` records each raw file's byte count and SHA256; extraction preserves the original evidence hashes. From `code/`, run `uv sync --locked --extra ml`; load `../checkpoint` using the native Stackcraft loader. No access to a private GitHub repository is required. | |