nima1's picture
Use Nima Karimi as technical report author
292ce27 verified
|
Raw History Blame Contribute Delete
9.68 kB
metadata
license: apache-2.0
base_model: Cloudflare/clef-flash
base_model_relation: adapter
library_name: stackcraft
language:
  - en
tags:
  - clef
  - decision-model
  - lora
  - imitation-learning
  - stackcraft

Stackcraft Clef-Flash

Technical report

Read the technical report (PDF) · Read online · Source and reproducible build

Stackcraft: Adapting Clef-flash to a Falling-Block Decision Task — Nima Karimi, 6 October 2026. A seven-page technical report with methods, measured results, figures, limitations and AI-assistance disclosure. Not peer reviewed. The paper reports 16.81 mean lines for the trained policy versus 0.07 for native base and 76.53 for the heuristic. It is an additive document release; the model weights, dataset and original experimental evidence are unchanged.

Selected checkpoint: epoch-02, chosen on validation before final tests. This is a small, synthetic imitation-learning study of a falling-block placement policy. It is not a generally capable game agent, reinforcement-learning result, or competitive Tetris implementation.

The artifact contains a rank-4 LoRA adapter and a separately trained FP32 native joint decision head for Cloudflare/clef-flash, pinned revision 17f0b0ad64efb65d273590632833508766b2aae6. Load both. The frozen BF16 backbone is downloaded separately; these files are not a standalone model. Clef-flash is based on Qwen3.5-9B. The native schema head scores every supplied legal placement in one forward pass, with no generated text or command parsing.

Related releases: source and tutorials, model, dataset, and local playable demo. The owner deferred hosted deployment. The game and Docker image need no GPU or Hugging Face subscription when run locally.

Intended use

Learn how to build deterministic environments, generate search-teacher labels, train a custom decision model, preserve its native head, and evaluate complete games with paired uncertainty. The playable demo uses recorded model replays; it does not provide public live GPU inference.

The observation is a 10×20 occupied-cell grid, current tetromino, exactly one next piece, rules and every legal placement with its coordinates. The action selects an orientation and column, then drops vertically. There is no hold, wall kick, tuck, gravity timer, T-spin bonus or changing speed. Hidden future pieces and sequence seeds are excluded from neural input. Probabilities reflect the model's choice distribution; they are not calibrated chances of clearing a line or winning.

Training and data

The study uses 827 training and 215 validation positions generated from fixed seven-bag streams. Random, heuristic and search policies collect states. Every label comes from a fixed two-placement search using the current piece and one preview, with a board-height/holes/bumpiness/line-clear heuristic. These are synthetic teacher preferences, not human demonstrations or globally optimal moves. Exact duplicate observations are removed across splits. Six source hashes and exact file hashes accompany the dataset manifest.

Preregistered training: seed 42, two epochs, batch 1, accumulation 8, rank 4/alpha 8, no LoRA dropout, AdamW lr 1e-5/weight_decay 0.01, gradient norm cap 1.0, label smoothing 0.05 and Brier sum weight0.1. BF16 backbone, FP32 head, text-layer LoRA including both hybrid attention types and MLPs. Vision, output and other original parameters remain frozen. There are 132,582,404 trainable parameters: 121,762,820 in the joint head and 10,819,584 in LoRA. This is why the head is a substantial part of the saved artifact, even though most backbone weights stay frozen. Select only between epoch 01/02 using lowest finite target NLL on all 215 validation positions; exact ties choose the earlier epoch.

Evaluation and limitations

Final preregistered comparison: 200 untouched seeds 30000–30199, 200-piece cap, selected model versus unchanged native BF16-head Clef, unchanged FP32-head ablation, random legal placement and a cheap heuristic. Primary metric is lines cleared. Report score, survival, cap hits, failures and latency. Paired 95% bootstrap intervals use 10000 episode resamples with seed 2026. Failed episodes remain in the analysis, assigned zero primary outcomes with partial outcomes reported separately.

Verified final outcomes on all 200 paired seeds, with a 200-piece cap. Lines, score and placed-piece summaries use the preregistered zero-on-error policy. Cap hits indicate censored survival.

Player Lines mean (median) Score mean (median) Placed mean (median) Cap hits Errors
base 0.070 (0.0) 7.000 (0.0) 26.020 (26.0) 0.0% 0.0%
base-fp32 0.105 (0.0) 10.500 (0.0) 26.000 (26.0) 0.0% 0.0%
trained 16.810 (16.0) 1745.500 (1650.0) 83.565 (82.0) 0.0% 0.0%
random 0.140 (0.0) 14.000 (0.0) 26.645 (27.0) 0.0% 0.0%
heuristic 76.530 (77.0) 8306.500 (8300.0) 200.000 (200.0) 100.0% 0.0%

Each paired 95% bootstrap interval resamples complete episode differences. Trained versus native base is the primary comparison; the heuristic and FP32-head comparisons are separate checks. Intervals are not adjusted for multiple comparisons.

Comparison (first minus second) Mean lines difference Paired 95% interval
trained − base 16.740 [15.740, 17.770]
trained − heuristic -59.720 [-60.805, -58.620]
trained − base-fp32 16.705 [15.705, 17.740]
base-fp32 − base 0.035 [-0.020, 0.090]

Decision latency includes recorded first-call effects, tokenization and policy overhead, but excludes game rendering and replay playback.

Player Decision mean ms Median ms p95 ms Invalid decisions
base 149.110 139.381 244.916 0
base-fp32 150.188 140.968 247.880 0
trained 206.814 171.844 309.000 0
random 0.001 0.001 0.003 0
heuristic 0.199 0.139 0.281 0

The full reports retain raw outcomes, all failed seeds and probability logs.

Fine-tuning improved complete-game play over both unchanged Clef references in this study. The fixed heuristic still cleared substantially more lines and used about one thousandth of the trained model's decision time on this workstation. All five policies had zero errors and invalid decisions. Every heuristic game hit the 200-piece cap; no neural or random game did. This is evidence of task adaptation, not a reason to prefer this 9B model over the heuristic for the simplified game. See the readable study report and the hardware/runtime snapshot.

Only one training seed, one small synthetic dataset, one simplified game ruleset and one hardware/software configuration are studied. The finite cap censors survival. Validation teacher agreement does not establish game quality. Test uncertainty does not capture variation across training seeds.

Loading

Use the released Stackcraft source and locked ML dependencies. Stackcraft uses PEFT internally but requires its custom native-head loader. Generic AutoModel or adapter-only loading omits the custom decision head and is incorrect. From the downloaded model repository's code/ directory, install the locked ML environment with uv sync --locked --extra ml and cache the pinned upstream snapshot as described in Tutorial 02. Run this example with uv run --locked --extra ml python; the sibling ../checkpoint directory contains the released adapter and head:

from stackcraft.clef import ClefPlayer
from stackcraft.training import load_checkpoint
from stackcraft.engine import new_game
from stackcraft.players import observe

player = ClefPlayer.from_pretrained(trust_pinned_code=True, local_files_only=True)
load_checkpoint(player.model, "../checkpoint")
print(player.choose(observe(new_game(42))))

trust_pinned_code=True executes reviewed, hash-checked upstream Python. It is not a sandbox. The measured RTX 5090 inference baseline allocated 19.7 GB; bounded LoRA training allocated 23.2 GB. Load only with sufficient free VRAM. Context is limited to 4096 tokens, with explicit rejection before any state truncation.

See the released tutorials, training configuration, artifact hashes and evaluation report for exact reproduction. Apache-2.0; see LICENSE and NOTICE for upstream attribution. This project is independent of Cloudflare, Qwen and Tetris.

Bundle layout

checkpoint/ preserves the selected checkpoint bytes. code/ contains the source, locked environment, tests, scripts and milestone tutorials. evidence.zip losslessly compresses the frozen selection, both validation candidates and complete final evaluation. From the model repository root, run python -m zipfile -e evidence.zip . after download verification to restore evidence/ and all original relative paths. evidence-files.json records each raw file's byte count and SHA256; extraction preserves the original evidence hashes. From code/, run uv sync --locked --extra ml; load ../checkpoint using the native Stackcraft loader. No access to a private GitHub repository is required.