File size: 19,819 Bytes
5b4da6c
 
292ce27
5b4da6c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
# Stackcraft: Adapting Clef-flash to a Falling-Block Decision Task

Nima Karimi · 2026-10-06 · Version 1.0

**Technical report · Not peer reviewed**


## Abstract

Can a general decision model learn a useful game policy from a small synthetic
dataset on one workstation? Stackcraft studies this question in a deterministic,
turn-based falling-block game. A Clef-flash model was adapted using rank-4 LoRA
and its full native decision head, with 827 training positions and 215 validation
positions labeled by a bounded search teacher. One training run produced two
checkpoint candidates; validation negative log likelihood selected the second.
Across 200 previously reserved piece sequences per policy, the selected model
cleared 16.81 lines per game, versus 0.07 for unchanged Clef-flash. The paired
improvement was 16.74 lines, with a 95% bootstrap interval of [15.74, 17.77].
However, a fixed arithmetic heuristic cleared 76.53 lines and was substantially
faster. All 1,000 evaluated episodes completed without inference errors or invalid
decisions. The study demonstrates task adaptation under a constrained local
budget, while showing why improvement over a weak base model is insufficient
evidence of practical superiority. Code, data, learned parameters, recorded games,
and verification evidence are public. This is a technical report, not a
peer-reviewed publication.

## 1. Motivation and scope

A falling-block game offers a concrete test of decision quality: choices alter
the board, mistakes accumulate, and complete games expose failures that isolated
classification scores can hide. Stackcraft asks whether a small imitation-learning
dataset can improve a general decision model, and whether that improvement is
competitive with a simple task-specific policy. The project also makes the
experiment inspectable through replayable games and milestone tutorials.

Clef-flash is a roughly nine-billion-parameter model derived from Qwen3.5-9B.
Its native interface accepts a state and typed questions, then produces
probabilities over supplied options through a joint decision head. This avoids
requiring a generated command or a text parser to select a legal move. Stackcraft
uses text and structured state only; it does not test the model's visual
capabilities. The implementation follows the pinned upstream interface. [[1]](https://huggingface.co/Cloudflare/clef-flash/blob/17f0b0ad64efb65d273590632833508766b2aae6/README.md)

The contribution is an engineering case study with released evidence. No new
adaptation algorithm, general game-playing advance, or superiority over
specialized game agents is claimed. The central distinction is between learning
to play better than the unchanged model and building the best bot for these rules.

## 2. Game and policy interface

The board has ten columns and twenty rows. Seven tetromino shapes arrive in
seeded seven-bag sequences: each shuffled bag contains every shape once. A turn
chooses a legal orientation and column, then drops the piece vertically. The
rules omit hold, wall kicks, tucks, T-spin bonuses, and real-time movement
requirements. Clearing one, two, three, or four lines awards 100, 300, 500, or
800 points. Lines cleared per episode is the primary outcome.

Every policy receives the board, current piece, one next-piece preview, and the
same legal placement options. Hidden future pieces, the episode seed, sequence
index, and generator state are excluded from the observation. Python owns the
authoritative transitions, so browser play, offline tournaments, and replay
validation use the same rules. This is a structured placement task rather than a
benchmark of full competitive Tetris.

Five policies were evaluated. **Native base** retains unchanged Clef-flash weights
and its original BF16 head. **FP32 base** uses unchanged weights with the FP32 head
wrapper used for training. **Trained** combines the selected LoRA adapter with the
learned FP32 head. All three share the pinned backbone, observation encoding,
option ordering, BF16 backbone precision, and 4,096-token limit. Full observations
are checked against that limit before inference.

**Random** samples uniformly from legal options using a recorded random generator.
**Heuristic** greedily evaluates the resulting board with a fixed formula:
negative aggregate column height, minus four times the number of holes, minus
adjacent-column height differences, plus eight times the number of cleared lines.
A hole is an empty cell below an occupied cell in its column. The heuristic does
not use the available preview; its weights were frozen before final testing. It
is distinct from the lookahead teacher that generated training labels. [[3]](https://github.com/kkarimi/stackcraft/blob/82a8854258f228a214f0964bbbd3d2b52e2be1e5/reports/final-study.md)

## 3. Synthetic data and adaptation

### Data construction

The dataset contains 827 training positions from seeds 10000–10023 and 215
validation positions from seeds 20000–20005. Collection cycles among random,
greedy-heuristic, and search-expert behavior, recording at most forty moves per
episode. This exposes the learner to boards produced by different policies rather
than only successful expert trajectories. Every episode belongs to one split.
Occupancy-normalized duplicate detection removed six repeated training positions
and two validation positions that overlapped the training split.

For each retained observation, the teacher enumerates legal placements of the
current piece and the one visible preview, then evaluates line clears and board
quality. It cannot inspect unseen pieces. Its chosen action is a deterministic,
bounded-search recommendation, not a globally optimal label. The released dataset
records legal choices, teacher values, labels, source identity, and split
provenance. It is programmatically generated synthetic data; it is not a corpus
of human gameplay or human-certified optimal decisions. The reserved test seeds
30000–30199 were excluded from collection and checkpoint selection. [[4]](https://huggingface.co/datasets/nima1/stackcraft-data/tree/211708dd56f1a8af711c060c2ab1ecab35e7166d)

### Training configuration

LoRA adds trainable low-rank weight updates while retaining frozen backbone
parameters. [[2]](https://arxiv.org/abs/2106.09685) Here it was applied to the text backbone alongside full training
of Clef's native joint head. A single seed-42 run used rank 4, alpha 8, zero
adapter dropout, a BF16 unquantized backbone, and an FP32 head. The combined
trainable parameter count was 132,582,404. This count includes the full head;
“LoRA” should not be read as implying that only a tiny adapter was trained.

Training ran for exactly two epochs. The objective combined cross-entropy with
0.05 label smoothing and a Brier-loss term weighted by 0.1. AdamW used a learning
rate of 0.00001 and weight decay 0.01. Batch size was one, gradient accumulation
was eight, and gradient norm clipping was 1.0. Each epoch processed all 827
positions in 104 optimizer steps. The final accumulation group used its actual
three examples for normalization. Recorded hashes confirmed that frozen
parameters remained unchanged while intended trainable parameters changed.

### Selection and reload checks

The two epoch checkpoints were the entire preregistered selection search.
Eligibility required complete, finite predictions on all 215 validation positions;
selection minimized mean target negative log likelihood (NLL), with exact ties
going to the earlier epoch. Both candidates qualified, and epoch 02 was selected.
Teacher agreement and Brier scores were diagnostics, not alternative selection
criteria. The local preregistration was preserved in source; it was not a signed
external registration.

| Condition | Teacher agreement | Mean NLL | Mean Brier |
| --- | --- | --- | --- |
| Native base | 21/215 (9.77%) | 2.952702 | 0.923173 |
| FP32 base | 22/215 (10.23%) | 2.952393 | 0.923102 |
| Epoch 01 | 89/215 (41.40%) | 1.860870 | 0.722190 |
| Epoch 02 (selected) | 92/215 (42.79%) | 1.775923 | 0.708980 |

*Table 1. Validation results on all 215 positions. Lower NLL and Brier values are
better. The second epoch was selected before test games were opened.*

The selected model agreed with 92 teacher actions, or 42.79%. That percentage is
not a game-quality score. Alternative moves may be useful, teacher tie-breaking
can affect agreement, and a wrong move changes the states encountered later.
Complete held-out games therefore provide the behavioral evaluation.

The checkpoint saves both the adapter and native head. A fresh-process reload
matched four fixed training-position reference distributions exactly, with a
maximum absolute probability difference of 0.0 against a declared tolerance of
0.0001. The same four-reference check later passed using anonymously downloaded
published code and weights. These checks establish serialization consistency on
those references, not correctness on every board. [[3]](https://github.com/kkarimi/stackcraft/blob/82a8854258f228a214f0964bbbd3d2b52e2be1e5/reports/final-study.md), [[5]](https://github.com/kkarimi/stackcraft/blob/ad348c20820061b130abfc2890ab8de00b62e7e4/reports/publication.md)

## 4. Held-out evaluation

Each policy played the same 200 held-out sequences, seeds 30000–30199, with a cap
of 200 placed pieces per episode. The tournament therefore contains 1,000
episodes. A cap hit means survival for at least 200 pieces; it is not a loss on
piece 200. No model, prompt, policy, dataset, or checkpoint was tuned after opening
this final test.

Comparisons subtract outcomes within each seed. The analysis resamples the 200
paired episode differences with replacement 10,000 times, using bootstrap seed
2026, and takes percentile 95% intervals. Resampling episodes retains dependence
among decisions in the same game. The primary contrast is trained minus native
base in lines cleared. Score and survival intervals are descriptive secondary
results without a multiplicity correction.

| Player | Mean lines | Median lines | Mean score | Mean pieces | Cap hits |
| --- | --- | --- | --- | --- | --- |
| Native base | 0.070 | 0 | 7.0 | 26.020 | 0/200 |
| FP32 base | 0.105 | 0 | 10.5 | 26.000 | 0/200 |
| Trained | 16.810 | 16 | 1745.5 | 83.565 | 0/200 |
| Random | 0.140 | 0 | 14.0 | 26.645 | 0/200 |
| Heuristic | 76.530 | 77 | 8306.5 | 200.000 | 200/200 |

*Table 2. Held-out outcomes for 200 episodes per policy. Survival is pieces
placed. All five policies recorded zero inference errors and invalid decisions.*

![Held-out line outcomes and paired differences](figures/outcomes.svg)

*Figure 1. Mean lines cleared and paired differences for the selected checkpoint.
Intervals describe variability across held-out sequences for this checkpoint,
not variation across independently trained models.*

The trained model cleared 16.81 lines on average, compared with 0.07 for native
base: a gain of 16.74 lines, 95% interval [15.74, 17.77]. The gain against FP32 base
was 16.705 [15.705, 17.740]. The FP32-minus-native difference was only 0.035
[−0.020, 0.090], so changing head precision alone does not explain the observed
adaptation gain. This comparison does not separate learned adapter effects from
learned head effects; both were optimized together.

The heuristic remained substantially stronger, clearing 76.53 lines on average.
The trained-minus-heuristic difference was −59.720 [−60.805, −58.619875]. Every
heuristic episode reached the cap; no other policy did. Its unrestricted survival
is consequently unknown. No failed episodes were removed. The frozen analysis
would assign zero lines, score, and survival to an errored episode while retaining
its partial outcome separately; because there were no errors, adjusted and
observed results coincide.

An independently implemented audit reconstructed all 1,000 saved episodes,
checked observations and legal choices, verified source and checkpoint bindings,
and recomputed all twelve paired intervals across lines, score, and survival.
It matched the published results. This was a separate code-path check by another
AI agent, not external scientific replication. [[3]](https://github.com/kkarimi/stackcraft/blob/82a8854258f228a214f0964bbbd3d2b52e2be1e5/reports/final-study.md)

## 5. Runtime and practical value

Experiments ran on an NVIDIA RTX 5090 with roughly 32 GB of VRAM, an AMD Ryzen 9
9950X3D, and about 123 GiB of host RAM. The recorded environment used Python
3.13.16, Torch 2.14.1+cu130, CUDA 13.0, Transformers 5.18.0, and PEFT 0.21.2.
Torch used eight CPU threads and TF32 was disabled.

| Player | Mean ms | Median ms | p95 ms | Decisions |
| --- | --- | --- | --- | --- |
| Native base | 149.109957 | 139.381274 | 244.916266 | 5,204 |
| FP32 base | 150.188366 | 140.968013 | 247.879591 | 5,200 |
| Trained | 206.814401 | 171.844287 | 309.000088 | 16,713 |
| Random | 0.001497 | 0.001250 | 0.002879 | 5,329 |
| Heuristic | 0.198510 | 0.139489 | 0.280528 | 40,000 |

*Table 3. Per-decision timings over each policy's actual tournament workload.
Initialization is excluded; no warmup calls within the tournament were removed.*

![Observed mean decision latency on a logarithmic scale](figures/latency.svg)

*Figure 2. Observed mean decision latency. Neural policies use the GPU; random and
heuristic policies use the CPU. Different policies visit different boards, so
this is not a controlled same-state speed benchmark.*

Decision timings include encoding and probability transfer, but exclude model
loading, rendering, and replay serialization. The trained mean was 206.814 ms,
versus approximately 0.198510 ms for the heuristic: about 1,042 times longer on
these workloads. The trained policy also saw longer inputs on average than either
unchanged model. These measurements support choosing the heuristic for a practical
bot under the tested rules; they do not isolate a hardware-independent inference
speed ratio.

Training took 1,511.25 seconds, or 25.2 minutes, including checkpoint and hash
work. Peak CUDA allocation was 23.23 GB (21.64 GiB), with 24.18 GB reserved.
The final evaluation wrapper took 5,103.11 seconds, or 85.1 minutes, including
loading, reporting, and service restoration checks. These durations exclude
previous development, validation, and data generation. Energy use and monetary
cost were not measured. [[3]](https://github.com/kkarimi/stackcraft/blob/82a8854258f228a214f0964bbbd3d2b52e2be1e5/reports/final-study.md)

## 6. Limitations and next experiments

The strongest limitation is the single training seed. The bootstrap intervals
measure variability across episode sequences for one selected model; they do not
quantify training instability. The dataset is small and synthetic, and its teacher
optimizes a short-horizon approximation. Collection stops after at most forty
moves, whereas test games can last two hundred. Errors may therefore lead the
policy into poorly represented states, although this study does not establish a
causal failure mechanism.

Only one structured encoding, one adaptation configuration, and two epoch
candidates were studied. The weak unchanged-model result characterizes that pinned
model under this interface, not every possible use of Clef-flash. The rules omit
many skills relevant to real-time Tetris. The heuristic's universal cap hits limit
claims about its full survival distribution. None of these findings establishes
performance on a different game, board representation, or deployment platform.

Useful follow-up studies could compare head-only training with joint head/LoRA
training, repeat training across seeds, and evaluate a compact game-specific
student against the same inexpensive heuristic. Gathering recovery states from
the learned policy could test whether broader state coverage improves complete
games. Those are proposed experiments, not measured improvements. Any follow-up
should declare its selection procedure and reserve fresh test sequences rather
than repeatedly optimizing against the published test set.

## 7. Reproducibility and availability

The public code repository includes the game, experiment scripts, tests, a locked
uv environment, and tutorials. The model release includes the adapter, learned
head, reproduction source, and a lossless evidence archive with 1,017 records.
The archive's inventory preserves uncompressed file hashes. Anonymous downloads
of the pinned model and dataset passed byte and schema verification; all archive
entries were read and hashed. The published reload check used downloaded source
and checkpoint files. [[5]](https://github.com/kkarimi/stackcraft/blob/ad348c20820061b130abfc2890ab8de00b62e7e4/reports/publication.md)

The frozen model release is
`c4272310bd6c63ee97a941255abf8f36ff162229`; the dataset release is
`211708dd56f1a8af711c060c2ab1ecab35e7166d`. The bundled release source is
`82a8854258f228a214f0964bbbd3d2b52e2be1e5`. These identify the experimental
artifacts, not the later additive PDF publication. The upstream Clef-flash
revision is `17f0b0ad64efb65d273590632833508766b2aae6`. [[1]](https://huggingface.co/Cloudflare/clef-flash/blob/17f0b0ad64efb65d273590632833508766b2aae6/README.md), [[4]](https://huggingface.co/datasets/nima1/stackcraft-data/tree/211708dd56f1a8af711c060c2ab1ecab35e7166d), [[6]](https://huggingface.co/nima1/stackcraft-clef-flash-lora/tree/c4272310bd6c63ee97a941255abf8f36ff162229)

Provenance distinguishes stages rather than assigning every result to the release
commit. Data generation used commit `70d84bd8dc5d4a60f3b96455a57d9e6f9416d109`.
Training recorded `d9fcd05c02e9327d4d791f2338725092cede8bdd` with a dirty working
tree and explicit source-file hashes; that commit alone is not the full training
source identity. Final evaluation used clean commit
`e524f0c15c2dd60ebfdc1374d66836c79e1e6dd6`. The retained reports and source hashes
are the authoritative record of these distinctions. [[3]](https://github.com/kkarimi/stackcraft/blob/82a8854258f228a214f0964bbbd3d2b52e2be1e5/reports/final-study.md)

The local browser demo pairs live human play with clearly labeled recordings from
the actual base, trained, and heuristic policies. Playback speed is not live
neural inference speed. Hosted CPU deployment is deferred; no public running
Space is claimed. Readers can inspect evidence and replay games without retraining
or renting a GPU.

## Assistance disclosure

AI coding agents assisted with implementation, experiment orchestration,
documentation, manuscript preparation, and computational review under the project
owner's direction. The separate audit used another agent and an independent
implementation, but was part of this project. Synthetic labels came from the
programmatic search teacher. No external peer review, independent laboratory
replication, or exhaustive human labeling review is claimed.

## References

1. Cloudflare. [Clef-Flash model card and native interface, pinned revision](https://huggingface.co/Cloudflare/clef-flash/blob/17f0b0ad64efb65d273590632833508766b2aae6/README.md).
2. Hu, E. J., et al. (2021). [LoRA: Low-Rank Adaptation of Large Language Models](https://arxiv.org/abs/2106.09685).
3. Stackcraft. [Frozen study report and linked experimental evidence](https://github.com/kkarimi/stackcraft/blob/82a8854258f228a214f0964bbbd3d2b52e2be1e5/reports/final-study.md).
4. Stackcraft. [Synthetic dataset and provenance, pinned release](https://huggingface.co/datasets/nima1/stackcraft-data/tree/211708dd56f1a8af711c060c2ab1ecab35e7166d).
5. Stackcraft. [Publication and fresh-download verification report](https://github.com/kkarimi/stackcraft/blob/ad348c20820061b130abfc2890ab8de00b62e7e4/reports/publication.md).
6. Stackcraft. [Adapter, native head, source, and evidence, pinned release](https://huggingface.co/nima1/stackcraft-clef-flash-lora/tree/c4272310bd6c63ee97a941255abf8f36ff162229).