--- license: mit library_name: onnxruntime tags: - chess - onnx - gan - computer-vision --- # King's Vision — models ONNX artifacts for [King's Vision](https://github.com/RyanMatthew04/kings-vision): mate-in-one puzzle generators steered by a Cross-Entropy Method search, and a board square classifier that reads a chessboard image into a FEN. Every model runs on `onnxruntime` alone — no TensorFlow, no PyTorch. ```python from huggingface_hub import hf_hub_download import onnxruntime as ort path = hf_hub_download("FlappingChance/kings-vision-models", "queen_mate_generator.onnx") session = ort.InferenceSession(path) ``` The application fetches these automatically; see `src/kings_vision/artifacts.py`. --- ## mate-type generators and CEM agents Five DCGAN generators that synthesise chess checkmate positions, each paired with a multi-modal Cross-Entropy Method agent that steers its latent input. | | | |---|---| | **Version** | 1.0 (weights migrated from the original Keras training run; not retrained) | | **Format** | ONNX opset 18, float32, single self-contained file | | **Size** | 20.8 MB per generator · 8.8 KB per CEM agent | | **Parameters** | 5,193,997 per generator | | **Licence** | MIT (this repository) | | **Contact** | https://github.com/RyanMatthew04/kings-vision/issues | ## Intended use Generating mate-in-one training puzzles for the King's Vision trainer. The generator produces a *finished* checkmate; `kings_vision.puzzles.backtrack` reconstructs the position one move earlier and verifies by replaying forward. **Out of scope.** These models do not play chess, do not evaluate positions, and have no notion of a game. They sample from a distribution over checkmate-shaped board states. They are not a substitute for a curated puzzle set such as the Lichess database, which carries human difficulty ratings these do not. ## Architecture Generator: 100-d latent → `Dense(32768)` → `Reshape(8, 8, 512)` → four `Conv2DTranspose` blocks (256 → 128 → 64 → 32, stride 1, `same` padding, kernels 3/3/5/5), each with `LeakyReLU(0.2)` and batch normalisation → `Conv2D(13, k=7)` → softmax over the channel axis. Output is an 8×8×13 tensor: channel 0 is an empty square, 1–6 are white pawn…king, 7–12 are black. `argmax` per square decodes it to a board. That encoding is shared with the board classifier — see `kings_vision.core.encoding`. ## Training data **The training data no longer exists.** The original run derived checkmate positions from Lichess PGN archives, split them by mating piece into five datasets, and trained one generator per mate type. Neither the intermediate CSV nor the extraction script survived into either predecessor repository or its history. This is a real limitation and the main reason these weights were migrated rather than retrained: they are not reproducible from anything in this repository. The pipeline in `ml/data/` reconstructs the *method*, not the exact dataset. ## Training procedure Recovered from telemetry embedded in the original notebook (`docs/data/gan_training_log.csv`): - 10,000 epochs, batch size 128, Adam - ≈67 minutes wall clock on CPU - Discriminator settles at 46–50% accuracy — near chance, the healthy adversarial equilibrium — after an unstable first ~800 epochs The inherited README claimed 200–300 epochs at batch 32 over 2–4 hours. None of those figures matched the logs. ## The CEM agents Each agent is five Gaussians over the generator's 100-dimensional input, trained by Cross-Entropy Method against a binary reward (is this a legal checkmate?) with a penalty for repeating a position. **The generator's weights are frozen.** Only the input distribution is optimised, which is why the learned artifact is 8.8 KB rather than another 20.8 MB. One shipped agent (queen) has a mode whose weight decayed to 2×10⁻¹⁴ — an effectively collapsed mode. `MultiModalCEM.collapsed_modes` reports this. ## Evaluation `python -m ml.eval.eval_gan --n 1000 --seed 0`, Wilson intervals on validity and bootstrap intervals on uniqueness: | Mate type | Valid, raw | Valid, + CEM | Δ | Unique, raw | Unique, + CEM | Δ | Puzzle yield | |---|--:|--:|--:|--:|--:|--:|--:| | Queen | 47.4% | 88.2% | +40.8 | 100.0% | 50.8% | −49.2 | 83.6% | | Rook | 62.1% | 77.6% | +15.5 | 100.0% | 73.7% | −26.3 | 73.6% | | Bishop | 40.8% | 73.6% | +32.8 | 99.3% | 62.6% | −36.6 | 72.5% | | Knight | 53.4% | 93.9% | +40.5 | 100.0% | 73.1% | −26.9 | 93.7% | | Pawn | 46.9% | 82.7% | +35.8 | 99.8% | 74.8% | −24.9 | 27.5% | | **Mean** | **50.1%** | **83.2%** | **+33.1** | **99.8%** | **67.0%** | **−32.8** | **70.2%** | "Valid" is the fraction of samples that are legal positions in which the side to move is genuinely checkmated. "Unique" is the distinct fraction among valid samples. "Puzzle yield" is the fraction of all samples that convert to a mate-in-one with exactly one solution. ## Limitations and honest caveats **Diversity is traded for validity, roughly one for one.** Concentrating probability mass on elite samples is simultaneously what raises the hit rate and what narrows the output distribution. The diversity-aware reward limits this but does not remove it. Anyone wanting maximum variety should sample unsteered and filter, accepting a ~50% rejection rate. **Validity overstates usefulness for pawn mates.** 82.7% valid but only 27.5% convert to a unique-solution puzzle, because a pawn has few squares it could have come from. Judge by puzzle yield, not validity. **Puzzles are non-capturing mates.** The backtracking construction moves a piece backwards to an empty square, so mates delivered by a capture are unreachable. Discovered and promotion mates are reachable. **No difficulty rating.** These puzzles are not calibrated to human skill. Every one is mate-in-one, but "mate in one" spans a wide difficulty range and nothing here estimates where a given puzzle sits. **Positions are synthetic.** They are legal and reachable-looking, but not drawn from real games, and some carry the slightly artificial texture of GAN output. **Numerical fidelity.** The ONNX artifacts agree with the original Keras models to 100% argmax agreement across fixed-seed reference batches — identical puzzles, verified by `ml/export/build_onnx_generator.py`. Probability residuals of ~1e−5 are float32 accumulation differences across frameworks. ## Ethical considerations Low risk. The models generate chess positions. They carry no personal data, and the training corpus was public game archives. The most plausible harm is a user being served an artificial-feeling puzzle, which the limitations above cover. --- ## board square classifier A 13-class CNN that reads one square of a chessboard image. Sixty-four invocations produce a position. | | | |---|---| | **Version** | 1.0 | | **Format** | ONNX opset 18, float32, single self-contained file | | **Size** | ~0.4 MB | | **Parameters** | 95,373 | | **Input** | `(N, 3, 32, 32)` uint8-valued float, **0–255** | | **Output** | `(N, 13)` logits | | **Licence** | MIT | Scaling to `[0, 1]` happens *inside* the graph. The serving code passes raw pixel values and there is no preprocessing contract to get subtly wrong, which is a common and silent source of train/serve skew. ## Intended use Turning a screenshot of a chessboard into a FEN, so a position can be analysed without retyping it. Output is a *placement* only — a picture cannot show whose move it is, and the API returns the placement field rather than inventing the rest. **Out of scope.** Photographs of physical boards (this is trained on rendered 2D boards), boards at an angle, 3D piece sets, and any board that is not axis-aligned and square in the crop. ## Classes Channel order is shared with the mate generators — the same 8×8×13 encoding they emit into. `0` empty, `1–6` white pawn…king, `7–12` black. ## Training data **Fully synthetic, and that is the interesting part.** No corpus of screenshotted boards with per-square labels exists, and none is needed: rendering a known position makes every label exact and free. `ml/data/render_boards.py` draws positions sampled from random legal playouts — real piece densities and structures, including near-empty endgames — then degrades them the way a screenshot is degraded between a website and a file: rendered at 3× and downsampled so pieces land on fractional pixels, JPEG ringing, last-move highlights, coordinate labels bleeding into edge squares, blur, and brightness/contrast drift. - 4,000 boards / 256,000 squares for training across 6 themes - 600 boards / 38,400 squares for validation across 2 unseen themes - 600 boards / 38,400 squares for test across 2 further unseen themes **Splits are by theme, three ways.** A validation set drawn from the training themes reports 100% and measures memorisation. One drawn from the test themes lets epoch selection peek at the reported number. Six themes train, two select, two are looked at once. ## Training procedure 8 epochs, AdamW with one-cycle scheduling, batch 512, label smoothing 0.02, ~22 minutes on CPU. Best epoch selected on validation board-exact match. **Colour augmentation is the whole ballgame.** Per-channel gain, channel permutation, random greyscale, brightness/contrast, and noise — all applied at batch time. Every one is label-preserving because piece identity in these sprites is carried by *luminance* (white pieces are light with a dark outline, black the reverse), never by hue. So hue can be attacked freely, and must be. ## Evaluation Board-level exact match is the metric that matters. A FEN wrong in one square is wrong, and per-square accuracy hides that. | Themes | Per-square | **Board exact** | |---|--:|--:| | Seen (training) | 100.0000% | 100.00% | | Unseen (validation) | 100.0000% | 100.00% | | **Unseen (test)** | **99.8724%** | **94.50%** | ### The augmentation ablation `--no-augment` reproduces the failure the augmentation was written to fix: | | Per-square (test) | **Board exact (test)** | |---|--:|--:| | Without augmentation | 96.1745% | **58.83%** | | With augmentation | 99.8724% | **94.50%** | | Δ | +3.70 pts | **+35.67 pts** | Without it the model reached 100.0000% per-square accuracy on the six themes it trained on and collapsed on an unfamiliar palette. It had learned colour schemes, not chess pieces. ## Limitations **94.5%, not 99%.** One board in eighteen is still wrong somewhere on an unfamiliar theme. The legality repair search recovers some of those, but it can only fix boards whose errors make the position *illegal* — a bishop misread as a queen usually leaves a perfectly legal board and passes through silently. **Evaluated on synthetic themes, not real screenshots.** Unseen renders are a proxy for transfer, not a measurement of it. A hand-labelled set of real captures from Lichess and Chess.com would be the honest test and does not exist here. **Two held-out splits disagreed sharply.** In the ablation the validation themes scored 99.17% board exact and the test themes 58.83% — same model, same run. Transfer depends heavily on *which* unfamiliar theme, so 94.50% should be read as one sample from a wide distribution rather than a guarantee. **Assumes a square, axis-aligned, tightly-cropped board.** `autocrop` handles uniform page background and deliberately gives up rather than guess. A skewed or partially occluded board is out of scope and will fail without saying so. **Piece sets.** Trained on one sprite set (the Wikipedia pieces) with geometric and colour augmentation. A visually distinct set — Lichess's `cburnett`, Chess.com's `neo` — is untested. ## Ethical considerations Low risk. The model reads chessboards. It processes user-supplied images, which are held in memory for the duration of a request and not stored; uploads are capped at 8 MB. ---