CHSM8 / README.md
LegumMagister's picture
docs: add model architecture training and data card
0b22ae8
|
Raw
History Blame Contribute Delete
2.78 kB
---
tags:
- chess
- pytorch
- language-model
license: unknown
---
# CHSM8
CHSM8 is a research chess move-prediction model. Given a game prefix, it
predicts the next chess move while applying a legal-move mask at inference.
It is a project-specific PyTorch model, not a standard Transformers model.
Source code, architecture, training scripts, and data pipeline:
[github.com/nov3o/llm_chess](https://github.com/nov3o/llm_chess).
## Architecture
- Llama-style causal chess decoder: 12 layers, width 960, 15 attention heads,
SwiGLU MLP (width 3,520), RoPE, and Q/K normalization.
- Factorized chess output heads predict move kind, source square, destination
square, piece, and promotion rather than a flat move vocabulary.
- About 165M trainable parameters. The architecture retains a frozen dummy
cross-attention scaffold for compatibility with later text-conditioned work.
- Trained in bf16 with FlashAttention, fused AdamW, gradient clipping 1.0, and
weight decay 0.1.
## Training data and schedule
Training uses packed Lichess standard-chess game records, represented as move
sequences. The final branch reads the `3p-s` packed corpus without replacement
in 512-position sequences; shards 77–127 were globally reshuffled before the
85% checkpoint to remove a temporal data-order artefact.
The planned one-pass budget is 78.32B chess positions. Optimisation uses a
position-indexed cosine learning-rate schedule: 5% warmup, peak LR `5e-4`, and
minimum LR equal to 10% of the peak. A 1% held-out split is reserved for
validation. The model's legal-move penalty weight is 0.5.
## Checkpoints
| File | Step | Training exposure | Description |
|---|---:|---:|---|
| `chsm8_best_pct85.pt` | 1,015,809 | 85.00% | Best checkpoint from the wide-shuffled lineage; it remained stable through the prior loss-bump window. |
| `chsm8_latest.pt` | 1,194,960 | 99.99% | Latest/final checkpoint from the training run. |
Both are weights-only PyTorch dictionaries containing `model_state_dict`,
`step`, `total_training_time`, `total_tokens_processed`, and `cfg`. Load them
with `FactoredChessDecoder` / `DecoderConfig` from this repository's
[`src/model.py`](https://github.com/nov3o/llm_chess/blob/main/src/model.py).
## Evaluation note
Playing evaluations use 200 colour-balanced games with greedy legal-move
decoding against Stockfish constrained to `UCI_Elo=2000` at 0.05 seconds per
move. This is a reproducible project metric, not an official FIDE Elo rating
or a general-purpose chess-engine benchmark.
## Limitations
CHSM8 is a research checkpoint with no search, opening book, endgame tablebase,
or standard Transformers interface. It has been evaluated only under the
project protocol above; use it for research and reproduce results from the
linked codebase.