metadata
license: apache-2.0
tags:
- chess
- reinforcement-learning
- grpo
model_20m_16B — RL (GRPO) checkpoints
RL post-training trajectory for the chess pre-to-post compute-allocation study.
The pretraining base and SFT init for this model are
model_20m_16B and
model_20m_16B.
| parameters | 20m |
| pretraining tokens | 15,824,042,105 (15.8B) |
| checkpoints here | 30 (steps 100–3000) |
| checkpoints saved by the run | 60 |
Steps
100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000
Loading
Each global_step_N/ folder is self-contained. The models use a custom
tokenizer (tokenizer.py), and the remote-code resolver ignores subfolder=,
so download the folder first and load the local path:
from huggingface_hub import snapshot_download
from transformers import AutoModelForCausalLM, AutoTokenizer
step = "global_step_3000"
p = snapshot_download("Pre2Post-Chess-RL/Chess-RL-Models", allow_patterns=f"model_20m_16B/{step}/*") + f"/model_20m_16B/{step}"
model = AutoModelForCausalLM.from_pretrained(p, trust_remote_code=True)
tok = AutoTokenizer.from_pretrained(p, trust_remote_code=True)