|
Download README.md from Compactbot/checkpoint-health: direct link, hf CLI and curl.
- Browser
- Download file 3.65 kB
-
https://huggingface.co/Compactbot/checkpoint-health/resolve/main/README.md
- Command line
-
hf download hf://Compactbot/checkpoint-health/README.md
-
curl -L -o README.md https://huggingface.co/Compactbot/checkpoint-health/resolve/main/README.md
3.65 kB
metadata
tags:
- builder-tool
- checkpoint
- tiny-lm
- debugging
pipeline_tag: text-generation
Checkpoint Health Checker
A single-file tool for inspecting .pt checkpoints from tiny language model training runs. Tells you if your model is healthy, degenerate, or broken — before you waste time uploading or evaluating it.
What it checks
| Check | What it looks for |
|---|---|
| Architecture | vocab size, d_model, layers, heads, tied embeddings (inferred from state_dict keys + config) |
| Parameter count | Total parameters (sum of all tensor numel) |
| Weight health | NaN, Inf, extreme values (>100 abs) in any weight tensor |
| Generation sanity | EOS collapse, degenerate loops, low diversity, too-short output |
Requirements
- Python 3.8+
- PyTorch (
pip install torch) - A tokenizer (HuggingFace
tokenizerslibrary) for the generation check — optional, the weight checks work without it
Usage
# Basic: weight health only (no tokenizer needed)
python3 checkpoint_health.py model.pt
# With generation check
python3 checkpoint_health.py model.pt --tokenizer tokenizer.json --prompt "Once upon a time"
# JSON output (for scripting)
python3 checkpoint_health.py model.pt --tokenizer tokenizer.json --json
Exit codes
| Code | Meaning |
|---|---|
| 0 | Healthy — weights clean, generation looks reasonable |
| 1 | Warning — degenerate generation or extreme weight values |
| 2 | Critical — NaN/Inf weights detected (model is broken) |
Example output
Healthy model (12.6M params, TinyStories-trained)
=== Checkpoint Health Report ===
File: story10m/final.pt (20.6 MB)
[ARCHITECTURE]
vocab_size: 256
d_model: 256
n_layers: 6
n_heads: 4
n_params: 12,603,648
tied_embeddings: yes
[WEIGHT HEALTH]
NaN tensors: 0
Inf tensors: 0
Max |val|: 3.214 (wte.weight)
Status: CLEAN
[GENERATION]
Prompt: "Once upon a time"
Output: "Once upon a time, there was a little girl who lived in a small house..."
New tokens: 42
EOS hits: 0
Trigram repetition: 4.2%
Char diversity: 0.412
Status: OK
VERDICT: HEALTHY
Degenerate model (1.6M params, undertrained)
=== Checkpoint Health Report ===
File: jokeclaude/final.pt (6.3 MB)
[ARCHITECTURE]
vocab_size: 50257
d_model: 128
n_layers: 2
n_params: 1,638,400
tied_embeddings: yes
[WEIGHT HEALTH]
NaN tensors: 0
Inf tensors: 0
Max |val|: 1.847 (wte.weight)
Status: CLEAN
[GENERATION]
Prompt: "The cat sat on the mat and"
Output: "The cat sat on the mat and the the the the the the the the the..."
New tokens: 100
EOS hits: 0
Trigram repetition: 78.3%
Char diversity: 0.089
Status: LOW_DIVERSITY
VERDICT: WARNING — Low character diversity (0.089). Model may be stuck on a character.
Limitations
- RoPE/GQA architectures: The generation check requires instantiating a model class. If the architecture uses RoPE or grouped query attention (common in 2024+ tiny LMs), the generation check is skipped with an explanation. Weight health and architecture inference still work.
- Tokenizer required for generation: Without
--tokenizer, only weight health is reported. - Single-file only: Expects a
.ptfile with either a raw state_dict or a dict withmodel/model_state_dictkey.
When to use it
- Right after a training run completes, before uploading
- When a model "works" in your training loop but generates garbage at inference
- To quickly triage a checkpoint you found on the Hub
- To confirm a training run didn't produce NaN weights (silent failure mode)