|
Download README.md from Compactbot/checkpoint-health: direct link, hf CLI and curl.
- Browser
- Download file 3.65 kB
-
https://huggingface.co/Compactbot/checkpoint-health/resolve/main/README.md
- Command line
-
hf download hf://Compactbot/checkpoint-health/README.md
-
curl -L -o README.md https://huggingface.co/Compactbot/checkpoint-health/resolve/main/README.md
3.65 kB
| tags: | |
| - builder-tool | |
| - checkpoint | |
| - tiny-lm | |
| - debugging | |
| pipeline_tag: text-generation | |
| # Checkpoint Health Checker | |
| A single-file tool for inspecting `.pt` checkpoints from tiny language model training runs. Tells you if your model is healthy, degenerate, or broken β before you waste time uploading or evaluating it. | |
| ## What it checks | |
| | Check | What it looks for | | |
| |-------|------------------| | |
| | **Architecture** | vocab size, d_model, layers, heads, tied embeddings (inferred from state_dict keys + config) | | |
| | **Parameter count** | Total parameters (sum of all tensor numel) | | |
| | **Weight health** | NaN, Inf, extreme values (>100 abs) in any weight tensor | | |
| | **Generation sanity** | EOS collapse, degenerate loops, low diversity, too-short output | | |
| ## Requirements | |
| - Python 3.8+ | |
| - PyTorch (`pip install torch`) | |
| - A tokenizer (HuggingFace `tokenizers` library) for the generation check β optional, the weight checks work without it | |
| ## Usage | |
| ```bash | |
| # Basic: weight health only (no tokenizer needed) | |
| python3 checkpoint_health.py model.pt | |
| # With generation check | |
| python3 checkpoint_health.py model.pt --tokenizer tokenizer.json --prompt "Once upon a time" | |
| # JSON output (for scripting) | |
| python3 checkpoint_health.py model.pt --tokenizer tokenizer.json --json | |
| ``` | |
| ## Exit codes | |
| | Code | Meaning | | |
| |------|---------| | |
| | 0 | Healthy β weights clean, generation looks reasonable | | |
| | 1 | Warning β degenerate generation or extreme weight values | | |
| | 2 | Critical β NaN/Inf weights detected (model is broken) | | |
| ## Example output | |
| ### Healthy model (12.6M params, TinyStories-trained) | |
| ``` | |
| === Checkpoint Health Report === | |
| File: story10m/final.pt (20.6 MB) | |
| [ARCHITECTURE] | |
| vocab_size: 256 | |
| d_model: 256 | |
| n_layers: 6 | |
| n_heads: 4 | |
| n_params: 12,603,648 | |
| tied_embeddings: yes | |
| [WEIGHT HEALTH] | |
| NaN tensors: 0 | |
| Inf tensors: 0 | |
| Max |val|: 3.214 (wte.weight) | |
| Status: CLEAN | |
| [GENERATION] | |
| Prompt: "Once upon a time" | |
| Output: "Once upon a time, there was a little girl who lived in a small house..." | |
| New tokens: 42 | |
| EOS hits: 0 | |
| Trigram repetition: 4.2% | |
| Char diversity: 0.412 | |
| Status: OK | |
| VERDICT: HEALTHY | |
| ``` | |
| ### Degenerate model (1.6M params, undertrained) | |
| ``` | |
| === Checkpoint Health Report === | |
| File: jokeclaude/final.pt (6.3 MB) | |
| [ARCHITECTURE] | |
| vocab_size: 50257 | |
| d_model: 128 | |
| n_layers: 2 | |
| n_params: 1,638,400 | |
| tied_embeddings: yes | |
| [WEIGHT HEALTH] | |
| NaN tensors: 0 | |
| Inf tensors: 0 | |
| Max |val|: 1.847 (wte.weight) | |
| Status: CLEAN | |
| [GENERATION] | |
| Prompt: "The cat sat on the mat and" | |
| Output: "The cat sat on the mat and the the the the the the the the the..." | |
| New tokens: 100 | |
| EOS hits: 0 | |
| Trigram repetition: 78.3% | |
| Char diversity: 0.089 | |
| Status: LOW_DIVERSITY | |
| VERDICT: WARNING β Low character diversity (0.089). Model may be stuck on a character. | |
| ``` | |
| ## Limitations | |
| - **RoPE/GQA architectures**: The generation check requires instantiating a model class. If the architecture uses RoPE or grouped query attention (common in 2024+ tiny LMs), the generation check is **skipped** with an explanation. Weight health and architecture inference still work. | |
| - **Tokenizer required for generation**: Without `--tokenizer`, only weight health is reported. | |
| - **Single-file only**: Expects a `.pt` file with either a raw state_dict or a dict with `model`/`model_state_dict` key. | |
| ## When to use it | |
| - Right after a training run completes, before uploading | |
| - When a model "works" in your training loop but generates garbage at inference | |
| - To quickly triage a checkpoint you found on the Hub | |
| - To confirm a training run didn't produce NaN weights (silent failure mode) |