| --- |
| language: en |
| license: cc-by-sa-4.0 |
| tags: |
| - nanogpt |
| - gpt2 |
| - from-scratch |
| - pretraining |
| - wikipedia |
| - project-gutenberg |
| - research |
| pipeline_tag: text-generation |
| --- |
| |
| # WikipediaGutenberg-0.5B |
|
|
| A **from-scratch** GPT-2-style language model (nanoGPT), **530M params**, trained for **1 epoch** on a combination of **English Wikipedia** ([wikimedia/wikipedia](https://huggingface.co/datasets/wikimedia/wikipedia), 20231101.en) and **Project Gutenberg** ([common-pile/project_gutenberg_filtered](https://huggingface.co/datasets/common-pile/project_gutenberg_filtered)) — ~9.0B tokens, ~17 tokens/parameter (near the [Chinchilla](https://arxiv.org/abs/2203.15556) compute-optimal ratio). |
|
|
| This is a **research model** from a study of how from-scratch pretraining behaves across capacity / data / exposure. It is **not** instruction-tuned, not safety-aligned, and not intended for production use. |
|
|
| ## Model details |
| - **Architecture:** GPT-2 (learned positional embeddings, GeLU MLP, full multi-head attention). **16 layers, n_embd 1536, 24 heads (head_dim 64) — 530.30M params.** |
| - **Training data:** Wikipedia (4.17B tok) + Project Gutenberg (4.82B tok) = **8.99B train tokens (1 epoch)**; 1.00B val tokens. |
| - **Optimizer:** AdamW, LR 1.2e-3 → 1.2e-4 cosine, warmup 5000, batch 256 (65,536 tokens/iter), bf16, dropout 0.1. |
| - **Hardware:** 1× NVIDIA GH200, ~16.7 h. |
| - **Status:** training in progress. `ckpt.pt` is the **latest** checkpoint (mirrored here every ~30 min); it carries `iter_num` + optimizer state for clean resumption. |
|
|
| ## ⚠️ Format: nanoGPT, not HF Transformers |
| This is a **raw nanoGPT checkpoint** (a torch `state_dict` + config), **not** loadable via `AutoModelForCausalLM.from_pretrained`. Load it with [nanoGPT](https://github.com/karpathy/nanoGPT): |
| ```python |
| import torch |
| from model import GPTConfig, GPT |
| ckpt = torch.load("ckpt.pt", map_location="cuda", weights_only=False) |
| model = GPT(GPTConfig(**ckpt["model_args"])) |
| model.load_state_dict(ckpt["model"]) |
| ``` |
|
|
| ## Intended use & limitations |
| - **Intended:** ML research / education on pretraining dynamics and small-model knowledge elicitation. |
| - **Not intended:** deployment, factual QA, instruction-following. At 530M it has limited and weakly-elicitable factual knowledge (see the project's elicitation evals). |
|
|
| ## Evaluation |
| The full **elicitation scoreboard** (latent completion-probe recall, multiple-choice, free-form QA, few-shot/CoT) vs the 253M baselines **will be added here on training completion**. |
|
|
| ## Data & license |
| Wikipedia (CC BY-SA 4.0) + Project Gutenberg (public domain). Per Wikipedia's share-alike, model weights are released under **CC BY-SA 4.0**. |
|
|
| ## Sources |
| - nanoGPT: [karpathy/nanoGPT](https://github.com/karpathy/nanoGPT) |
| - Chinchilla (compute-optimal scaling): [Hoffmann et al., 2022](https://arxiv.org/abs/2203.15556) |
| - Datasets: [wikimedia/wikipedia](https://huggingface.co/datasets/wikimedia/wikipedia) · [common-pile/project_gutenberg_filtered](https://huggingface.co/datasets/common-pile/project_gutenberg_filtered) |
|
|