| --- |
| language: |
| - en |
| license: apache-2.0 |
| library_name: pytorch |
| inference: false |
| datasets: |
| - roneneldan/TinyStories |
| tags: |
| - tinystories |
| - language-modeling |
| - from-scratch |
| - small-language-model |
| --- |
| |
| # v11-tinystories-115m-base |
|
|
| A 115.1M-parameter decoder-only transformer trained **from scratch** on |
| TinyStories for 16M tokens, using the |
| [v11 tokenizer](https://huggingface.co/chrishayuk/v11-tokenizer). |
|
|
| This is the **base pretrain only**: phase 1, no frozen-FFN attention retrain, and |
| **no maths mid-training**. It writes competent children's-story English and it |
| **cannot do arithmetic** — it narrates straight past the place a number belongs |
| rather than putting a wrong number there. That is the intended state of this |
| checkpoint, not a defect: it is the starting point for the mid-training |
| experiments in |
| [tinystories-train-video](https://github.com/chrishayuk/tinystories-train-video). |
|
|
| ## Loading |
|
|
| Not an `AutoModel` — `TinyModel` is a 3-file Gemma-shaped decoder (RMSNorm, RoPE, |
| GQA, gated FFN, tied embeddings), shipped in this repo under `tiny_model_v11/`. |
|
|
| ```python |
| import json, torch |
| from huggingface_hub import snapshot_download |
| from safetensors.torch import load_file |
| from transformers import AutoTokenizer |
| |
| path = snapshot_download("chrishayuk/v11-tinystories-115m-base") |
| |
| import sys; sys.path.insert(0, path) |
| from tiny_model_v11 import TinyModel |
| |
| cfg = json.load(open(f"{path}/config.json")) |
| model = TinyModel( |
| vocab_size=cfg["vocab_size"], dim=cfg["dim"], n_layers=cfg["n_layers"], |
| ffn_dim=cfg["ffn_dim"], n_heads=cfg["n_heads"], n_kv_heads=cfg["n_kv_heads"], |
| max_seq=cfg["max_seq"], |
| ) |
| model.load_state_dict(load_file(f"{path}/model.safetensors")) |
| model.eval() |
| |
| tok = AutoTokenizer.from_pretrained("chrishayuk/v11-tinystories-115m-base") |
| ids = [tok.convert_tokens_to_ids("<s>")] + tok("Once upon a time", add_special_tokens=False)["input_ids"] |
| for _ in range(40): |
| logits = model(torch.tensor([ids[-cfg["max_seq"]:]]))[0, -1] |
| ids.append(int(logits.argmax())) |
| print(tok.decode(ids)) |
| ``` |
|
|
| That exact round-trip — download, build, load, generate — is run against the |
| uploaded files at publish time. It is not an untested snippet. |
|
|
| ## Architecture |
|
|
| | | | |
| |---|---| |
| | Parameters | 115,149,312 | |
| | Layers | 20 | |
| | Model dim | 512 | |
| | Attention heads | 8 (4 KV heads, GQA) | |
| | FFN dim | 2048 | |
| | Context | 256 tokens | |
| | Vocabulary | 71,260 | |
| | Embeddings | tied (`lm_head.weight is embed.weight`) | |
|
|
| The embedding table is 36.5M parameters — **32% of the |
| whole model is its vocabulary lookup**, which is what a 71k vocabulary costs at |
| this width. |
|
|
| `rope_freqs` is a complex64 buffer in the state dict. It is derived from |
| `dim`/`n_heads`/`max_seq` and recomputed on construction, so it round-trips but |
| carries no learned information. |
|
|
| ## Identity |
|
|
| **Identity is the content hash, not the Hub revision** — re-pushing identical |
| bytes mints a new commit oid, and a README edit does too. Join on these: |
|
|
| | | | |
| |---|---| |
| | `model.safetensors` sha256 | `1841e0581574629716b646dacd4e70feaca153a8adc5ecb0b77e0e2ebdf78d9c` | |
| | `tokenizer.json` sha256 | `10dd51100331ab503115db23eee7e8dc3e360e3aed697c8a2e1b12b8f46031ae` | |
| | Tokenizer repo | [`chrishayuk/v11-tokenizer`](https://huggingface.co/chrishayuk/v11-tokenizer) | |
| | Source | [tinystories-train-video](https://github.com/chrishayuk/tinystories-train-video) | |
| | Source commit | `fbb5418` | |
|
|
| The tokenizer sha is the **same value** this checkpoint's training run wrote into |
| its `meta.json` as `tokenizer_hash`, and the same one |
| [`chrishayuk/v11-tokenizer`](https://huggingface.co/chrishayuk/v11-tokenizer) publishes. A |
| checkpoint driven by a different tokenizer produces fluent nonsense rather than |
| an error, so that join is checked mechanically at publish time, not asserted here. |
|
|
| ## Training |
|
|
| | | | |
| |---|---| |
| | Corpus | [`roneneldan/TinyStories`](https://huggingface.co/datasets/roneneldan/TinyStories) @ `f54c09fd2331…` | |
| | Tokens | 16M | |
| | Steps | 15,625 | |
| | Batch × context | 4 × 256 | |
| | Optimiser | AdamW, lr 0.0003, weight decay 0.01, grad clip 1.0 | |
| | LR schedule | linear warmup then linear decay to 5% | |
| | Seed | 42 | |
| | Precision | fp32 | |
|
|
| The dataset revision is pinned, so the document set is reproducible and held-out |
| text can be shown never to have been trained on. |
|
|
| ## Capability emergence |
|
|
| `emergence.json` carries the generations captured at each milestone, plus the |
| loss trace. Greedy, 30 new tokens, same prompts throughout: |
|
|
| | Tokens | Step | First sample continuation | |
| |---|---|---| |
| | 0.00M | 0 | simple_expressionesauesauesauesauesauesauesauesauesauesauesauesauesauesauesaumarredmarredmarredmarredmarredmar | |
| | 0.10M | 97 | .............................. | |
| | 1.00M | 976 | , there was a little girl named Lily. She loved to play with her toys and her friends. One day, L | |
| | 2.00M | 1953 | , there was a little girl named Lily. She loved to play outside and play with her friends. One day, Li | |
| | 5.00M | 4882 | , there was a little girl named Lily. She loved to play outside in the park. One day, she saw a big, | |
| | 8.00M | 7812 | , there was a little girl named Lily. She loved to play outside in the sunshine. One day, she saw a big, | |
| | 12.00M | 11718 | , there was a little girl named Lily. She loved to play outside in the sun. One day, she saw a big, | |
| | 16.00M | 15625 | , there was a little girl named Lily. She loved to play outside in the sun. One day, she saw a big, | |
| |
| Only the final row's weights are published here; the earlier rows are the same |
| run mid-flight. |
| |
| ## What this model cannot do |
| |
| It cannot do arithmetic. Nothing in TinyStories teaches addition, and number words |
| in that corpus are narrative texture rather than quantities — "once upon a time |
| there were two" is an idiom the model learns the way it learns "happily ever |
| after". |
| |
| The interesting part is *how* it fails. It does not answer with a wrong number; it |
| carries on telling the story, straight past the place a number belongs. |
| |
| Its own generation at the final milestone, greedy, from this run: |
| |
| > *Lily had three apples. Tom gave her four more. Now Lily has* **a big piece of paper. She likes the color. Tom likes the color. He likes the** |
| |
| It has also seen 16M tokens, which is **not converged**. Continued |
| training on more or less anything improves it, so do not read an improvement after |
| mid-training as evidence that the mid-training data helped specifically. |
| |
| ## Intended use |
| |
| Research and teaching on compact language models: mid-training, tool-use / |
| tool-call training, and measuring the difference between memorising a |
| distribution and learning an algorithm. Not intended for any production use, and |
| it has no safety training of any kind. |
| |
| ## Provenance |
| |
| `provenance.json` records the model and tokenizer hashes, the corpus pin, and the |
| run config. The chuk-datasets catalog holds a content-addressed, bit-reproducible |
| tokenization of the same corpus (`tiny-model/v11-rust-tokenized-phase1`); these |
| weights came from the pinned HF revision instead, which is reproducible given the |
| seed but not bit-identical. That distinction is recorded rather than glossed. |
| |