TinyStories SLM

A ~12.4M-parameter decoder-only transformer, written from scratch β€” no transformers, no tokenizers, no lightning. Tokenizer, model, training loop, and sampler are all hand-written PyTorch. Full source, training logs, and design writeup: github.com/sur950/tinystories-slm.

It trains on TinyStories β€” Ronen Eldan and Yuanzhi Li's dataset of GPT-4-generated children's stories β€” and asks how small a language model can be and still write coherent English.

Not a transformers model

This architecture (RMSNorm, tied embeddings, hand-rolled KV cache) isn't a transformers-compatible class, so AutoModel.from_pretrained(...) won't work here. To actually run it:

git clone https://github.com/sur950/tinystories-slm.git
cd tinystories-slm
./setup.sh

Download stage2.pt/stage3.pt/tokenizer.json from this repo into ./checkpoints and ./data respectively, then:

python chat.py --stage 2   # base model: plain text continuation, no chat format
python chat.py --stage 3   # instruction-tuned: greetings, "tell me a story about X",
                            # and a redirect for anything outside TinyStories' domain

Two checkpoints

file what it is
stage2.pt Base language model. Trained in two passes over the ~560M-token corpus (Chinchilla-sized for this model). No chat format β€” continues whatever text precedes it, same as the original TinyStories models.
stage3.pt stage2.pt fine-tuned on a small (~4k example) instruction set β€” see instruct_raw.jsonl in this repo for the exact data. Responds to greetings, takes "tell me a story about X" requests, and gives a fixed redirect for anything outside the TinyStories domain instead of hallucinating an answer.

stage3.pt is the one to use for anything chat-like β€” not the lower-val-loss checkpoint from the same run (not included here). Val loss on this fine-tune is a single number blended across three very unevenly sized categories, so "lowest aggregate val" mostly tracks the largest category (story requests) and says little about whether the smaller, demo-critical categories (greetings, redirects) are reliable. Tested directly: the final checkpoint answered "Hi" correctly on-template ~4/5 times; the lowest-val checkpoint from partway through the same run only ~2/5. Full writeup in the GitHub repo's DEVELOPMENT.md.

Architecture

vocab 4096   d_model 384   layers 6   heads 6   context 512   tied embeddings
12,391,296 parameters total

Byte-level BPE tokenizer, vocab size 4096, fit on this same corpus β€” not the standard GPT-2 vocabulary, so token IDs from other tokenizers won't map correctly onto this model.

Limitations

  • English-only, children's-story register β€” no world knowledge, no code, no reasoning beyond what a simple short story requires.
  • stage3.pt's instruction-following is narrow by design: greetings, story requests within a small set of trained keywords, and off-domain redirects. It is not a general chatbot.
  • 512-token context window.

Citation

@misc{eldan2023tinystories,
  title         = {TinyStories: How Small Can Language Models Be and Still
                   Speak Coherent English?},
  author        = {Ronen Eldan and Yuanzhi Li},
  year          = {2023},
  eprint        = {2305.07759},
  archivePrefix = {arXiv},
}

Built by Suresh. MIT licensed.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Dataset used to train sur950/tinystories-slm

Paper for sur950/tinystories-slm