--- license: apache-2.0 pipeline_tag: text-generation language: en tags: - tiny-lm - tiny - slm - small-language-model - sub-1m - from-scratch - text-generation metrics: - perplexity --- # CompactLM-5M A **from-scratch** ~5M-parameter language model, trained from zero on a 300 MB slice of diverse real web text (chemistry, code, literature, general web). This is an independent small-model build in the "fits on a floppy disk" range — not a fine-tune of a bigger model. ## Architecture LLaMA-style decoder, built from scratch: | Field | Value | |---|---| | Parameters | **4,912,992** | | Layers | 6 | | Hidden size | 224 | | Attention | GQA — 7 query heads, 2 KV heads, head_dim 32 | | FFN | SwiGLU, intermediate 576 | | Norm | RMSNorm (eps 1e-6) | | Positional | RoPE (theta 10000) | | Embeddings | Tied (input = output head) | | Vocab | 8192 (BPE, trained on the corpus) | | Context | 512 | | Dtype | float32 | ## Training - **Data:** `/corpus_slice300m` — 300 MB of diverse real web text, tokenized to ~76.06M tokens (BPE, 8192 vocab). - **Steps:** 10,000 @ batch 32 × seq 512 - **Optimizer:** AdamW, betas (0.9, 0.95), weight decay 0.1 - **LR:** 3e-4, cosine decay with 10% warmup, floor 10% - **Grad clip:** 1.0 - **Hardware:** NVIDIA RTX 5090 (32 GB), CUDA ## Measured results (independently recomputed) - **Val perplexity: 58.91** — computed on the 1.52M-token held-out tail (last 2% of the corpus), token-level, by the author. - **Unigram baseline: 1453.67** on the same held-out tail. - The model beats the unigram floor by ~25×, i.e. it genuinely learned context, not just token frequencies. Note: the in-training `val_loss` (≈0.008) is **not** a reliable number — the training loop's validation slice leaked from the training stream. The 58.91 above is the honest held-out figure. ## What it is good at / not At 5M parameters this model produces grammatical first sentences but is far from fluent: greedy decoding is coherent for roughly the first sentence and then collapses into repetition loops (e.g. "the world's largest city in the world is the world's largest city…"), and sampled decoding (top-p 0.9, temp 0.7) is more varied but still drifts into incoherence within a few sentences. Factual recall is weak. It is a demonstration of from-scratch small-model training, not a useful general assistant. ### Sample outputs (greedy, temp 0.0) > **The capital of France is** → "The capital of France is a very important > part of the world's economy." > **In machine learning, a neural network** → "In machine learning, a neural > network is a very important part of the development of the system." > **To make a cup of tea, you need** → "To make a cup of tea, you need to be > sure to use a cup of coffee." ### Sample outputs (temp 0.7, top-p 0.9) > **Once upon a time, there was a** → "Once upon a time, there was a great > chance to have a lot of life." > **The sun rises in the** → "The sun rises in the sun. Collecting a land, > which is known for its brightness in a space, has been caused by the dirt > and swords." ## Files ## Version This is **v2** of CompactLM-5M. It replaces an earlier 6,162,688-param build (d256/4L/4H, vocab 12288) whose card honestly noted it produced grammatical word-salad. This v2 retrain (d224/6L/GQA, vocab 8192) produces grammatical first sentences (better than the v1 word-salad) but still collapses into repetition loops under greedy decoding, so it is published as a from-scratch training demonstration, not as a coherent generator. - `model.safetensors` — weights (27 MB, F32). `head.weight` and `tok.weight` are tied (identical values); both keys are present for loaders that expect an untied head. - `tokenizer.json` — BPE tokenizer (8192 vocab) - `config.json` — architecture config ## Reproduction Trained with a custom GQA LLaMA script (6 layers, d224, SwiGLU). The exact script, tokenizer and training log are not bundled here; the architecture is fully specified in `config.json` and above.