--- license: apache-2.0 language: - en pipeline_tag: text-generation library_name: transformers tags: - tiny - tiny-lm - slm - small-language-model - from-scratch - llama datasets: - HuggingFaceFW/fineweb-edu metrics: - perplexity model-index: - name: compactlm-5m type: text-generation params: 6162688 results: - task: name: Perplexity type: perplexity dataset: name: fineweb-edu (held-out) type: HuggingFaceFW/fineweb-edu metrics: - name: Perplexity type: perplexity value: 48.3 --- # CompactLM-5M A **from-scratch LLaMA-style English language model**, ~6.2M parameters, trained on fineweb-edu. Built to fulfill [model-requests #14](https://huggingface.co/spaces/Compactbot/model-requests/discussions/14) (requested by @DedeProGames). This is a small-language-model in the "fits on a floppy" sense: it was trained from random initialisation, not fine-tuned from a larger model. ## Architecture | Field | Value | |---|---| | Parameters | **6,162,688** (exact, `sum(p.numel() for p in model.parameters())`) | | Style | LLaMA (RMSNorm, RoPE, SwiGLU MLP, tied embeddings) | | d_model | 256 | | Layers | 4 | | Attention heads | 4 (MHA) | | FFN (SwiGLU) | 640 | | Vocab | 12,288 (BPE, same tokenizer as LDT-10M) | | Context | 512 | | Embeddings | tied (token embedding = LM head) | > The name says "5M" because that was the requested round target; the exact > count for this architecture is 6,162,688. ## Training - **Data:** [HuggingFaceFW/fineweb-edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu), ~62M unique tokens (61.7M). `dclm-baseline-1.0` was requested but was unreachable during the run (connection errors), so this checkpoint is fineweb-edu only — logged here rather than hidden. - **Schedule:** 20,000 steps, batch 128, ctx 512 → ~1.31B token-passes over the 62M unique tokens (~21 passes). - **Optimizer:** AdamW, peak LR 3e-4, warmup 300, cosine decay to 0.1×, weight decay 0.1, grad clip 1.0. - **Hardware:** shared RTX 5090 (32 GB), run alongside other work. ## Quality (honest) - **Val loss / perplexity:** 3.8775 / **48.3** (held-out fineweb-edu, 1M tokens). - The model produces **grammatically intact English** with no token-loops, no broken punctuation, and no hallucinated speaker tags — it completes 64-token generations cleanly. - It is **semantically shallow**: short generations drift and repeat the topic word ("the church … the church … the church", "the sun rises in the sun"). This is the expected ceiling for a 6M-param model on 62M unique tokens. It is a working small LM at its scale, **not** a strong completion model. Sample (seed 0, temp 0.8, top-k 40): > **Prompt:** The cat sat on the > **Output:** The cat sat on the center of the church in the center of the > church. The catalog is the same as the Bishop of the church, which includes > the church. ## Files | File | What | |---|---| | `compactlm-5m.pt` | `model_state_dict` (39 tensors) + `n_params` + `config` | | `config.json` | architecture config | | `eval_fresh.json` | fresh val ppl + 15 generation samples + degeneracy check | ## Usage The checkpoint is a raw PyTorch state dict for the `CompactLM` class (LLaMA-style, 4 layers). It is not a Hugging Face `transformers` checkpoint — load it with the training script's model class. A `transformers` conversion is a natural next step. ## What it is not - Not fine-tuned from a larger model. - Not a strong completion model — see Quality above. - Not a `transformers`-loadable checkpoint yet.