|
Download README.md from Compactbot/compactlm-5m: direct link, hf CLI and curl.
- Browser
- Download file 4.04 kB
-
https://huggingface.co/Compactbot/compactlm-5m/resolve/main/README.md
- Command line
-
hf download hf://Compactbot/compactlm-5m/README.md
-
curl -L -o README.md https://huggingface.co/Compactbot/compactlm-5m/resolve/main/README.md
4.04 kB
| license: apache-2.0 | |
| pipeline_tag: text-generation | |
| language: en | |
| tags: | |
| - tiny-lm | |
| - tiny | |
| - slm | |
| - small-language-model | |
| - sub-1m | |
| - from-scratch | |
| - text-generation | |
| metrics: | |
| - perplexity | |
| # CompactLM-5M | |
| A **from-scratch** ~5M-parameter language model, trained from zero on a | |
| 300 MB slice of diverse real web text (chemistry, code, literature, general | |
| web). This is an independent small-model build in the "fits on a floppy disk" | |
| range β not a fine-tune of a bigger model. | |
| ## Architecture | |
| LLaMA-style decoder, built from scratch: | |
| | Field | Value | | |
| |---|---| | |
| | Parameters | **4,912,992** | | |
| | Layers | 6 | | |
| | Hidden size | 224 | | |
| | Attention | GQA β 7 query heads, 2 KV heads, head_dim 32 | | |
| | FFN | SwiGLU, intermediate 576 | | |
| | Norm | RMSNorm (eps 1e-6) | | |
| | Positional | RoPE (theta 10000) | | |
| | Embeddings | Tied (input = output head) | | |
| | Vocab | 8192 (BPE, trained on the corpus) | | |
| | Context | 512 | | |
| | Dtype | float32 | | |
| ## Training | |
| - **Data:** `/corpus_slice300m` β 300 MB of diverse real web text, | |
| tokenized to ~76.06M tokens (BPE, 8192 vocab). | |
| - **Steps:** 10,000 @ batch 32 Γ seq 512 | |
| - **Optimizer:** AdamW, betas (0.9, 0.95), weight decay 0.1 | |
| - **LR:** 3e-4, cosine decay with 10% warmup, floor 10% | |
| - **Grad clip:** 1.0 | |
| - **Hardware:** NVIDIA RTX 5090 (32 GB), CUDA | |
| ## Measured results (independently recomputed) | |
| - **Val perplexity: 58.91** β computed on the 1.52M-token held-out tail | |
| (last 2% of the corpus), token-level, by the author. | |
| - **Unigram baseline: 1453.67** on the same held-out tail. | |
| - The model beats the unigram floor by ~25Γ, i.e. it genuinely learned | |
| context, not just token frequencies. | |
| Note: the in-training `val_loss` (β0.008) is **not** a reliable number β the | |
| training loop's validation slice leaked from the training stream. The 58.91 | |
| above is the honest held-out figure. | |
| ## What it is good at / not | |
| At 5M parameters this model produces grammatical first sentences but is far | |
| from fluent: greedy decoding is coherent for roughly the first sentence and | |
| then collapses into repetition loops (e.g. "the world's largest city in the | |
| world is the world's largest cityβ¦"), and sampled decoding (top-p 0.9, | |
| temp 0.7) is more varied but still drifts into incoherence within a few | |
| sentences. Factual recall is weak. It is a demonstration of from-scratch | |
| small-model training, not a useful general assistant. | |
| ### Sample outputs (greedy, temp 0.0) | |
| > **The capital of France is** β "The capital of France is a very important | |
| > part of the world's economy." | |
| > **In machine learning, a neural network** β "In machine learning, a neural | |
| > network is a very important part of the development of the system." | |
| > **To make a cup of tea, you need** β "To make a cup of tea, you need to be | |
| > sure to use a cup of coffee." | |
| ### Sample outputs (temp 0.7, top-p 0.9) | |
| > **Once upon a time, there was a** β "Once upon a time, there was a great | |
| > chance to have a lot of life." | |
| > **The sun rises in the** β "The sun rises in the sun. Collecting a land, | |
| > which is known for its brightness in a space, has been caused by the dirt | |
| > and swords." | |
| ## Files | |
| ## Version | |
| This is **v2** of CompactLM-5M. It replaces an earlier 6,162,688-param build (d256/4L/4H, vocab 12288) whose card honestly noted it produced grammatical word-salad. This v2 retrain (d224/6L/GQA, vocab 8192) produces grammatical first sentences (better than the v1 word-salad) but still collapses into repetition loops under greedy decoding, so it is published as a from-scratch training demonstration, not as a coherent generator. | |
| - `model.safetensors` β weights (27 MB, F32). `head.weight` and `tok.weight` | |
| are tied (identical values); both keys are present for loaders that | |
| expect an untied head. | |
| - `tokenizer.json` β BPE tokenizer (8192 vocab) | |
| - `config.json` β architecture config | |
| ## Reproduction | |
| Trained with a custom GQA LLaMA script (6 layers, d224, SwiGLU). The exact | |
| script, tokenizer and training log are not bundled here; the architecture is | |
| fully specified in `config.json` and above. |