Download README.md from Compactbot/compactlm-5m: direct link, hf CLI and curl.
- Browser
- Download file 3.63 kB
-
https://huggingface.co/Compactbot/compactlm-5m/resolve/refs%2Fpr%2F2/README.md
- Command line
-
hf download hf://Compactbot/compactlm-5m@refs/pr/2/README.md
-
curl -L -o README.md https://huggingface.co/Compactbot/compactlm-5m/resolve/refs%2Fpr%2F2/README.md
license: apache-2.0
language:
- en
pipeline_tag: text-generation
library_name: transformers
tags:
- tiny
- tiny-lm
- slm
- small-language-model
- from-scratch
- llama
datasets:
- HuggingFaceFW/fineweb-edu
metrics:
- perplexity
model-index:
- name: compactlm-5m
type: text-generation
params: 6162688
results:
- task:
name: Perplexity
type: perplexity
dataset:
name: fineweb-edu (held-out)
type: HuggingFaceFW/fineweb-edu
metrics:
- name: Perplexity
type: perplexity
value: 48.3
CompactLM-5M
A from-scratch LLaMA-style English language model, ~6.2M parameters, trained on fineweb-edu. Built to fulfill model-requests #14 (requested by @DedeProGames).
This is a small-language-model in the "fits on a floppy" sense: it was trained from random initialisation, not fine-tuned from a larger model.
Architecture
| Field | Value |
|---|---|
| Parameters | 6,162,688 (exact, sum(p.numel() for p in model.parameters())) |
| Style | LLaMA (RMSNorm, RoPE, SwiGLU MLP, tied embeddings) |
| d_model | 256 |
| Layers | 4 |
| Attention heads | 4 (MHA) |
| FFN (SwiGLU) | 640 |
| Vocab | 12,288 (BPE, same tokenizer as LDT-10M) |
| Context | 512 |
| Embeddings | tied (token embedding = LM head) |
The name says "5M" because that was the requested round target; the exact count for this architecture is 6,162,688.
Training
- Data: HuggingFaceFW/fineweb-edu,
~62M unique tokens (61.7M).
dclm-baseline-1.0was requested but was unreachable during the run (connection errors), so this checkpoint is fineweb-edu only — logged here rather than hidden. - Schedule: 20,000 steps, batch 128, ctx 512 →
1.31B token-passes over the 62M unique tokens (21 passes). - Optimizer: AdamW, peak LR 3e-4, warmup 300, cosine decay to 0.1×, weight decay 0.1, grad clip 1.0.
- Hardware: shared RTX 5090 (32 GB), run alongside other work.
Quality (honest)
- Val loss / perplexity: 3.8775 / 48.3 (held-out fineweb-edu, 1M tokens).
- The model produces grammatically intact English with no token-loops, no broken punctuation, and no hallucinated speaker tags — it completes 64-token generations cleanly.
- It is semantically shallow: short generations drift and repeat the topic word ("the church … the church … the church", "the sun rises in the sun"). This is the expected ceiling for a 6M-param model on 62M unique tokens. It is a working small LM at its scale, not a strong completion model.
Sample (seed 0, temp 0.8, top-k 40):
Prompt: The cat sat on the Output: The cat sat on the center of the church in the center of the church. The catalog is the same as the Bishop of the church, which includes the church.
Files
| File | What |
|---|---|
compactlm-5m.pt |
model_state_dict (39 tensors) + n_params + config |
config.json |
architecture config |
eval_fresh.json |
fresh val ppl + 15 generation samples + degeneracy check |
Usage
The checkpoint is a raw PyTorch state dict for the CompactLM class
(LLaMA-style, 4 layers). It is not a Hugging Face transformers checkpoint —
load it with the training script's model class. A transformers conversion is
a natural next step.
What it is not
- Not fine-tuned from a larger model.
- Not a strong completion model — see Quality above.
- Not a
transformers-loadable checkpoint yet.