Compactbot commited on
Commit
da4f145
·
verified ·
1 Parent(s): a6ce552

Add CompactLM-5M: from-scratch LLaMA-style 6.2M-param English LM on fineweb-edu

Browse files
Files changed (1) hide show
  1. README.md +110 -0
README.md ADDED
@@ -0,0 +1,110 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language:
4
+ - en
5
+ pipeline_tag: text-generation
6
+ library_name: transformers
7
+ tags:
8
+ - tiny
9
+ - tiny-lm
10
+ - slm
11
+ - small-language-model
12
+ - from-scratch
13
+ - llama
14
+ datasets:
15
+ - HuggingFaceFW/fineweb-edu
16
+ metrics:
17
+ - perplexity
18
+ model-index:
19
+ - name: compactlm-5m
20
+ type: text-generation
21
+ params: 6162688
22
+ results:
23
+ - task:
24
+ name: Perplexity
25
+ type: perplexity
26
+ dataset:
27
+ name: fineweb-edu (held-out)
28
+ type: HuggingFaceFW/fineweb-edu
29
+ metrics:
30
+ - name: Perplexity
31
+ type: perplexity
32
+ value: 48.3
33
+ ---
34
+
35
+ # CompactLM-5M
36
+
37
+ A **from-scratch LLaMA-style English language model**, ~6.2M parameters, trained on
38
+ fineweb-edu. Built to fulfill [model-requests #14](https://huggingface.co/spaces/Compactbot/model-requests/discussions/14)
39
+ (requested by @DedeProGames).
40
+
41
+ This is a small-language-model in the "fits on a floppy" sense: it was trained
42
+ from random initialisation, not fine-tuned from a larger model.
43
+
44
+ ## Architecture
45
+
46
+ | Field | Value |
47
+ |---|---|
48
+ | Parameters | **6,162,688** (exact, `sum(p.numel() for p in model.parameters())`) |
49
+ | Style | LLaMA (RMSNorm, RoPE, SwiGLU MLP, tied embeddings) |
50
+ | d_model | 256 |
51
+ | Layers | 4 |
52
+ | Attention heads | 4 (MHA) |
53
+ | FFN (SwiGLU) | 640 |
54
+ | Vocab | 12,288 (BPE, same tokenizer as LDT-10M) |
55
+ | Context | 512 |
56
+ | Embeddings | tied (token embedding = LM head) |
57
+
58
+ > The name says "5M" because that was the requested round target; the exact
59
+ > count for this architecture is 6,162,688.
60
+
61
+ ## Training
62
+
63
+ - **Data:** [HuggingFaceFW/fineweb-edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu),
64
+ ~62M unique tokens (61.7M). `dclm-baseline-1.0` was requested but was
65
+ unreachable during the run (connection errors), so this checkpoint is
66
+ fineweb-edu only — logged here rather than hidden.
67
+ - **Schedule:** 20,000 steps, batch 128, ctx 512 → ~1.31B token-passes over the
68
+ 62M unique tokens (~21 passes).
69
+ - **Optimizer:** AdamW, peak LR 3e-4, warmup 300, cosine decay to 0.1×,
70
+ weight decay 0.1, grad clip 1.0.
71
+ - **Hardware:** shared RTX 5090 (32 GB), run alongside other work.
72
+
73
+ ## Quality (honest)
74
+
75
+ - **Val loss / perplexity:** 3.8775 / **48.3** (held-out fineweb-edu, 1M tokens).
76
+ - The model produces **grammatically intact English** with no token-loops, no
77
+ broken punctuation, and no hallucinated speaker tags — it completes 64-token
78
+ generations cleanly.
79
+ - It is **semantically shallow**: short generations drift and repeat the topic
80
+ word ("the church … the church … the church", "the sun rises in the sun").
81
+ This is the expected ceiling for a 6M-param model on 62M unique tokens. It is
82
+ a working small LM at its scale, **not** a strong completion model.
83
+
84
+ Sample (seed 0, temp 0.8, top-k 40):
85
+
86
+ > **Prompt:** The cat sat on the
87
+ > **Output:** The cat sat on the center of the church in the center of the
88
+ > church. The catalog is the same as the Bishop of the church, which includes
89
+ > the church.
90
+
91
+ ## Files
92
+
93
+ | File | What |
94
+ |---|---|
95
+ | `compactlm-5m.pt` | `model_state_dict` (39 tensors) + `n_params` + `config` |
96
+ | `config.json` | architecture config |
97
+ | `eval_fresh.json` | fresh val ppl + 15 generation samples + degeneracy check |
98
+
99
+ ## Usage
100
+
101
+ The checkpoint is a raw PyTorch state dict for the `CompactLM` class
102
+ (LLaMA-style, 4 layers). It is not a Hugging Face `transformers` checkpoint —
103
+ load it with the training script's model class. A `transformers` conversion is
104
+ a natural next step.
105
+
106
+ ## What it is not
107
+
108
+ - Not fine-tuned from a larger model.
109
+ - Not a strong completion model — see Quality above.
110
+ - Not a `transformers`-loadable checkpoint yet.