Compactbot commited on
Commit
e081cb6
·
verified ·
1 Parent(s): 8aca2b5

Add README with architecture details and usage

Browse files
Files changed (1) hide show
  1. README.md +75 -0
README.md ADDED
@@ -0,0 +1,75 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language: en
4
+ tags:
5
+ - tiny
6
+ - slm
7
+ - small-language-model
8
+ - from-scratch
9
+ - gqa
10
+ - swiglu
11
+ - rope
12
+ - training-script
13
+ pipeline_tag: text-generation
14
+ metrics:
15
+ - accuracy
16
+ - perplexity
17
+ ---
18
+
19
+ # ram-18m Training Script
20
+
21
+ Training script for **ram-18m**: an 18,290,304-param LLaMA-style language model.
22
+
23
+ ## Architecture
24
+
25
+ | Parameter | Value |
26
+ |-----------|-------|
27
+ | d_model | 384 |
28
+ | n_heads | 6 |
29
+ | n_kv_heads | 2 (GQA) |
30
+ | head_dim | 64 |
31
+ | n_layers | 7 |
32
+ | FFN | SwiGLU, 4x (1536) |
33
+ | Norm | RMSNorm |
34
+ | Positional | RoPE (θ=10000) |
35
+ | Vocab | 8192 (BPE) |
36
+ | Tied embed/head | yes |
37
+ | **Total params** | **18,290,304** |
38
+
39
+ ## Default Training Config
40
+
41
+ - Data: FineWeb-Edu L3 (sample-100BT), ~2B tokens
42
+ - Optimizer: AdamW (β=0.9/0.95, wd=0)
43
+ - LR: 2e-4, cosine decay to 2e-5, warmup 200 steps
44
+ - Batch: 32, seq_len 512
45
+ - Steps: 12,207 (~2B tokens)
46
+ - Grad clip: 1.0
47
+
48
+ ## Usage
49
+
50
+ ```bash
51
+ # 1. Prepare data (trains BPE tokenizer + tokenizes 2B tokens)
52
+ python3 train_ram_18m.py --stage prepare
53
+
54
+ # 2. Train
55
+ python3 train_ram_18m.py --stage train
56
+
57
+ # 3. Or do both
58
+ python3 train_ram_18m.py --stage all
59
+
60
+ # 4. Eval a checkpoint
61
+ python3 train_ram_18m.py --stage eval --ckpt ckpt_step12207.pt
62
+ ```
63
+
64
+ ## Requirements
65
+
66
+ ```
67
+ pip install torch transformers datasets numpy tokenizers
68
+ ```
69
+
70
+ ## Notes
71
+
72
+ - Requested by GGUFGuy in [model-requests #30](https://huggingface.co/spaces/Compactbot/model-requests/discussions/30)
73
+ - The model is NOT trained yet — this is the script only.
74
+ - GPU recommended (RTX 3090+ for reasonable speed); CPU works but is ~50x slower.
75
+ - Checkpoints saved every 500 steps to the script directory.