File size: 10,991 Bytes
8e382fd
0fcfdf0
8e382fd
0fcfdf0
 
 
 
 
8e382fd
0fcfdf0
 
 
 
 
8e382fd
0fcfdf0
8e382fd
0fcfdf0
 
 
8e382fd
0fcfdf0
8e382fd
0fcfdf0
 
 
 
 
8e382fd
0fcfdf0
8e382fd
0fcfdf0
8e382fd
 
 
 
 
0fcfdf0
8e382fd
0fcfdf0
 
8e382fd
0fcfdf0
 
 
 
8e382fd
 
 
0fcfdf0
8e382fd
0fcfdf0
 
 
 
8e382fd
0fcfdf0
 
 
 
 
 
8e382fd
0fcfdf0
8e382fd
0fcfdf0
 
8e382fd
 
 
 
0fcfdf0
 
8e382fd
0fcfdf0
 
 
 
 
8e382fd
0fcfdf0
8e382fd
0fcfdf0
 
 
 
 
8e382fd
0fcfdf0
 
 
 
 
8e382fd
0fcfdf0
 
 
 
8e382fd
0fcfdf0
8e382fd
0fcfdf0
 
 
 
8e382fd
0fcfdf0
8e382fd
0fcfdf0
8e382fd
0fcfdf0
 
 
8e382fd
 
 
 
 
 
 
 
0fcfdf0
 
8e382fd
0fcfdf0
 
 
8e382fd
0fcfdf0
8e382fd
0fcfdf0
8e382fd
0fcfdf0
8e382fd
0fcfdf0
8e382fd
 
0fcfdf0
8e382fd
0fcfdf0
8e382fd
0fcfdf0
8e382fd
0fcfdf0
 
8e382fd
 
 
 
0fcfdf0
 
8e382fd
0fcfdf0
8e382fd
0fcfdf0
 
8e382fd
0fcfdf0
8e382fd
0fcfdf0
8e382fd
0fcfdf0
 
 
 
8e382fd
0fcfdf0
8e382fd
0fcfdf0
 
 
8e382fd
0fcfdf0
8e382fd
0fcfdf0
 
8e382fd
0fcfdf0
 
8e382fd
 
0fcfdf0
 
 
8e382fd
0fcfdf0
 
 
8e382fd
 
0fcfdf0
8e382fd
0fcfdf0
 
 
8e382fd
0fcfdf0
 
 
8e382fd
0fcfdf0
8e382fd
0fcfdf0
8e382fd
0fcfdf0
8e382fd
 
0fcfdf0
 
8e382fd
0fcfdf0
8e382fd
0fcfdf0
8e382fd
 
0fcfdf0
 
8e382fd
0fcfdf0
8e382fd
0fcfdf0
8e382fd
 
0fcfdf0
 
 
 
 
 
 
 
8e382fd
 
 
 
 
 
 
 
 
 
 
0fcfdf0
 
 
 
8e382fd
0fcfdf0
 
 
8e382fd
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
# The Beginner's Guide to Understanding an LLM Training Run

*Tailored for the GLM-5.2 Baby 120M pretraining on RTX 4050*

---

## Part 1: What Do All These Numbers Mean?

When a log line appears like this:

```
step 52000 | train 4.1230 | val 4.1071 | lr 5.51e-04
```

Here is what each piece indicates:

### Loss (train & val)

**Loss** = how wrong the model is. Lower = better.

The model attempts to predict the next word (token) in a sequence. The loss measures how surprised the model is by the correct answer. Think of it like a quiz score, but inverted: **4.10 is better than 4.30**.

| Term | What it means | Example value |
|------|--------------|------------|
| **Training loss** | Error on data the model is actively learning from | `4.1230` |
| **Validation loss** | Error on data the model has **never seen** (the real test) | `4.1071` |

> [!IMPORTANT]
> **Always trust validation loss over training loss.** Training loss can go down just because the model memorizes data. Validation loss indicates if it is actually *learning patterns*.

#### What is considered "good" for a 120M model?

- **Loss > 5.0** : The model is essentially guessing randomly
- **Loss ~4.0-4.5** : It is learning basic grammar and common patterns
- **Loss ~3.5-4.0** : It understands sentence structure well
- **Loss ~3.0-3.5** : Coherent paragraphs, factual fragments
- **Loss < 3.0** : Very strong for this model size (hard to achieve with 120M params)

### Learning Rate (lr)

```
lr 5.51e-04  ->  this means 0.000551
```

**Learning rate** = how big of a step the model takes when adjusting its weights after each batch.

- **Too high** : Model overshoots and loss spikes or goes to NaN (crashes)
- **Too low** : Model barely changes and training takes extremely long
- **Just right** : Steady decrease in loss

The training uses a **cosine schedule with warmup**. Here is what that looks like:

```
Learning Rate over time:

 6e-4 |    /__________________________\
      |   /                              \
      |  /                                \
      | /                                  \
 6e-5 |/                                    \___
      +----+--------+--------+--------+--------+
      0   1.5k    65k     130k     195k     260k
       warmup              steps ->

      Phase: [WARM] [--- COSINE DECAY -------]  [MIN]
```

**What this means for the training run:**
- Steps 0-1,500: LR ramps up from 0 to `6e-4` (warmup)
- Steps 1,500-260,000: LR slowly decays from `6e-4` to `6e-5` following a cosine curve
- At step 52,000, it is at `5.51e-04` β€” still very close to the peak.
- **The big improvements come later** when the LR drops significantly (after ~130k steps)

### Tokens per Second (tok/s)

```
4,259 tok/s
```

This is the training speed β€” how many tokens (approximate words) the model processes per second. Higher = faster training. An RTX 4050 pushing ~4,000-5,000 tok/s is solid for a laptop GPU.

### VRAM

```
VRAM 4.78GB
```

How much GPU memory is being used. A GPU with 6GB total running at 4.78GB means there is a ~1.2GB safety buffer. If this hits 6GB, an Out of Memory crash (OOM) occurs.

---

## Part 2: How to Tell if Training is Going Well

### Signs of Healthy Training

1. **Val loss is trending down over thousands of steps** (not every single eval, but the overall trend)
2. **Train loss and val loss are close together** (no big gap)
3. **No NaN or Inf in loss** (that would mean the training exploded)
4. **No sudden loss spikes** that do not recover

### Warning Signs

| Symptom | What it means | What to do |
|---------|--------------|------------|
| Val loss goes **up** while train loss goes **down** | **Overfitting** β€” memorizing instead of learning | Stop training, use more data, or add regularization |
| Val loss **flat for 20,000+ steps** | **Plateau** β€” model may be stuck | Often resolves when LR decays |
| Loss suddenly shoots to **100+** or **NaN** | **Training instability** | Reduce learning rate, check data for corruption |
| Train loss **very noisy** (swings of +-1.0) | **Batch size too small** | Increase `gradient_accumulation_steps` |

### Health Check Example

```
Val Loss Trajectory:
  Step 20k: 4.31 #######################...
  Step 24k: 4.27 #####################....
  Step 28k: 4.22 ####################.....
  Step 30k: 4.19 ###################......
  Step 36k: 4.16 ##################.......
  Step 40k: 4.12 #################........  <- plateau started here
  Step 50k: 4.18 ##################.......  <- bouncing around
  Step 52k: 4.11 #################........  <- NEW BEST! Plateau broken
```

**Verdict: The training is healthy.** A 10k-step plateau (40k-50k) is normal and typically breaks through eventually, as seen at step 52k.

---

## Part 3: Key Concepts

### Chinchilla Scaling Laws (The 20:1 Rule)

DeepMind discovered in 2022 that for optimal training, roughly **20 tokens of data per parameter** should be used.

**For this model:**
- Parameters: ~120M
- Chinchilla-optimal data: 120M * 20 = **2.4B tokens**
- Planned data: **2.4B tokens** (2.0B Phase 2 + 0.4B Phase 3)

This means the model should be trained to near its full potential by the time it finishes.

### Gradient Accumulation (Why batch_size 6 * grad_accum 3)

A laptop GPU can only fit 6 sequences in VRAM at once. But training works better with larger "effective" batches (more stable gradients). **Gradient accumulation** is a trick to achieve this:

```
Step 1: Process 6 sequences -> compute gradients (no weight updates yet)
Step 2: Process 6 more sequences -> add gradients to step 1
Step 3: Process 6 more sequences -> add gradients again
-> NOW update weights using all 18 sequences' worth of gradients!
```

Effective batch = `6 * 3 = 18 sequences * 512 tokens = 9,216 tokens per weight update`

### Why Mixed Precision (bfloat16) Matters

Normally, numbers in the model use 32 bits (float32). **bfloat16** uses only 16 bits:
- **Pro:** Uses roughly half the VRAM, runs roughly 2x faster on Tensor Cores
- **Con:** Slightly less precise math
- **Net result:** Massive performance win, essential for fitting a 120M model in 6GB.

### Gradient Checkpointing

Normally, the GPU stores ALL intermediate calculations during the forward pass (to use during backpropagation). With gradient checkpointing:
- GPU **throws away** intermediate results to save VRAM
- During backpropagation, it **recomputes** them on the fly
- **Trade-off:** ~30% slower training, but ~40% less VRAM

This is enabled and is essential for fitting the model in 6GB.

---

## Part 4: The Cosine Schedule

Here is something critical to understand about this specific run:

```
The LR decay spans 260,000 steps, but this run goes to 110,000.

At step 52,000:
  - It is 20% through the cosine schedule
  - LR has only dropped from 6.00e-4 to 5.51e-4 (an 8% decrease)
  - The model is still taking BIG learning steps

At step 110,000 (end of this run):
  - It will be 42% through the cosine schedule
  - LR will be around ~4.4e-4
  - Model will be learning at a moderate pace

The BIGGEST gains happen in Phase 3 (steps 110k -> 260k):
  - LR drops dramatically from 4.4e-4 to 6e-5
  - This is where the model "settles in" and polishes its knowledge
  - Early training = rough sketch, late training = fine details
```

> [!TIP]
> **This explains the plateaus.** The LR is high enough that the model is "bouncing around" in the loss landscape. As LR decreases, these oscillations shrink and the loss drops more smoothly.

---

## Part 5: Must-Read Resources (Ordered by Difficulty)

### Beginner

| # | Resource | What to Learn | Link |
|---|----------|-------------------|------|
| 1 | **Karpathy: "Let's build GPT: from scratch"** (YouTube) | How transformers work, attention, the training loop β€” coded live | [YouTube](https://www.youtube.com/watch?v=kCc8FmEb1nY) |
| 2 | **Karpathy: "A Recipe for Training Neural Networks"** (Blog) | The definitive guide for debugging training. | [karpathy.github.io](https://karpathy.github.io/2019/04/25/recipe/) |
| 3 | **3Blue1Brown: "But what is a neural network?"** (YouTube) | Visual intuition for how neural nets learn | [YouTube](https://www.youtube.com/watch?v=aircAruvnKk) |

### Intermediate

| # | Resource | What to Learn | Link |
|---|----------|-------------------|------|
| 4 | **Karpathy: "Let's reproduce GPT-2 (124M)"** (YouTube) | Pretraining a GPT from scratch, optimized | [YouTube](https://www.youtube.com/watch?v=l8pRSuU81PU) |
| 5 | **Sebastian Raschka: "Build a Large Language Model (From Scratch)"** (Book + GitHub) | Full pipeline: tokenization -> training -> finetuning | [GitHub](https://github.com/rasbt/LLMs-from-scratch) |
| 6 | **Chinchilla Paper (Summary)** | Why 20 tokens/param matters, scaling laws | Search: "Chinchilla scaling laws explained" |

### Advanced

| # | Resource | What to Learn | Link |
|---|----------|-------------------|------|
| 7 | **EleutherAI: LM Evaluation Harness** | Proper benchmarking beyond just val loss | [GitHub](https://github.com/EleutherAI/lm-evaluation-harness) |
| 8 | **DeepSeek-V3 Technical Report** | The MLA + MoE architecture the model is based on | [arXiv](https://arxiv.org/abs/2412.19437) |
| 9 | **Sebastian Raschka: "Ahead of AI" newsletter** | Weekly updates on LLM research and training techniques | [Substack](https://magazine.sebastianraschka.com/) |

---

## Part 6: Quick Glossary

| Term | Plain English |
|------|--------------|
| **Token** | A piece of a word. "training" -> ["train", "ing"]. Models see tokens, not words. |
| **Epoch** | One full pass through all training data. |
| **Perplexity** | `e^loss` β€” another way to express loss. Val loss 4.11 = perplexity ~61. Means "the model is choosing between ~61 equally likely next tokens." |
| **Overfitting** | Model memorizes training data instead of learning general patterns. Val loss goes up while train loss goes down. |
| **Underfitting** | Model has not learned enough yet. Both losses are still high. |
| **Cosine decay** | LR schedule that follows a cosine curve from high to low. |
| **AdamW** | The optimizer (algorithm that updates weights) used for most LLMs. |
| **Gradient clipping** | Caps the size of gradient updates to prevent explosions. |
| **MoE** | Mixture of Experts β€” only some "expert" sub-networks activate per token, scaling capacity without proportional compute costs. |
| **MLA** | Multi-Latent Attention β€” compresses attention using LoRA-style projections to save VRAM. |
| **DSA** | DeepSeek Sparse Attention β€” selects only the most relevant tokens to attend to. |

---

> [!NOTE]
> The single most important resource is Karpathy's ["Let's reproduce GPT-2 (124M)"](https://www.youtube.com/watch?v=l8pRSuU81PU) video. It covers the same workflow: pretraining a ~124M parameter model from scratch on a single GPU with the same optimizer, LR schedule, and training loop design. The GLM-5.2 script is heavily inspired by nanoGPT.

---

*Training is proceeding nominally. Val loss 4.1071 at step 52k indicates solid progress.*