Spaces:
Paused
The Beginner's Guide to Understanding an LLM Training Run
Tailored for the GLM-5.2 Baby 120M pretraining on RTX 4050
Part 1: What Do All These Numbers Mean?
When a log line appears like this:
step 52000 | train 4.1230 | val 4.1071 | lr 5.51e-04
Here is what each piece indicates:
Loss (train & val)
Loss = how wrong the model is. Lower = better.
The model attempts to predict the next word (token) in a sequence. The loss measures how surprised the model is by the correct answer. Think of it like a quiz score, but inverted: 4.10 is better than 4.30.
| Term | What it means | Example value |
|---|---|---|
| Training loss | Error on data the model is actively learning from | 4.1230 |
| Validation loss | Error on data the model has never seen (the real test) | 4.1071 |
Always trust validation loss over training loss. Training loss can go down just because the model memorizes data. Validation loss indicates if it is actually learning patterns.
What is considered "good" for a 120M model?
- Loss > 5.0 : The model is essentially guessing randomly
- Loss ~4.0-4.5 : It is learning basic grammar and common patterns
- Loss ~3.5-4.0 : It understands sentence structure well
- Loss ~3.0-3.5 : Coherent paragraphs, factual fragments
- Loss < 3.0 : Very strong for this model size (hard to achieve with 120M params)
Learning Rate (lr)
lr 5.51e-04 -> this means 0.000551
Learning rate = how big of a step the model takes when adjusting its weights after each batch.
- Too high : Model overshoots and loss spikes or goes to NaN (crashes)
- Too low : Model barely changes and training takes extremely long
- Just right : Steady decrease in loss
The training uses a cosine schedule with warmup. Here is what that looks like:
Learning Rate over time:
6e-4 | /__________________________\
| / \
| / \
| / \
6e-5 |/ \___
+----+--------+--------+--------+--------+
0 1.5k 65k 130k 195k 260k
warmup steps ->
Phase: [WARM] [--- COSINE DECAY -------] [MIN]
What this means for the training run:
- Steps 0-1,500: LR ramps up from 0 to
6e-4(warmup) - Steps 1,500-260,000: LR slowly decays from
6e-4to6e-5following a cosine curve - At step 52,000, it is at
5.51e-04β still very close to the peak. - The big improvements come later when the LR drops significantly (after ~130k steps)
Tokens per Second (tok/s)
4,259 tok/s
This is the training speed β how many tokens (approximate words) the model processes per second. Higher = faster training. An RTX 4050 pushing ~4,000-5,000 tok/s is solid for a laptop GPU.
VRAM
VRAM 4.78GB
How much GPU memory is being used. A GPU with 6GB total running at 4.78GB means there is a ~1.2GB safety buffer. If this hits 6GB, an Out of Memory crash (OOM) occurs.
Part 2: How to Tell if Training is Going Well
Signs of Healthy Training
- Val loss is trending down over thousands of steps (not every single eval, but the overall trend)
- Train loss and val loss are close together (no big gap)
- No NaN or Inf in loss (that would mean the training exploded)
- No sudden loss spikes that do not recover
Warning Signs
| Symptom | What it means | What to do |
|---|---|---|
| Val loss goes up while train loss goes down | Overfitting β memorizing instead of learning | Stop training, use more data, or add regularization |
| Val loss flat for 20,000+ steps | Plateau β model may be stuck | Often resolves when LR decays |
| Loss suddenly shoots to 100+ or NaN | Training instability | Reduce learning rate, check data for corruption |
| Train loss very noisy (swings of +-1.0) | Batch size too small | Increase gradient_accumulation_steps |
Health Check Example
Val Loss Trajectory:
Step 20k: 4.31 #######################...
Step 24k: 4.27 #####################....
Step 28k: 4.22 ####################.....
Step 30k: 4.19 ###################......
Step 36k: 4.16 ##################.......
Step 40k: 4.12 #################........ <- plateau started here
Step 50k: 4.18 ##################....... <- bouncing around
Step 52k: 4.11 #################........ <- NEW BEST! Plateau broken
Verdict: The training is healthy. A 10k-step plateau (40k-50k) is normal and typically breaks through eventually, as seen at step 52k.
Part 3: Key Concepts
Chinchilla Scaling Laws (The 20:1 Rule)
DeepMind discovered in 2022 that for optimal training, roughly 20 tokens of data per parameter should be used.
For this model:
- Parameters: ~120M
- Chinchilla-optimal data: 120M * 20 = 2.4B tokens
- Planned data: 2.4B tokens (2.0B Phase 2 + 0.4B Phase 3)
This means the model should be trained to near its full potential by the time it finishes.
Gradient Accumulation (Why batch_size 6 * grad_accum 3)
A laptop GPU can only fit 6 sequences in VRAM at once. But training works better with larger "effective" batches (more stable gradients). Gradient accumulation is a trick to achieve this:
Step 1: Process 6 sequences -> compute gradients (no weight updates yet)
Step 2: Process 6 more sequences -> add gradients to step 1
Step 3: Process 6 more sequences -> add gradients again
-> NOW update weights using all 18 sequences' worth of gradients!
Effective batch = 6 * 3 = 18 sequences * 512 tokens = 9,216 tokens per weight update
Why Mixed Precision (bfloat16) Matters
Normally, numbers in the model use 32 bits (float32). bfloat16 uses only 16 bits:
- Pro: Uses roughly half the VRAM, runs roughly 2x faster on Tensor Cores
- Con: Slightly less precise math
- Net result: Massive performance win, essential for fitting a 120M model in 6GB.
Gradient Checkpointing
Normally, the GPU stores ALL intermediate calculations during the forward pass (to use during backpropagation). With gradient checkpointing:
- GPU throws away intermediate results to save VRAM
- During backpropagation, it recomputes them on the fly
- Trade-off: ~30% slower training, but ~40% less VRAM
This is enabled and is essential for fitting the model in 6GB.
Part 4: The Cosine Schedule
Here is something critical to understand about this specific run:
The LR decay spans 260,000 steps, but this run goes to 110,000.
At step 52,000:
- It is 20% through the cosine schedule
- LR has only dropped from 6.00e-4 to 5.51e-4 (an 8% decrease)
- The model is still taking BIG learning steps
At step 110,000 (end of this run):
- It will be 42% through the cosine schedule
- LR will be around ~4.4e-4
- Model will be learning at a moderate pace
The BIGGEST gains happen in Phase 3 (steps 110k -> 260k):
- LR drops dramatically from 4.4e-4 to 6e-5
- This is where the model "settles in" and polishes its knowledge
- Early training = rough sketch, late training = fine details
This explains the plateaus. The LR is high enough that the model is "bouncing around" in the loss landscape. As LR decreases, these oscillations shrink and the loss drops more smoothly.
Part 5: Must-Read Resources (Ordered by Difficulty)
Beginner
| # | Resource | What to Learn | Link |
|---|---|---|---|
| 1 | Karpathy: "Let's build GPT: from scratch" (YouTube) | How transformers work, attention, the training loop β coded live | YouTube |
| 2 | Karpathy: "A Recipe for Training Neural Networks" (Blog) | The definitive guide for debugging training. | karpathy.github.io |
| 3 | 3Blue1Brown: "But what is a neural network?" (YouTube) | Visual intuition for how neural nets learn | YouTube |
Intermediate
| # | Resource | What to Learn | Link |
|---|---|---|---|
| 4 | Karpathy: "Let's reproduce GPT-2 (124M)" (YouTube) | Pretraining a GPT from scratch, optimized | YouTube |
| 5 | Sebastian Raschka: "Build a Large Language Model (From Scratch)" (Book + GitHub) | Full pipeline: tokenization -> training -> finetuning | GitHub |
| 6 | Chinchilla Paper (Summary) | Why 20 tokens/param matters, scaling laws | Search: "Chinchilla scaling laws explained" |
Advanced
| # | Resource | What to Learn | Link |
|---|---|---|---|
| 7 | EleutherAI: LM Evaluation Harness | Proper benchmarking beyond just val loss | GitHub |
| 8 | DeepSeek-V3 Technical Report | The MLA + MoE architecture the model is based on | arXiv |
| 9 | Sebastian Raschka: "Ahead of AI" newsletter | Weekly updates on LLM research and training techniques | Substack |
Part 6: Quick Glossary
| Term | Plain English |
|---|---|
| Token | A piece of a word. "training" -> ["train", "ing"]. Models see tokens, not words. |
| Epoch | One full pass through all training data. |
| Perplexity | e^loss β another way to express loss. Val loss 4.11 = perplexity ~61. Means "the model is choosing between ~61 equally likely next tokens." |
| Overfitting | Model memorizes training data instead of learning general patterns. Val loss goes up while train loss goes down. |
| Underfitting | Model has not learned enough yet. Both losses are still high. |
| Cosine decay | LR schedule that follows a cosine curve from high to low. |
| AdamW | The optimizer (algorithm that updates weights) used for most LLMs. |
| Gradient clipping | Caps the size of gradient updates to prevent explosions. |
| MoE | Mixture of Experts β only some "expert" sub-networks activate per token, scaling capacity without proportional compute costs. |
| MLA | Multi-Latent Attention β compresses attention using LoRA-style projections to save VRAM. |
| DSA | DeepSeek Sparse Attention β selects only the most relevant tokens to attend to. |
The single most important resource is Karpathy's "Let's reproduce GPT-2 (124M)" video. It covers the same workflow: pretraining a ~124M parameter model from scratch on a single GPU with the same optimizer, LR schedule, and training loop design. The GLM-5.2 script is heavily inspired by nanoGPT.
Training is proceeding nominally. Val loss 4.1071 at step 52k indicates solid progress.