Spaces:
Paused
Paused
| # The Beginner's Guide to Understanding an LLM Training Run | |
| *Tailored for the GLM-5.2 Baby 120M pretraining on RTX 4050* | |
| --- | |
| ## Part 1: What Do All These Numbers Mean? | |
| When a log line appears like this: | |
| ``` | |
| step 52000 | train 4.1230 | val 4.1071 | lr 5.51e-04 | |
| ``` | |
| Here is what each piece indicates: | |
| ### Loss (train & val) | |
| **Loss** = how wrong the model is. Lower = better. | |
| The model attempts to predict the next word (token) in a sequence. The loss measures how surprised the model is by the correct answer. Think of it like a quiz score, but inverted: **4.10 is better than 4.30**. | |
| | Term | What it means | Example value | | |
| |------|--------------|------------| | |
| | **Training loss** | Error on data the model is actively learning from | `4.1230` | | |
| | **Validation loss** | Error on data the model has **never seen** (the real test) | `4.1071` | | |
| > [!IMPORTANT] | |
| > **Always trust validation loss over training loss.** Training loss can go down just because the model memorizes data. Validation loss indicates if it is actually *learning patterns*. | |
| #### What is considered "good" for a 120M model? | |
| - **Loss > 5.0** : The model is essentially guessing randomly | |
| - **Loss ~4.0-4.5** : It is learning basic grammar and common patterns | |
| - **Loss ~3.5-4.0** : It understands sentence structure well | |
| - **Loss ~3.0-3.5** : Coherent paragraphs, factual fragments | |
| - **Loss < 3.0** : Very strong for this model size (hard to achieve with 120M params) | |
| ### Learning Rate (lr) | |
| ``` | |
| lr 5.51e-04 -> this means 0.000551 | |
| ``` | |
| **Learning rate** = how big of a step the model takes when adjusting its weights after each batch. | |
| - **Too high** : Model overshoots and loss spikes or goes to NaN (crashes) | |
| - **Too low** : Model barely changes and training takes extremely long | |
| - **Just right** : Steady decrease in loss | |
| The training uses a **cosine schedule with warmup**. Here is what that looks like: | |
| ``` | |
| Learning Rate over time: | |
| 6e-4 | /__________________________\ | |
| | / \ | |
| | / \ | |
| | / \ | |
| 6e-5 |/ \___ | |
| +----+--------+--------+--------+--------+ | |
| 0 1.5k 65k 130k 195k 260k | |
| warmup steps -> | |
| Phase: [WARM] [--- COSINE DECAY -------] [MIN] | |
| ``` | |
| **What this means for the training run:** | |
| - Steps 0-1,500: LR ramps up from 0 to `6e-4` (warmup) | |
| - Steps 1,500-260,000: LR slowly decays from `6e-4` to `6e-5` following a cosine curve | |
| - At step 52,000, it is at `5.51e-04` — still very close to the peak. | |
| - **The big improvements come later** when the LR drops significantly (after ~130k steps) | |
| ### Tokens per Second (tok/s) | |
| ``` | |
| 4,259 tok/s | |
| ``` | |
| This is the training speed — how many tokens (approximate words) the model processes per second. Higher = faster training. An RTX 4050 pushing ~4,000-5,000 tok/s is solid for a laptop GPU. | |
| ### VRAM | |
| ``` | |
| VRAM 4.78GB | |
| ``` | |
| How much GPU memory is being used. A GPU with 6GB total running at 4.78GB means there is a ~1.2GB safety buffer. If this hits 6GB, an Out of Memory crash (OOM) occurs. | |
| --- | |
| ## Part 2: How to Tell if Training is Going Well | |
| ### Signs of Healthy Training | |
| 1. **Val loss is trending down over thousands of steps** (not every single eval, but the overall trend) | |
| 2. **Train loss and val loss are close together** (no big gap) | |
| 3. **No NaN or Inf in loss** (that would mean the training exploded) | |
| 4. **No sudden loss spikes** that do not recover | |
| ### Warning Signs | |
| | Symptom | What it means | What to do | | |
| |---------|--------------|------------| | |
| | Val loss goes **up** while train loss goes **down** | **Overfitting** — memorizing instead of learning | Stop training, use more data, or add regularization | | |
| | Val loss **flat for 20,000+ steps** | **Plateau** — model may be stuck | Often resolves when LR decays | | |
| | Loss suddenly shoots to **100+** or **NaN** | **Training instability** | Reduce learning rate, check data for corruption | | |
| | Train loss **very noisy** (swings of +-1.0) | **Batch size too small** | Increase `gradient_accumulation_steps` | | |
| ### Health Check Example | |
| ``` | |
| Val Loss Trajectory: | |
| Step 20k: 4.31 #######################... | |
| Step 24k: 4.27 #####################.... | |
| Step 28k: 4.22 ####################..... | |
| Step 30k: 4.19 ###################...... | |
| Step 36k: 4.16 ##################....... | |
| Step 40k: 4.12 #################........ <- plateau started here | |
| Step 50k: 4.18 ##################....... <- bouncing around | |
| Step 52k: 4.11 #################........ <- NEW BEST! Plateau broken | |
| ``` | |
| **Verdict: The training is healthy.** A 10k-step plateau (40k-50k) is normal and typically breaks through eventually, as seen at step 52k. | |
| --- | |
| ## Part 3: Key Concepts | |
| ### Chinchilla Scaling Laws (The 20:1 Rule) | |
| DeepMind discovered in 2022 that for optimal training, roughly **20 tokens of data per parameter** should be used. | |
| **For this model:** | |
| - Parameters: ~120M | |
| - Chinchilla-optimal data: 120M * 20 = **2.4B tokens** | |
| - Planned data: **2.4B tokens** (2.0B Phase 2 + 0.4B Phase 3) | |
| This means the model should be trained to near its full potential by the time it finishes. | |
| ### Gradient Accumulation (Why batch_size 6 * grad_accum 3) | |
| A laptop GPU can only fit 6 sequences in VRAM at once. But training works better with larger "effective" batches (more stable gradients). **Gradient accumulation** is a trick to achieve this: | |
| ``` | |
| Step 1: Process 6 sequences -> compute gradients (no weight updates yet) | |
| Step 2: Process 6 more sequences -> add gradients to step 1 | |
| Step 3: Process 6 more sequences -> add gradients again | |
| -> NOW update weights using all 18 sequences' worth of gradients! | |
| ``` | |
| Effective batch = `6 * 3 = 18 sequences * 512 tokens = 9,216 tokens per weight update` | |
| ### Why Mixed Precision (bfloat16) Matters | |
| Normally, numbers in the model use 32 bits (float32). **bfloat16** uses only 16 bits: | |
| - **Pro:** Uses roughly half the VRAM, runs roughly 2x faster on Tensor Cores | |
| - **Con:** Slightly less precise math | |
| - **Net result:** Massive performance win, essential for fitting a 120M model in 6GB. | |
| ### Gradient Checkpointing | |
| Normally, the GPU stores ALL intermediate calculations during the forward pass (to use during backpropagation). With gradient checkpointing: | |
| - GPU **throws away** intermediate results to save VRAM | |
| - During backpropagation, it **recomputes** them on the fly | |
| - **Trade-off:** ~30% slower training, but ~40% less VRAM | |
| This is enabled and is essential for fitting the model in 6GB. | |
| --- | |
| ## Part 4: The Cosine Schedule | |
| Here is something critical to understand about this specific run: | |
| ``` | |
| The LR decay spans 260,000 steps, but this run goes to 110,000. | |
| At step 52,000: | |
| - It is 20% through the cosine schedule | |
| - LR has only dropped from 6.00e-4 to 5.51e-4 (an 8% decrease) | |
| - The model is still taking BIG learning steps | |
| At step 110,000 (end of this run): | |
| - It will be 42% through the cosine schedule | |
| - LR will be around ~4.4e-4 | |
| - Model will be learning at a moderate pace | |
| The BIGGEST gains happen in Phase 3 (steps 110k -> 260k): | |
| - LR drops dramatically from 4.4e-4 to 6e-5 | |
| - This is where the model "settles in" and polishes its knowledge | |
| - Early training = rough sketch, late training = fine details | |
| ``` | |
| > [!TIP] | |
| > **This explains the plateaus.** The LR is high enough that the model is "bouncing around" in the loss landscape. As LR decreases, these oscillations shrink and the loss drops more smoothly. | |
| --- | |
| ## Part 5: Must-Read Resources (Ordered by Difficulty) | |
| ### Beginner | |
| | # | Resource | What to Learn | Link | | |
| |---|----------|-------------------|------| | |
| | 1 | **Karpathy: "Let's build GPT: from scratch"** (YouTube) | How transformers work, attention, the training loop — coded live | [YouTube](https://www.youtube.com/watch?v=kCc8FmEb1nY) | | |
| | 2 | **Karpathy: "A Recipe for Training Neural Networks"** (Blog) | The definitive guide for debugging training. | [karpathy.github.io](https://karpathy.github.io/2019/04/25/recipe/) | | |
| | 3 | **3Blue1Brown: "But what is a neural network?"** (YouTube) | Visual intuition for how neural nets learn | [YouTube](https://www.youtube.com/watch?v=aircAruvnKk) | | |
| ### Intermediate | |
| | # | Resource | What to Learn | Link | | |
| |---|----------|-------------------|------| | |
| | 4 | **Karpathy: "Let's reproduce GPT-2 (124M)"** (YouTube) | Pretraining a GPT from scratch, optimized | [YouTube](https://www.youtube.com/watch?v=l8pRSuU81PU) | | |
| | 5 | **Sebastian Raschka: "Build a Large Language Model (From Scratch)"** (Book + GitHub) | Full pipeline: tokenization -> training -> finetuning | [GitHub](https://github.com/rasbt/LLMs-from-scratch) | | |
| | 6 | **Chinchilla Paper (Summary)** | Why 20 tokens/param matters, scaling laws | Search: "Chinchilla scaling laws explained" | | |
| ### Advanced | |
| | # | Resource | What to Learn | Link | | |
| |---|----------|-------------------|------| | |
| | 7 | **EleutherAI: LM Evaluation Harness** | Proper benchmarking beyond just val loss | [GitHub](https://github.com/EleutherAI/lm-evaluation-harness) | | |
| | 8 | **DeepSeek-V3 Technical Report** | The MLA + MoE architecture the model is based on | [arXiv](https://arxiv.org/abs/2412.19437) | | |
| | 9 | **Sebastian Raschka: "Ahead of AI" newsletter** | Weekly updates on LLM research and training techniques | [Substack](https://magazine.sebastianraschka.com/) | | |
| --- | |
| ## Part 6: Quick Glossary | |
| | Term | Plain English | | |
| |------|--------------| | |
| | **Token** | A piece of a word. "training" -> ["train", "ing"]. Models see tokens, not words. | | |
| | **Epoch** | One full pass through all training data. | | |
| | **Perplexity** | `e^loss` — another way to express loss. Val loss 4.11 = perplexity ~61. Means "the model is choosing between ~61 equally likely next tokens." | | |
| | **Overfitting** | Model memorizes training data instead of learning general patterns. Val loss goes up while train loss goes down. | | |
| | **Underfitting** | Model has not learned enough yet. Both losses are still high. | | |
| | **Cosine decay** | LR schedule that follows a cosine curve from high to low. | | |
| | **AdamW** | The optimizer (algorithm that updates weights) used for most LLMs. | | |
| | **Gradient clipping** | Caps the size of gradient updates to prevent explosions. | | |
| | **MoE** | Mixture of Experts — only some "expert" sub-networks activate per token, scaling capacity without proportional compute costs. | | |
| | **MLA** | Multi-Latent Attention — compresses attention using LoRA-style projections to save VRAM. | | |
| | **DSA** | DeepSeek Sparse Attention — selects only the most relevant tokens to attend to. | | |
| --- | |
| > [!NOTE] | |
| > The single most important resource is Karpathy's ["Let's reproduce GPT-2 (124M)"](https://www.youtube.com/watch?v=l8pRSuU81PU) video. It covers the same workflow: pretraining a ~124M parameter model from scratch on a single GPU with the same optimizer, LR schedule, and training loop design. The GLM-5.2 script is heavily inspired by nanoGPT. | |
| --- | |
| *Training is proceeding nominally. Val loss 4.1071 at step 52k indicates solid progress.* | |