nano-glm / docs /llm_training_guide.md
P1yansh
Reorganize directory structure, add FastAPI server and web UI
d2aafc6
|
Raw
History Blame Contribute Delete
11 kB

The Beginner's Guide to Understanding an LLM Training Run

Tailored for the GLM-5.2 Baby 120M pretraining on RTX 4050


Part 1: What Do All These Numbers Mean?

When a log line appears like this:

step 52000 | train 4.1230 | val 4.1071 | lr 5.51e-04

Here is what each piece indicates:

Loss (train & val)

Loss = how wrong the model is. Lower = better.

The model attempts to predict the next word (token) in a sequence. The loss measures how surprised the model is by the correct answer. Think of it like a quiz score, but inverted: 4.10 is better than 4.30.

Term What it means Example value
Training loss Error on data the model is actively learning from 4.1230
Validation loss Error on data the model has never seen (the real test) 4.1071

Always trust validation loss over training loss. Training loss can go down just because the model memorizes data. Validation loss indicates if it is actually learning patterns.

What is considered "good" for a 120M model?

  • Loss > 5.0 : The model is essentially guessing randomly
  • Loss ~4.0-4.5 : It is learning basic grammar and common patterns
  • Loss ~3.5-4.0 : It understands sentence structure well
  • Loss ~3.0-3.5 : Coherent paragraphs, factual fragments
  • Loss < 3.0 : Very strong for this model size (hard to achieve with 120M params)

Learning Rate (lr)

lr 5.51e-04  ->  this means 0.000551

Learning rate = how big of a step the model takes when adjusting its weights after each batch.

  • Too high : Model overshoots and loss spikes or goes to NaN (crashes)
  • Too low : Model barely changes and training takes extremely long
  • Just right : Steady decrease in loss

The training uses a cosine schedule with warmup. Here is what that looks like:

Learning Rate over time:

 6e-4 |    /__________________________\
      |   /                              \
      |  /                                \
      | /                                  \
 6e-5 |/                                    \___
      +----+--------+--------+--------+--------+
      0   1.5k    65k     130k     195k     260k
       warmup              steps ->

      Phase: [WARM] [--- COSINE DECAY -------]  [MIN]

What this means for the training run:

  • Steps 0-1,500: LR ramps up from 0 to 6e-4 (warmup)
  • Steps 1,500-260,000: LR slowly decays from 6e-4 to 6e-5 following a cosine curve
  • At step 52,000, it is at 5.51e-04 β€” still very close to the peak.
  • The big improvements come later when the LR drops significantly (after ~130k steps)

Tokens per Second (tok/s)

4,259 tok/s

This is the training speed β€” how many tokens (approximate words) the model processes per second. Higher = faster training. An RTX 4050 pushing ~4,000-5,000 tok/s is solid for a laptop GPU.

VRAM

VRAM 4.78GB

How much GPU memory is being used. A GPU with 6GB total running at 4.78GB means there is a ~1.2GB safety buffer. If this hits 6GB, an Out of Memory crash (OOM) occurs.


Part 2: How to Tell if Training is Going Well

Signs of Healthy Training

  1. Val loss is trending down over thousands of steps (not every single eval, but the overall trend)
  2. Train loss and val loss are close together (no big gap)
  3. No NaN or Inf in loss (that would mean the training exploded)
  4. No sudden loss spikes that do not recover

Warning Signs

Symptom What it means What to do
Val loss goes up while train loss goes down Overfitting β€” memorizing instead of learning Stop training, use more data, or add regularization
Val loss flat for 20,000+ steps Plateau β€” model may be stuck Often resolves when LR decays
Loss suddenly shoots to 100+ or NaN Training instability Reduce learning rate, check data for corruption
Train loss very noisy (swings of +-1.0) Batch size too small Increase gradient_accumulation_steps

Health Check Example

Val Loss Trajectory:
  Step 20k: 4.31 #######################...
  Step 24k: 4.27 #####################....
  Step 28k: 4.22 ####################.....
  Step 30k: 4.19 ###################......
  Step 36k: 4.16 ##################.......
  Step 40k: 4.12 #################........  <- plateau started here
  Step 50k: 4.18 ##################.......  <- bouncing around
  Step 52k: 4.11 #################........  <- NEW BEST! Plateau broken

Verdict: The training is healthy. A 10k-step plateau (40k-50k) is normal and typically breaks through eventually, as seen at step 52k.


Part 3: Key Concepts

Chinchilla Scaling Laws (The 20:1 Rule)

DeepMind discovered in 2022 that for optimal training, roughly 20 tokens of data per parameter should be used.

For this model:

  • Parameters: ~120M
  • Chinchilla-optimal data: 120M * 20 = 2.4B tokens
  • Planned data: 2.4B tokens (2.0B Phase 2 + 0.4B Phase 3)

This means the model should be trained to near its full potential by the time it finishes.

Gradient Accumulation (Why batch_size 6 * grad_accum 3)

A laptop GPU can only fit 6 sequences in VRAM at once. But training works better with larger "effective" batches (more stable gradients). Gradient accumulation is a trick to achieve this:

Step 1: Process 6 sequences -> compute gradients (no weight updates yet)
Step 2: Process 6 more sequences -> add gradients to step 1
Step 3: Process 6 more sequences -> add gradients again
-> NOW update weights using all 18 sequences' worth of gradients!

Effective batch = 6 * 3 = 18 sequences * 512 tokens = 9,216 tokens per weight update

Why Mixed Precision (bfloat16) Matters

Normally, numbers in the model use 32 bits (float32). bfloat16 uses only 16 bits:

  • Pro: Uses roughly half the VRAM, runs roughly 2x faster on Tensor Cores
  • Con: Slightly less precise math
  • Net result: Massive performance win, essential for fitting a 120M model in 6GB.

Gradient Checkpointing

Normally, the GPU stores ALL intermediate calculations during the forward pass (to use during backpropagation). With gradient checkpointing:

  • GPU throws away intermediate results to save VRAM
  • During backpropagation, it recomputes them on the fly
  • Trade-off: ~30% slower training, but ~40% less VRAM

This is enabled and is essential for fitting the model in 6GB.


Part 4: The Cosine Schedule

Here is something critical to understand about this specific run:

The LR decay spans 260,000 steps, but this run goes to 110,000.

At step 52,000:
  - It is 20% through the cosine schedule
  - LR has only dropped from 6.00e-4 to 5.51e-4 (an 8% decrease)
  - The model is still taking BIG learning steps

At step 110,000 (end of this run):
  - It will be 42% through the cosine schedule
  - LR will be around ~4.4e-4
  - Model will be learning at a moderate pace

The BIGGEST gains happen in Phase 3 (steps 110k -> 260k):
  - LR drops dramatically from 4.4e-4 to 6e-5
  - This is where the model "settles in" and polishes its knowledge
  - Early training = rough sketch, late training = fine details

This explains the plateaus. The LR is high enough that the model is "bouncing around" in the loss landscape. As LR decreases, these oscillations shrink and the loss drops more smoothly.


Part 5: Must-Read Resources (Ordered by Difficulty)

Beginner

# Resource What to Learn Link
1 Karpathy: "Let's build GPT: from scratch" (YouTube) How transformers work, attention, the training loop β€” coded live YouTube
2 Karpathy: "A Recipe for Training Neural Networks" (Blog) The definitive guide for debugging training. karpathy.github.io
3 3Blue1Brown: "But what is a neural network?" (YouTube) Visual intuition for how neural nets learn YouTube

Intermediate

# Resource What to Learn Link
4 Karpathy: "Let's reproduce GPT-2 (124M)" (YouTube) Pretraining a GPT from scratch, optimized YouTube
5 Sebastian Raschka: "Build a Large Language Model (From Scratch)" (Book + GitHub) Full pipeline: tokenization -> training -> finetuning GitHub
6 Chinchilla Paper (Summary) Why 20 tokens/param matters, scaling laws Search: "Chinchilla scaling laws explained"

Advanced

# Resource What to Learn Link
7 EleutherAI: LM Evaluation Harness Proper benchmarking beyond just val loss GitHub
8 DeepSeek-V3 Technical Report The MLA + MoE architecture the model is based on arXiv
9 Sebastian Raschka: "Ahead of AI" newsletter Weekly updates on LLM research and training techniques Substack

Part 6: Quick Glossary

Term Plain English
Token A piece of a word. "training" -> ["train", "ing"]. Models see tokens, not words.
Epoch One full pass through all training data.
Perplexity e^loss β€” another way to express loss. Val loss 4.11 = perplexity ~61. Means "the model is choosing between ~61 equally likely next tokens."
Overfitting Model memorizes training data instead of learning general patterns. Val loss goes up while train loss goes down.
Underfitting Model has not learned enough yet. Both losses are still high.
Cosine decay LR schedule that follows a cosine curve from high to low.
AdamW The optimizer (algorithm that updates weights) used for most LLMs.
Gradient clipping Caps the size of gradient updates to prevent explosions.
MoE Mixture of Experts β€” only some "expert" sub-networks activate per token, scaling capacity without proportional compute costs.
MLA Multi-Latent Attention β€” compresses attention using LoRA-style projections to save VRAM.
DSA DeepSeek Sparse Attention β€” selects only the most relevant tokens to attend to.

The single most important resource is Karpathy's "Let's reproduce GPT-2 (124M)" video. It covers the same workflow: pretraining a ~124M parameter model from scratch on a single GPU with the same optimizer, LR schedule, and training loop design. The GLM-5.2 script is heavily inspired by nanoGPT.


Training is proceeding nominally. Val loss 4.1071 at step 52k indicates solid progress.