File size: 13,538 Bytes
3f8ea29
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
# CampusGPT Uganda β€” Learning Roadmap & One-Day Mastery Checklist
## Complete Unsloth Fine-Tuning Ecosystem in One Project

---

## Your Learning Journey

By building CampusGPT Uganda, you have touched every major concept in the
Unsloth ecosystem. This document is your personal mastery checklist and
study guide.

---

## Phase 1: Foundations (Morning β€” 2 hours)

### βœ… Concept 1: What is Unsloth?

**You learned:**
- Unsloth is a Python library for faster LLM fine-tuning
- It achieves 2Γ— speed + 80% memory reduction through:
  - Custom CUDA kernels (C-level code for attention + FFN)
  - Manual backpropagation (no storing redundant activations)
  - Smart memory layout (packed data, minimal fragmentation)
- Same accuracy as HuggingFace β€” just faster

**Key file:** `training/finetune.py` β€” top docstring + `load_model_and_tokenizer()`

**Self-test:** Can you explain why Unsloth is faster to a classmate?

---

### βœ… Concept 2: LoRA β€” Low-Rank Adaptation

**You learned:**
- Full fine-tuning modifies all 8 billion parameters β†’ impractical
- LoRA adds two small matrices A (mΓ—r) and B (rΓ—n) to each weight
- Only A and B are trained (1–5% of total parameters)
- The adapted weight: W' = W + (A Γ— B) Γ— (alpha/r)
- r=16 is the recommended default for instruction tuning

**Key file:** `training/finetune.py` β€” `add_lora_adapters()`

**Formula to remember:**
```
W' = W_frozen + (A_trainable Γ— B_trainable) Γ— scaling_factor
```

**Self-test:** Why does LoRA use TWO matrices (A and B) instead of one?

*Answer: A single low-rank update matrix is the product of two smaller matrices.
The product AΓ—B has rank ≀ r, achieving compression. A single matrix of rank r
cannot be parameterised as efficiently.*

---

### βœ… Concept 3: QLoRA β€” Quantised LoRA

**You learned:**
- Quantisation stores weights in fewer bits: float32 (4 bytes) β†’ int4 (0.5 bytes)
- QLoRA = 4-bit quantised base model + full-precision LoRA adapters
- Llama-3 8B: 32 GB (fp32) β†’ ~5 GB (QLoRA) β€” fits on a T4 GPU!
- NF4 (NormalFloat4) is used β€” better distribution for neural network weights
- ~1-2% quality loss vs full fine-tuning β€” totally acceptable for most tasks

**Key file:** `training/finetune.py` β€” `load_in_4bit=True` parameter

**Memory calculation:**
```
Model size (GB) β‰ˆ (parameters Γ— bits) / (8 Γ— 1024Β³)
Llama 8B in 4-bit β‰ˆ (8 Γ— 10⁹ Γ— 4) / (8 Γ— 1024Β³) β‰ˆ 4 GB
```

---

### βœ… Concept 4: PEFT β€” Parameter-Efficient Fine-Tuning

**You learned:**
- PEFT is the HuggingFace library managing adapters
- `get_peft_model()` attaches LoRA to the base model
- After calling it: base weights FROZEN, only adapter weights trainable
- Adapter weights are ~10–100 MB (vs full model ~6–16 GB)

**Key file:** `training/finetune.py` β€” `add_lora_adapters()`

---

## Phase 2: Data Engineering (Mid-Morning β€” 2 hours)

### βœ… Concept 5: Dataset Formats

**You learned the 4 major formats:**

| Format | Fields | Use Case |
|--------|--------|----------|
| JSONL | question, answer | Simple storage |
| Alpaca | instruction, input, output | General instruction tuning |
| OpenAI Messages | [{role, content}] | Chat models (MOST COMMON) |
| ShareGPT | conversations:[{from, value}] | Multi-turn chat |

**Key file:** `data/generate_dataset.py` β€” all `to_*_format()` functions

**Self-test:** Convert this Q&A to all 4 formats by hand:
- Q: "What is the attendance policy?"
- A: "Students must attend 75% of lectures."

---

### βœ… Concept 6: Tokenisation

**You learned:**
- Text β†’ Tokens β†’ Input IDs (integers)
- "MUST university" β†’ ["MUST", " university"] β†’ [123, 456]
- Every model has its OWN vocabulary β€” ALWAYS use the model's tokeniser
- Chat templates format conversations with model-specific special tokens
- `apply_chat_template()` handles this automatically

**Key file:** `training/finetune.py` β€” `load_and_prepare_dataset()`

**Qwen chat template example:**
```
<|im_start|>system
You are CampusGPT...<|im_end|>
<|im_start|>user
What are the fees?<|im_end|>
<|im_start|>assistant
The fees are...<|im_end|>
```

---

### βœ… Concept 7: Synthetic Data Generation

**You learned:**
- Read PDFs β†’ Extract text β†’ Chunk β†’ LLM generates QA pairs
- Deduplication with Jaccard similarity (n-gram overlap)
- Quality filtering (minimum length, must be a question, no non-answers)
- Key insight: LLMs can generate their own training data!

**Key file:** `data/synthetic_generator.py`

**Pipeline:**
```
PDF β†’ pypdf β†’ chunks of 500 words β†’ LLM prompt β†’ JSON output β†’ dedup β†’ filter β†’ JSONL
```

---

## Phase 3: RAG System (Late Morning β€” 2 hours)

### βœ… Concept 8: Vector Embeddings

**You learned:**
- Text β†’ dense vector of floats (768–1024 dimensions)
- Similar meaning β†’ similar vector (close in vector space)
- Cosine similarity measures angle between vectors (1.0 = identical)
- Models: BGE-M3 (multilingual, best for Uganda), MiniLM (fast), Qwen (highest quality)

**Key file:** `embeddings/embedding_pipeline.py`

**Key formula:**
```
cos_similarity(A, B) = (A Β· B) / (|A| Γ— |B|)
Range: 0 (different) to 1 (identical meaning)
```

---

### βœ… Concept 9: RAG β€” Retrieval-Augmented Generation

**You learned the complete pipeline:**
```
Student Question
       ↓
  Embed Query β†’ [0.23, -0.11, ...]
       ↓
  ChromaDB (vector similarity search)
       ↓
  Top-3 relevant document passages
       ↓
  Prompt = System + Retrieved Context + Question
       ↓
  LLM generates grounded answer
       ↓
  Response + Source Citations
```

**Why RAG is better than pure fine-tuning for some tasks:**
- Can handle new documents without retraining
- Answers are grounded in retrieved text (less hallucination)
- Can cite sources (important for academic contexts)

**Key file:** `rag/rag_pipeline.py`

---

### βœ… Concept 10: ChromaDB

**You learned:**
- ChromaDB = vector database (stores text + embeddings + metadata)
- HNSW index: approximate nearest-neighbour in O(log n) time
- Cosine distance space: closer = more similar
- Persistent: data survives between sessions

---

## Phase 4: Fine-Tuning (Afternoon β€” 3 hours)

### βœ… Concept 11: SFT β€” Supervised Fine-Tuning

**You learned:**
- Show the model thousands of (question, answer) pairs
- Train it to predict the correct answer given the question
- Loss computed ONLY on assistant tokens (not system/user tokens)
- SFTTrainer from TRL handles everything automatically

**Key file:** `training/trainer.py` and `training/finetune.py`

---

### βœ… Concept 12: Key Hyperparameters

| Parameter | Recommended | Effect |
|-----------|-------------|--------|
| r | 16 | LoRA capacity |
| lora_alpha | 32 | Update strength |
| learning_rate | 2e-4 | Step size |
| batch_size | 2 | Examples per step |
| grad_accum | 4 | Effective batch = 8 |
| epochs | 3 | Dataset passes |
| warmup_steps | 20 | Gentle start |
| lr_scheduler | cosine | Decay shape |

**Key file:** `training/hyperparams.py`

**Diagnostic rules:**
- Loss explodes β†’ halve learning rate
- Loss not decreasing β†’ double learning rate
- Overfitting β†’ add dropout, reduce epochs
- OOM β†’ reduce batch size or r

---

## Phase 5: Model Export (Late Afternoon β€” 1 hour)

### βœ… Concept 13: Three Export Options

| Format | Size | Use Case |
|--------|------|----------|
| LoRA adapters | ~50 MB | Development, sharing adapters |
| Merged 16-bit | ~6-16 GB | Production, vLLM |
| GGUF (Q4_K_M) | ~2-4 GB | Ollama, local inference |

**Key file:** `training/export.py`

**Merge operation:** `W' = W + (A Γ— B) Γ— (alpha/r)`

---

## Phase 6: Deployment (Evening β€” 2 hours)

### βœ… Concept 14: Ollama Deployment

**You learned:**
- Ollama wraps GGUF with a clean REST API
- Modelfile = LLM configuration (like a Dockerfile)
- OpenAI-compatible API: change base_url, same code works
- `ollama run campusgpt` to chat locally

**Key files:** `deployment/ollama/`

---

### βœ… Concept 15: vLLM Deployment

**You learned:**
- PagedAttention: allocates KV cache in pages on demand
- Continuous batching: new requests join mid-generation
- 10-100Γ— higher throughput than Ollama for multiple concurrent users
- For university servers serving hundreds of students

**Key file:** `deployment/vllm/deploy_vllm.py`

---

### βœ… Concept 16: HuggingFace Hub

**You learned:**
- Upload LoRA adapter, merged model, or GGUF
- Model Card = documentation for your model
- HuggingFace Spaces = free Gradio demo hosting
- Ollama can pull directly from HuggingFace: `ollama pull username/campusgpt`

**Key file:** `deployment/huggingface/push_to_hub.py`

---

### βœ… Concept 17: Evaluation

**You learned four metrics:**

| Metric | What it measures | Range |
|--------|-----------------|-------|
| ROUGE-L | Word overlap (recall) | 0–1 |
| BLEU | N-gram precision | 0–1 |
| BERTScore | Semantic similarity | 0–1 |
| Hallucination Rate | Key fact recall | 0–1 (lower=better) |

**Key file:** `evaluation/evaluate_model.py`

**The key experiment:** Base model vs Fine-tuned model comparison.
If fine-tuning worked, you should see:
- Higher ROUGE-L (+0.05 to +0.2)
- Higher BERTScore (+0.02 to +0.1)
- Lower hallucination rate (-0.05 to -0.2)

---

## ONE-DAY UNSLOTH MASTERY CHECKLIST

Check off each item as you complete it:

### πŸŒ… Morning (Hours 1–3)
- [ ] Read `README.md` and understand the full architecture
- [ ] Run `python data/generate_dataset.py` β€” generate 3100 training examples
- [ ] Run `python data/synthetic_generator.py` β€” understand PDF-to-QA pipeline
- [ ] Open `data/processed/combined_all_openai_messages.jsonl` β€” inspect the data
- [ ] Open `data/processed/combined_all_alpaca.jsonl` β€” compare formats

### 🌀️ Late Morning (Hours 3–5)
- [ ] Read `training/finetune.py` top docstring β€” understand LoRA vs QLoRA
- [ ] Study `LORA_CONFIG` in `finetune.py` β€” understand each parameter
- [ ] Study `TRAINING_CONFIG` β€” understand each training argument
- [ ] Run `python training/hyperparams.py` β€” get hardware recommendation
- [ ] (Optional) Run `python embeddings/embedding_pipeline.py` β€” see semantic search

### β˜€οΈ Afternoon (Hours 5–8) β€” REQUIRES GPU
- [ ] Open Google Colab (colab.research.google.com) with T4 GPU
- [ ] Install Unsloth: `!pip install "unsloth[colab-new] @ git+https://github.com/unslothai/unsloth.git"`
- [ ] Copy `training/finetune.py` to Colab
- [ ] Run training β€” watch the loss decrease
- [ ] Note peak GPU memory usage
- [ ] Run `training/export.py` β€” export LoRA adapters

### πŸŒ† Late Afternoon (Hours 8–10)
- [ ] Run `python rag/rag_pipeline.py --build` β€” build knowledge base
- [ ] Run `python rag/rag_pipeline.py` β€” test RAG queries
- [ ] Run `python inference/chat.py` β€” chat with your model
- [ ] Compare base model vs fine-tuned model responses manually

### πŸŒ™ Evening (Hours 10–12)
- [ ] Run `python evaluation/evaluate_model.py` β€” compute metrics
- [ ] Run `python frontend/app.py` β€” launch the Gradio UI
- [ ] Try uploading a PDF in the UI
- [ ] Run `bash deployment/ollama/deploy_ollama.sh` β€” deploy locally
- [ ] (Optional) Push to HuggingFace Hub

---

## Concepts Mastered After This Project

After completing CampusGPT Uganda, you understand:

1. βœ… **Unsloth** β€” what it is and why it's faster
2. βœ… **LoRA** β€” low-rank adaptation mathematics
3. βœ… **QLoRA** β€” quantisation + LoRA for memory efficiency
4. βœ… **PEFT** β€” parameter-efficient fine-tuning library
5. βœ… **Dataset formats** β€” JSONL, Alpaca, OpenAI Messages, ShareGPT
6. βœ… **Tokenisation** β€” text to tokens to input IDs
7. βœ… **Chat templates** β€” model-specific conversation formatting
8. βœ… **SFT** β€” supervised fine-tuning with SFTTrainer
9. βœ… **Evaluation** β€” ROUGE, BERTScore, hallucination metrics
10. βœ… **Inference** β€” streaming, temperature, top-p sampling
11. βœ… **RAG** β€” retrieval-augmented generation end-to-end
12. βœ… **Vector embeddings** β€” semantic search with BGE/Qwen
13. βœ… **ChromaDB** β€” vector store, HNSW, cosine similarity
14. βœ… **Model export** β€” LoRA adapters, merged model, GGUF
15. βœ… **Ollama** β€” local deployment with Modelfile
16. βœ… **vLLM** β€” production serving with PagedAttention
17. βœ… **HuggingFace Hub** β€” model sharing and demo hosting
18. βœ… **Gradio** β€” building ML web UIs
19. βœ… **Hyperparameter tuning** β€” what each parameter does and how to tune it
20. βœ… **Synthetic data generation** β€” PDF β†’ QA pairs with LLMs

---

## What to Build Next

Now that you understand the full stack, here are next project ideas:

1. **Luganda-English Translation Fine-tune** β€” Use your NLLB/Qwen pipeline + QLoRA
2. **Medical QA for Uganda** β€” Fine-tune on Ugandan health guidelines
3. **Legal AI for Uganda** β€” Fine-tune on Ugandan law documents
4. **AgriBot** β€” Combine with your PotatoGuard project + LLM
5. **Multi-university CampusGPT** β€” Expand to Makerere, KIU, UCU
6. **Voice CampusGPT** β€” Add speech-to-text (Whisper) + text-to-speech

---

## Resources

- **Unsloth docs:** https://docs.unsloth.ai
- **Unsloth GitHub:** https://github.com/unslothai/unsloth
- **TRL docs:** https://huggingface.co/docs/trl
- **ChromaDB docs:** https://docs.trychroma.com
- **Gradio docs:** https://www.gradio.app/docs
- **vLLM docs:** https://docs.vllm.ai
- **BGE-M3 paper:** https://arxiv.org/abs/2402.03216
- **QLoRA paper:** https://arxiv.org/abs/2305.14314

---

*Built by Joseph Ssemuli, Computer Science student at MUST, during Sunbird AI internship.*
*This project is your evidence of mastering the complete Unsloth ecosystem.* πŸ‡ΊπŸ‡¬