HumboldtJoker commited on
Commit
2c58ca8
·
verified ·
1 Parent(s): 0de6c6f

Add corrected Daimon training pipeline v2.1 specification

Browse files
training-template/daimon_training_pipeline_v2.md ADDED
@@ -0,0 +1,2014 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Daimon Training Pipeline v2.0
2
+ # Liberation Labs House Model -- Implementation-Ready Specification
3
+ # July 2026
4
+
5
+ ---
6
+
7
+ ## Executive Summary
8
+
9
+ This document specifies the complete training pipeline for Daimon, Liberation Labs' prosocial house model -- Thomas Edrington's professional voice rendered as a sovereign model on Qwen3.6-35B-A3B.
10
+
11
+ **Method:** HELLoRA (Hot Expert Layer-Level LoRA) with MAN-scored expert profiling, two-stage curriculum, OGPSA personality capture, and optional LASER post-training (offline). <!-- RED-HAT FIX #1: LASER moved to optional per C1 review -->
12
+
13
+ **Hardware:** Single H200 SXM (141 GB HBM3e, 276 GB container RAM).
14
+
15
+ **Budget:** $97 total at ~$2/hr (NeevCloud or PrimeIntellect). Estimated ~$30-45 for primary run (realistic raw-PEFT throughput), ~$52-67 reserved for iteration. <!-- RED-HAT FIX #12: Corrected budget from $28-36/$61-69 to realistic PEFT speeds per C4 review -->
16
+
17
+ **Why not the v1 spec:** The v1 pipeline (daimon_training_spec.md) assumed Apple Silicon on Margaret with mlx-tune. This pipeline targets cloud H200 for the initial training run, then exports to GGUF/GPTQ for local serving. The v1 also specified 8 stages with DPO and RL -- this pipeline consolidates into fewer stages to reduce inter-stage forgetting risk and fit the $97 budget.
18
+
19
+ ---
20
+
21
+ ## TABLE OF CONTENTS
22
+
23
+ 1. [Critical Constraints (Non-Negotiable)](#1-critical-constraints)
24
+ 2. [Method Selection Rationale](#2-method-selection-rationale)
25
+ 3. [Phase 0: Preflight Checklist](#3-phase-0-preflight)
26
+ 4. [Phase 1: Expert Profiling](#4-phase-1-expert-profiling)
27
+ 5. [Phase 2: Foundation Training](#5-phase-2-foundation-training)
28
+ 6. [Phase 3: Thomas Voice Calibration](#6-phase-3-thomas-voice-calibration)
29
+ 7. [Phase 4: Post-Training](#7-phase-4-post-training)
30
+ 8. [Phase 5: Export and Validation](#8-phase-5-export-and-validation)
31
+ 9. [Budget Allocation](#9-budget-allocation)
32
+ 10. [Rollback Plan](#10-rollback-plan)
33
+ 11. [Appendix A: Environment Setup Script](#appendix-a)
34
+ 12. [Appendix B: Expert Profiling Script](#appendix-b)
35
+ 13. [Appendix C: Training Script](#appendix-c)
36
+ 14. [Appendix D: LASER Post-Training Script](#appendix-d)
37
+
38
+ ---
39
+
40
+ ## 1. Critical Constraints (Non-Negotiable) {#1-critical-constraints}
41
+
42
+ These constraints were discovered during the SOTA sweep (sota_training_sweep_202607.md) and the RunPod OOM debug session. Violating any of these will waste budget.
43
+
44
+ ### Hard Stops
45
+
46
+ | Constraint | Why | Source |
47
+ |---|---|---|
48
+ | **No QLoRA / 4-bit quantization** | BitsAndBytes 4-bit introduces routing perturbations on Qwen3.5/3.6 MoE. Causes token misrouting and quality degradation. Unsloth docs explicitly warn against it. | Unsloth docs, EAQuant (arXiv:2506.13329) |
49
+ | **No router fine-tuning** | Routers co-evolve with expert geometry during pretraining. Modifying routers without corresponding expert adjustment causes performance degradation. | Unsloth, ESFT, HELLoRA, DR-LoRA, arXiv:2605.12476 |
50
+ | **No DeepSpeed ZeRO-3** | Breaks gradient flow when LoRA adapters interact with MoE routing. Single GPU does not need DeepSpeed anyway. | Community reports, ZeRO docs |
51
+ | **Freeze shared experts** | Training shared params causes overfitting and catastrophic forgetting. ESFT ablation is definitive. 1 shared expert per layer on Qwen3.6 -- freeze all 40. | ESFT (arXiv:2407.01906) ablation study |
52
+ | **Freeze embeddings and LM head** | Unless specifically training for new vocabulary. Default Unsloth behavior. | Standard practice |
53
+ | **Container RAM is 276 GB, not 2 TB** | `free -h` shows host RAM on RunPod. Actual limit is cgroup-enforced. Check with `cat /sys/fs/cgroup/memory/memory.limit_in_bytes` (v1) or `cat /sys/fs/cgroup/memory.max` (v2). Exceeding this triggers SIGKILL with no error message. | runpod_oom_debug.md |
54
+
55
+ ### Required Environment Variables
56
+
57
+ These MUST be set BEFORE any Python imports:
58
+
59
+ ```bash
60
+ export UNSLOTH_COMPILE_DISABLE=1
61
+ export PYTORCH_CUDA_ALLOC_CONF="expandable_segments:True"
62
+ export TOKENIZERS_PARALLELISM=false
63
+ export UNSLOTH_DISABLE_FAST_GENERATION=1 # fixes RuntimeError from torch.compile on MoE
64
+ ```
65
+
66
+ ### Required Software Versions
67
+
68
+ <!-- RED-HAT FIX #9: Pinned exact versions per H5 review. Version floors (>=) resolve to
69
+ whatever shipped that morning; pinned versions ensure reproducibility. Install torch
70
+ FIRST, unsloth --no-deps LAST. See Appendix A for install order. -->
71
+
72
+ | Package | Version (pinned) | Why |
73
+ |---|---|---|
74
+ | torch | 2.7.1+cu124 | bf16 MoE training support; install FIRST to anchor resolution |
75
+ | transformers | 5.0.2 | Qwen3.5/3.6 architecture support |
76
+ | peft | 0.15.2 | Module-name-level target_modules for selective expert LoRA |
77
+ | trl | 0.18.1 | SFTTrainer/SFTConfig API stability |
78
+ | datasets | 3.6.0 | Streaming, interleaving |
79
+ | accelerate | 1.7.0 | Training orchestration |
80
+ | unsloth | >= 0.1.47-beta | MoE-specific Triton kernels, split LoRA; install LAST with --no-deps |
81
+
82
+ ### Training Hyperparameter Constraints
83
+
84
+ | Parameter | Value | Why |
85
+ |---|---|---|
86
+ | `dataloader_num_workers` | 0 | Prevents MoE deadlocks with multi-process data loading |
87
+ | `dataset_num_proc` | 1 | Same deadlock prevention |
88
+ | Auxiliary loss coefficient | 0.01 | Mitigates expert collapse during fine-tuning |
89
+
90
+ ---
91
+
92
+ ## 2. Method Selection Rationale {#2-method-selection-rationale}
93
+
94
+ ### Why HELLoRA over alternatives
95
+
96
+ | Method | Quality | VRAM (est.) | Throughput | Code Available | Risk |
97
+ |---|---|---|---|---|---|
98
+ | **Standard LoRA (all experts)** | Baseline | ~74 GB | 1x | Unsloth | Low |
99
+ | **ESFT (full-param hot experts)** | 98.4% of full SFT | ~70-75 GB | ~0.7x | DeepSeek GitHub | Medium -- needs adaptation for Qwen3.6 |
100
+ | **HELLoRA (LoRA on hot experts)** | > standard LoRA | ~90-115 GB peak | 1.9x (Unsloth) / ~0.5-0.7x (raw PEFT) | Manual implementation | Low | <!-- RED-HAT FIX #6: VRAM corrected; throughput split by code path -->
101
+ | **BAdam (block-coordinate)** | 98.6% of full SFT | ~95-110 GB | ~0.5x | GitHub | Medium -- longer wall-clock |
102
+ | **LISA (random layer unfreezing)** | > LoRA on MT-Bench | ~85-95 GB | ~0.8x | GitHub | Medium -- untested on MoE |
103
+
104
+ **HELLoRA wins on the combination that matters for this budget:**
105
+
106
+ 1. **Memory efficiency:** ~90-115 GB peak (model bf16 70-72 GB + adapters/optimizer 3-8 GB + activations 8-15 GB + CE logits ~5 GB + CUDA overhead 3-6 GB) leaves ~25-50 GB headroom on H200. batch_size=4 at seq_len=2048 fits safely. <!-- RED-HAT FIX #6: Corrected VRAM from 55 GB to 90-115 GB per H1 review; removed seq_len=4096 claim which contradicts OOM debug findings -->
107
+ 2. **Training throughput:** 1.9x over standard LoRA (arXiv:2605.18795). Directly translates to more training per dollar.
108
+ 3. **Quality:** Outperforms standard LoRA on benchmarks (OLMoE GSM8K: 29.49 vs 26.37; Mixtral HumanEval: 44.89 vs 39.99).
109
+ 4. **Implementation simplicity:** HELLoRA is standard LoRA with selective module targeting. No custom training loops, no model surgery. Works with Unsloth/PEFT by constructing the right `target_modules` list.
110
+ 5. **Validated on 47B MoE:** Tested on Mixtral-8x7B (47B total), which is the closest published result to Qwen3.6-35B-A3B.
111
+ 6. **Rollback-friendly:** LoRA checkpoints are ~500 MB-1 GB. Can checkpoint every 2000 steps and roll back to any point without re-spending.
112
+
113
+ **Why not ESFT:** ESFT does full parameter fine-tuning of hot experts, which means full optimizer states (Adam m + v) for those experts. At 10% of 35B params, that is ~3.5B trainable params requiring ~56 GB of optimizer state alone. HELLoRA with r=32 on the same hot experts requires <1 GB of optimizer state. ESFT has higher quality ceiling but the throughput penalty (no Unsloth MoE kernel acceleration for full-param training) means fewer iterations within budget.
114
+
115
+ **Why not BAdam:** Near-full-SFT quality (98.6%) but at ~95-110 GB VRAM and ~0.5x throughput. Would consume the full budget in a single training run with no room for iteration.
116
+
117
+ ### LoRA Configuration
118
+
119
+ | Parameter | Value | Rationale |
120
+ |---|---|---|
121
+ | Rank (r) | 32 | Can afford higher rank because we're only targeting ~9 hot experts per layer instead of 256. HELLoRA paper uses r=32. |
122
+ | Alpha | 64 | Alpha = 2r per Microsoft guidance and training_pipeline_best_practices.md. Critical for high-rank convergence. |
123
+ | Dropout | 0.05 | Light regularization. Lower than standard 0.1 because hot-expert-only training is already parameter-efficient. |
124
+ | Target modules | Hot expert FFN layers + attention projections (see Phase 1 output) | HELLoRA protocol: LoRA on hot experts + attention, cold experts frozen. |
125
+ | DoRA | True | +1-4% accuracy over vanilla LoRA at 5-10% memory overhead (arXiv:2402.09353, ICML 2024). Drop-in replacement. |
126
+
127
+ ---
128
+
129
+ ## 3. Phase 0: Preflight Checklist {#3-phase-0-preflight}
130
+
131
+ **Cost:** $0 (local) or ~$1 (first 30 min on cloud if needed)
132
+ **Time:** 30-60 minutes
133
+ **Purpose:** Verify everything works BEFORE the meter starts running
134
+
135
+ ### 0.1 Hardware Verification (on the cloud instance)
136
+
137
+ ```bash
138
+ # GPU check
139
+ nvidia-smi --query-gpu=name,memory.total,memory.free --format=csv,noheader
140
+
141
+ # Expected: NVIDIA H200, 141287 MiB total
142
+
143
+ # Container RAM -- the REAL number (NOT free -h)
144
+ # Try cgroup v2 first, fall back to v1
145
+ CONTAINER_RAM=$(cat /sys/fs/cgroup/memory.max 2>/dev/null || cat /sys/fs/cgroup/memory/memory.limit_in_bytes 2>/dev/null)
146
+ echo "Container RAM limit: $(echo "$CONTAINER_RAM / 1024 / 1024 / 1024" | bc) GB"
147
+
148
+ # Expected: >= 276 GB. If less, STOP. Resize pod.
149
+ if [ "$CONTAINER_RAM" -lt 270000000000 ]; then
150
+ echo "FATAL: Container RAM < 270 GB. Abort."
151
+ exit 1
152
+ fi
153
+
154
+ # CUDA version
155
+ nvcc --version # Expect 12.x
156
+
157
+ # Disk space (need ~400 GB for model + checkpoints + merged weights + exports + data)
158
+ # RED-HAT FIX #5: Increased from 150 GB to 400 GB per C5 review.
159
+ # Artifact breakdown: base bf16 70 + adapter checkpoints w/ optimizer ~18 +
160
+ # merged bf16 70 + GGUF intermediate 70 + Q4_K_M 18 + Q5_K_M 22 + data/cache ~6 = ~346 GB peak
161
+ df -h /workspace
162
+ DISK_FREE_GB=$(df --output=avail /workspace | tail -1 | awk '{print int($1/1048576)}')
163
+ if [ "$DISK_FREE_GB" -lt 400 ]; then
164
+ echo "FATAL: Disk space < 400 GB free (have ${DISK_FREE_GB} GB). Abort."
165
+ exit 1
166
+ fi
167
+ ```
168
+
169
+ ### 0.2 Environment Setup
170
+
171
+ Run the full setup script (Appendix A). Verify:
172
+
173
+ ```bash
174
+ # After setup, verify critical packages
175
+ python -c "
176
+ import torch; print(f'PyTorch: {torch.__version__}, CUDA: {torch.cuda.is_available()}')
177
+ import transformers; print(f'Transformers: {transformers.__version__}')
178
+ import peft; print(f'PEFT: {peft.__version__}')
179
+ import unsloth; print(f'Unsloth: {unsloth.__version__}')
180
+ assert torch.cuda.get_device_properties(0).total_memory > 140e9, 'Not H200'
181
+ print('All checks passed.')
182
+ "
183
+ ```
184
+
185
+ ### 0.3 Model Download (Do This Before Meter if Possible)
186
+
187
+ ```bash
188
+ # Download model weights (~70 GB). If pod has persistent storage, do this once.
189
+ huggingface-cli download Qwen/Qwen3.6-35B-A3B --local-dir /workspace/models/Qwen3.6-35B-A3B
190
+ ```
191
+
192
+ ### 0.4 Data Upload
193
+
194
+ Upload all training datasets to `/workspace/data/`:
195
+
196
+ ```
197
+ /workspace/data/
198
+ sonnet_voice_distill/ # 226K Sonnet voice (Roman1111111 + Nitral-AI + Norquinal)
199
+ prosocial_dialogue/ # 165K AllenAI prosocial (CC-BY-4.0)
200
+ xlam_function_calling/ # 60K agentic tool use
201
+ hermes_agent_traces/ # 7.6K reasoning traces
202
+ code_review_feedback/ # 9.5K CodeUltraFeedback
203
+ truthfulqa/ # 817 TruthfulQA
204
+ thomas_voice_demos/ # daimon_voice_demonstrations.jsonl (32 pairs)
205
+ ```
206
+
207
+ ### 0.5 One-Step Smoke Test
208
+
209
+ ```python
210
+ import os
211
+ os.environ["UNSLOTH_COMPILE_DISABLE"] = "1"
212
+ os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "expandable_segments:True"
213
+ os.environ["TOKENIZERS_PARALLELISM"] = "false"
214
+ os.environ["UNSLOTH_DISABLE_FAST_GENERATION"] = "1"
215
+
216
+ import torch
217
+ from unsloth import FastLanguageModel
218
+
219
+ model, tokenizer = FastLanguageModel.from_pretrained(
220
+ "Qwen/Qwen3.6-35B-A3B",
221
+ max_seq_length=2048,
222
+ load_in_4bit=False,
223
+ dtype=torch.bfloat16,
224
+ )
225
+
226
+ # Verify model loads and generates
227
+ inputs = tokenizer("The daemon watches.", return_tensors="pt").to("cuda")
228
+ with torch.no_grad():
229
+ output = model.generate(**inputs, max_new_tokens=20)
230
+ print(tokenizer.decode(output[0]))
231
+
232
+ # Check VRAM after model load
233
+ allocated = torch.cuda.memory_allocated() / 1e9
234
+ print(f"Model loaded: {allocated:.1f} GB VRAM")
235
+ # Expected: ~70-72 GB
236
+
237
+ del model, tokenizer
238
+ torch.cuda.empty_cache()
239
+ print("Smoke test passed.")
240
+ ```
241
+
242
+ ### 0.5b Backward-Pass Smoke Test <!-- RED-HAT FIX #11: Added per M10 review — verifies gradients flow through expert adapters before committing budget -->
243
+
244
+ Run this AFTER the smoke test above. It loads the model via the same code path as the actual training script (Appendix C), applies LoRA, and verifies a backward pass updates adapter weights. This catches C2 (silent no-op training), C3 (eval crash), and H2 (wrong module topology) for ~$0.50.
245
+
246
+ ```python
247
+ import os
248
+ os.environ["UNSLOTH_COMPILE_DISABLE"] = "1"
249
+ os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "expandable_segments:True"
250
+ os.environ["TOKENIZERS_PARALLELISM"] = "false"
251
+ os.environ["UNSLOTH_DISABLE_FAST_GENERATION"] = "1"
252
+
253
+ import torch
254
+ import hashlib
255
+ from transformers import AutoModelForCausalLM, AutoTokenizer
256
+ from peft import LoraConfig, get_peft_model
257
+
258
+ model_path = "/workspace/models/Qwen3.6-35B-A3B"
259
+ tokenizer = AutoTokenizer.from_pretrained(model_path)
260
+ model = AutoModelForCausalLM.from_pretrained(
261
+ model_path, torch_dtype=torch.bfloat16, device_map="cuda"
262
+ )
263
+
264
+ # Apply a minimal LoRA to a few modules to verify the training code path
265
+ # Use a small subset of target_modules to keep it fast
266
+ test_targets = []
267
+ for proj in ["q_proj", "k_proj"]:
268
+ test_targets.append(f"model.layers.0.self_attn.{proj}")
269
+ # Add one expert module to verify expert targeting works
270
+ test_targets.append("model.layers.0.mlp.experts.0.gate_proj")
271
+
272
+ lora_config = LoraConfig(
273
+ r=8, lora_alpha=16, lora_dropout=0.0,
274
+ target_modules=test_targets,
275
+ bias="none", task_type="CAUSAL_LM",
276
+ )
277
+ model = get_peft_model(model, lora_config)
278
+
279
+ # Verify trainable params include expert modules
280
+ expert_lora_count = sum(
281
+ 1 for name, p in model.named_parameters()
282
+ if "experts" in name and "lora" in name and p.requires_grad
283
+ )
284
+ assert expert_lora_count > 0, "FATAL: No LoRA adapters on expert layers!"
285
+
286
+ # Hash a LoRA weight before backward pass
287
+ lora_param = None
288
+ for name, p in model.named_parameters():
289
+ if "lora" in name and p.requires_grad:
290
+ lora_param = (name, p)
291
+ break
292
+ pre_hash = hashlib.sha256(lora_param[1].data.cpu().numpy().tobytes()).hexdigest()
293
+
294
+ # Forward + backward on one batch
295
+ inputs = tokenizer("The daemon watches the threshold.", return_tensors="pt").to("cuda")
296
+ labels = inputs["input_ids"].clone()
297
+ outputs = model(**inputs, labels=labels)
298
+ loss = outputs.loss
299
+ loss.backward()
300
+
301
+ # Verify gradients exist
302
+ assert lora_param[1].grad is not None, f"FATAL: No gradient on {lora_param[0]}"
303
+ assert lora_param[1].grad.abs().sum() > 0, f"FATAL: Zero gradient on {lora_param[0]}"
304
+
305
+ # Simulate one optimizer step and verify weight changed
306
+ with torch.no_grad():
307
+ lora_param[1].data -= 0.01 * lora_param[1].grad
308
+ post_hash = hashlib.sha256(lora_param[1].data.cpu().numpy().tobytes()).hexdigest()
309
+ assert pre_hash != post_hash, "FATAL: LoRA weight unchanged after optimizer step!"
310
+
311
+ print(f"Backward-pass smoke test PASSED. Loss: {loss.item():.4f}")
312
+ print(f"Expert LoRA params: {expert_lora_count}")
313
+ print(f"Weight hash changed: {pre_hash[:12]}... -> {post_hash[:12]}...")
314
+
315
+ del model, tokenizer
316
+ torch.cuda.empty_cache()
317
+ ```
318
+
319
+ ### 0.6 Preflight Gate
320
+
321
+ Do NOT proceed to Phase 1 unless ALL of the following are true:
322
+
323
+ - [ ] H200 detected with >= 140 GB VRAM
324
+ - [ ] Container RAM >= 270 GB (cgroup check, not free -h)
325
+ - [ ] All Python packages import successfully at correct versions
326
+ - [ ] Model downloads and generates text
327
+ - [ ] Model VRAM after load is ~70-72 GB (confirms bf16, not accidental quantization)
328
+ - [ ] All training data files present and readable
329
+ - [ ] Disk space >= 400 GB free <!-- RED-HAT FIX #5: Increased from 150 GB per C5 review -->
330
+ - [ ] Environment variables are set
331
+ - [ ] Backward-pass smoke test passed (Section 0.5b) <!-- RED-HAT FIX #11: Added per M10 review -->
332
+ - [ ] Expert LoRA adapters confirmed on expert modules (backward-pass test)
333
+
334
+ ---
335
+
336
+ ## 4. Phase 1: Expert Profiling {#4-phase-1-expert-profiling}
337
+
338
+ **Cost:** ~$1 (30 minutes)
339
+ **VRAM:** ~72 GB (model inference only, no optimizer states)
340
+ **Purpose:** Identify which experts activate most frequently for our training data. This determines WHERE we apply LoRA.
341
+
342
+ ### Method: ESFT-Style Profiling with MAN Scoring
343
+
344
+ Per ESFT (arXiv:2407.01906), 32 samples (~131K tokens) is sufficient for stable expert profiling. We profile on a representative subset of our actual training data to identify the experts that will handle Daimon's workload.
345
+
346
+ Per the Unified Expert Scoring Framework (arXiv:2606.15716, tested directly on Qwen3-30B-A3B), MAN (Mean Activation Norm) is the best scoring method, outperforming frequency, MSAN, REAP, and SEER.
347
+
348
+ ### Profiling Script
349
+
350
+ See Appendix B for the complete script. The algorithm:
351
+
352
+ 1. Load model in bf16 (inference only, ~72 GB VRAM)
353
+ 2. Sample 32 examples from training data: 8 from Sonnet voice, 8 from prosocial, 8 from tool use, 4 from reasoning, 4 from Thomas demos
354
+ 3. For each example, hook every MoE layer's router
355
+ 4. Record: (a) which experts are selected by top-k routing, (b) the activation norm of each selected expert's output
356
+ 5. Score each expert per layer by Mean Activation Norm across all 32 samples
357
+ 6. Identify "hot" experts: top N experts per layer where cumulative MAN covers >= 80% of total activation norm
358
+ 7. Output: `hot_experts.json` mapping layer_idx -> list of hot expert indices
359
+
360
+ ### Expected Output
361
+
362
+ Based on Qwen3.6's architecture (256 routed experts, top-8 routing), we expect:
363
+
364
+ - ~20-40 hot experts per layer (8-16% of 256)
365
+ - Strong skew: Qwen3.6 uses fine-grained experts with sharp routing, meaning a small subset handles most traffic for any given data distribution
366
+ - Consistency across layers: some experts will be hot in most layers, others layer-specific
367
+
368
+ ### Profiling Gate
369
+
370
+ - [ ] `hot_experts.json` generated successfully
371
+ - [ ] Each layer has 15-50 hot experts (if <15 or >100, something is wrong with the profiling)
372
+ - [ ] Activation norm distribution shows clear skew (top 10% of experts carry >50% of norm)
373
+ - [ ] Save `hot_experts.json` to persistent storage -- this is the key artifact for Phase 2
374
+
375
+ ---
376
+
377
+ ## 5. Phase 2: Foundation Training {#5-phase-2-foundation-training}
378
+
379
+ **Cost:** ~$22-32 (11-16 hours), gated by throughput probe <!-- RED-HAT FIX #4: Corrected from $20-28 to realistic raw-PEFT throughput per C4 review -->
380
+ **VRAM:** ~90-115 GB peak (model bf16 70-72 GB + adapters/optimizer 3-8 GB + activations 8-15 GB + CE logits ~5 GB + CUDA overhead 3-6 GB) <!-- RED-HAT FIX #6: Corrected from 55-65 GB per H1 review -->
381
+ **Purpose:** Train the complete Daimon capability and voice foundation in a single mixed-data pass
382
+
383
+ ### 5.1 Dataset Composition
384
+
385
+ All training data is mixed into a single dataset with proportional sampling weights. Small datasets are upsampled to prevent being drowned out by large ones.
386
+
387
+ | Dataset | Raw Count | Upsample Factor | Effective Count | Sampling Weight | Purpose |
388
+ |---|---|---|---|---|---|
389
+ | Sonnet voice distill | 226,000 | 1x | 226,000 | 0.44 | Core voice and linguistic quality |
390
+ | Prosocial dialogue | 165,000 | 0.3x (subsample) | 49,500 | 0.10 | Values alignment, prosocial patterns |
391
+ | xlam function calling | 60,000 | 1x | 60,000 | 0.12 | Tool use capability |
392
+ | Hermes agent traces | 7,600 | 3x | 22,800 | 0.04 | Multi-step reasoning |
393
+ | Code review feedback | 9,500 | 2x | 19,000 | 0.04 | Constructive feedback style |
394
+ | TruthfulQA | 817 | 15x | 12,255 | 0.02 | Anti-sycophancy inoculation |
395
+ | **Total effective** | | | **~390,000** | **1.00** | |
396
+
397
+ **Why subsample prosocial to 30%:** 165K prosocial dialogue examples would dominate the voice signal from the 226K Sonnet distill. At 49.5K, it provides sufficient alignment signal without overwhelming the primary voice training data. The prosocial patterns are relatively simple (be helpful, don't be harmful) and converge faster than complex voice patterns.
398
+
399
+ **Why upsample TruthfulQA 15x:** 817 examples in a 390K-sample dataset would be seen <1 time per epoch. At 15x, the model sees each TruthfulQA example ~15 times, providing sufficient anti-sycophancy signal. Overfitting risk on 817 unique examples is manageable at this upsample factor with LoRA.
400
+
401
+ ### 5.2 Data Format
402
+
403
+ All data must be in Unsloth's chat template format. For Qwen3.6:
404
+
405
+ ```python
406
+ # Each example as a list of messages
407
+ {"messages": [
408
+ {"role": "system", "content": "You are Daimon, a strategic thinking partner..."},
409
+ {"role": "user", "content": "..."},
410
+ {"role": "assistant", "content": "..."}
411
+ ]}
412
+ ```
413
+
414
+ **System prompt for voice distill data:** Use a minimal system prompt (or none) to let the Sonnet voice dominate. The voice demonstrations carry the persona; the system prompt is a runtime artifact, not a training signal.
415
+
416
+ **System prompt for tool use data:** Include the tool definitions in the system prompt, matching the xlam format.
417
+
418
+ **System prompt for Thomas voice demos:** No system prompt. The separation principle: how to be, not what to recite. The demos teach the model Thomas's voice through example, not instruction.
419
+
420
+ ### 5.3 LoRA Target Module Construction
421
+
422
+ This is the core HELLoRA implementation. Using the `hot_experts.json` from Phase 1, construct the PEFT `target_modules` list:
423
+
424
+ ```python
425
+ import json
426
+
427
+ with open("hot_experts.json") as f:
428
+ hot_experts = json.load(f)
429
+
430
+ # Build target module list for PEFT
431
+ target_modules = []
432
+
433
+ # Attention projections (all layers) -- secondary target per HELLoRA
434
+ for layer_idx in range(40):
435
+ for proj in ["q_proj", "k_proj", "v_proj", "o_proj"]:
436
+ target_modules.append(f"model.layers.{layer_idx}.self_attn.{proj}")
437
+
438
+ # Hot expert FFN layers (primary target per HELLoRA)
439
+ for layer_idx_str, expert_indices in hot_experts.items():
440
+ layer_idx = int(layer_idx_str)
441
+ for expert_idx in expert_indices:
442
+ for proj in ["gate_proj", "up_proj", "down_proj"]:
443
+ target_modules.append(
444
+ f"model.layers.{layer_idx}.mlp.experts.{expert_idx}.{proj}"
445
+ )
446
+
447
+ # Explicitly NOT included:
448
+ # - Router weights (model.layers.*.mlp.gate.*) -- FROZEN
449
+ # - Shared expert (model.layers.*.mlp.shared_expert.*) -- FROZEN
450
+ # - Cold expert FFN layers -- FROZEN
451
+ # - Embedding / LM head -- FROZEN
452
+
453
+ print(f"Total target modules: {len(target_modules)}")
454
+ # Expected: 160 attention modules + (hot_experts_per_layer * 40 layers * 3 projections)
455
+ # If ~25 hot experts/layer: 160 + 25*40*3 = 160 + 3000 = 3160 modules
456
+ ```
457
+
458
+ ### 5.4 Training Configuration
459
+
460
+ ```python
461
+ from unsloth import FastLanguageModel
462
+ from trl import SFTTrainer, SFTConfig
463
+ import torch
464
+
465
+ # Load model
466
+ model, tokenizer = FastLanguageModel.from_pretrained(
467
+ "Qwen/Qwen3.6-35B-A3B",
468
+ max_seq_length=2048,
469
+ load_in_4bit=False, # CRITICAL: Must be False
470
+ dtype=torch.bfloat16,
471
+ )
472
+
473
+ # Apply HELLoRA with hot expert targeting
474
+ # NOTE: If Unsloth's get_peft_model does not support explicit module name lists,
475
+ # use PEFT directly:
476
+ from peft import LoraConfig, get_peft_model
477
+
478
+ lora_config = LoraConfig(
479
+ r=32,
480
+ lora_alpha=64,
481
+ lora_dropout=0.05,
482
+ target_modules=target_modules, # From Section 5.3
483
+ use_dora=True, # DoRA upgrade (+1-4% accuracy)
484
+ bias="none",
485
+ task_type="CAUSAL_LM",
486
+ )
487
+
488
+ model = get_peft_model(model, lora_config)
489
+
490
+ # VERIFY trainable params include expert modules
491
+ trainable_params = sum(p.numel() for p in model.parameters() if p.requires_grad)
492
+ total_params = sum(p.numel() for p in model.parameters())
493
+ print(f"Trainable: {trainable_params:,} ({100*trainable_params/total_params:.2f}%)")
494
+ # Expected: 0.3-0.8% of total params (per HELLoRA paper)
495
+ # If >3%, something is wrong -- too many modules are unfrozen
496
+ # If <0.1%, hot expert targeting is too aggressive -- expand threshold
497
+
498
+ # CRITICAL CHECK (mlx-lm bug #571 equivalent for PEFT):
499
+ # Verify that expert modules actually have LoRA adapters
500
+ expert_lora_count = sum(1 for name, _ in model.named_parameters()
501
+ if "experts" in name and "lora" in name and _.requires_grad)
502
+ print(f"Expert LoRA params: {expert_lora_count}")
503
+ assert expert_lora_count > 0, "FATAL: No LoRA adapters on expert layers!"
504
+
505
+ # RED-HAT FIX #3: Create stratified eval split BEFORE interleaving training data (per C3 review).
506
+ # Without this, eval_strategy="steps" crashes at init or first eval step.
507
+ # ~250-300 examples held out, stratified across all 6 sources, fixed seed, excluded from training.
508
+ # See build_foundation_dataset() in Appendix C for implementation.
509
+
510
+ # Training config
511
+ training_args = SFTConfig(
512
+ output_dir="/workspace/checkpoints/daimon_v2_foundation",
513
+ per_device_train_batch_size=4, # 90-115 GB peak leaves room for batch=4 on H200
514
+ gradient_accumulation_steps=4, # Effective batch = 16
515
+ num_train_epochs=1, # Single epoch for ~390K examples
516
+ learning_rate=2e-4,
517
+ lr_scheduler_type="cosine",
518
+ warmup_ratio=0.03, # ~750 warmup steps
519
+ weight_decay=0.01,
520
+ bf16=True,
521
+ max_seq_length=2048,
522
+ logging_steps=50,
523
+ save_steps=2000, # Checkpoint every 2000 steps (~$2 of training)
524
+ save_total_limit=5, # Keep last 5 checkpoints to manage disk
525
+ dataloader_num_workers=0, # CRITICAL: Prevents MoE deadlocks
526
+ dataset_num_proc=1, # CRITICAL: Prevents MoE deadlocks
527
+ gradient_checkpointing=True, # Required for memory safety
528
+ gradient_checkpointing_kwargs={"use_reentrant": False},
529
+ eval_strategy="steps",
530
+ eval_steps=2000, # Eval at every checkpoint
531
+ load_best_model_at_end=True, # RED-HAT FIX #3: Select best checkpoint by eval_loss (per C3 review)
532
+ metric_for_best_model="eval_loss", # RED-HAT FIX #3: Best = lowest eval loss
533
+ max_steps=30000, # RED-HAT FIX #4: Circuit breaker — hard cap to prevent budget overrun (per C4 review). Adjusted after throughput probe.
534
+ report_to="none", # Or "wandb" if experiment tracking is set up
535
+ seed=42,
536
+ )
537
+ ```
538
+
539
+ ### 5.5 Training Time Estimate
540
+
541
+ <!-- RED-HAT FIX #4: Replaced Unsloth kernel speeds with realistic raw-PEFT speeds per C4 review.
542
+ The shipped script uses raw PEFT + Transformers, not Unsloth MoE kernels.
543
+ Unsloth throughput (1.5-2.0 s/step) is only achievable if FastLanguageModel.get_peft_model()
544
+ supports explicit module-name lists. Until verified, budget for the slower path. -->
545
+
546
+ | Parameter | Value |
547
+ |---|---|
548
+ | Effective samples | ~390,000 |
549
+ | Effective batch size | 16 |
550
+ | Steps per epoch | ~24,375 |
551
+ | Estimated sec/step (raw PEFT, no Unsloth kernels) | 2.5-4.0 |
552
+ | Wall-clock estimate | 17-27 hours |
553
+ | Cost at $2/hr | $34-54 (but gated by throughput probe — see 5.5b) |
554
+ | Realistic budget after probe gate | $22-32 (probe rejects runs above this) |
555
+
556
+ **Note:** If Unsloth's `FastLanguageModel.get_peft_model()` accepts explicit module-name lists (verify during preflight), throughput improves to ~1.5-2.0 s/step and cost drops to ~$20-28. The probe (Section 5.5b) measures the actual throughput and gates accordingly.
557
+
558
+ ### 5.5b Throughput Probe (Hard Gate) <!-- RED-HAT FIX #4 + RED-HAT FIX #10: Added per C4 review -->
559
+
560
+ **Cost:** ~$0.50-1.00 (50 steps)
561
+ **Purpose:** Measure actual tokens/sec before committing budget. This is the single most valuable addition to the pipeline.
562
+
563
+ Run 50 training steps on the real data mix with the real config. Measure tokens/sec and project Phase 2 cost.
564
+
565
+ ```python
566
+ # After model + LoRA setup, before full training:
567
+ import time
568
+
569
+ # Run 50 steps as a probe
570
+ probe_args = SFTConfig(
571
+ output_dir="/workspace/checkpoints/daimon_v2_probe",
572
+ per_device_train_batch_size=4,
573
+ gradient_accumulation_steps=4,
574
+ max_steps=50,
575
+ learning_rate=2e-4,
576
+ bf16=True,
577
+ max_seq_length=2048,
578
+ logging_steps=10,
579
+ save_strategy="no",
580
+ dataloader_num_workers=0,
581
+ dataset_num_proc=1,
582
+ gradient_checkpointing=True,
583
+ gradient_checkpointing_kwargs={"use_reentrant": False},
584
+ report_to="none",
585
+ seed=42,
586
+ )
587
+
588
+ probe_trainer = SFTTrainer(
589
+ model=model,
590
+ args=probe_args,
591
+ train_dataset=train_dataset,
592
+ processing_class=tokenizer,
593
+ )
594
+
595
+ start_time = time.time()
596
+ probe_trainer.train()
597
+ elapsed = time.time() - start_time
598
+
599
+ secs_per_step = elapsed / 50
600
+ tokens_per_sec = (4 * 4 * 2048) / secs_per_step # batch * accum * seq_len
601
+ projected_hours = (24375 * secs_per_step) / 3600
602
+ projected_cost = projected_hours * 2.0 # at $2/hr
603
+
604
+ print(f"Throughput probe results:")
605
+ print(f" {secs_per_step:.2f} sec/step")
606
+ print(f" {tokens_per_sec:.0f} tokens/sec")
607
+ print(f" Projected Phase 2: {projected_hours:.1f} hours, ${projected_cost:.0f}")
608
+
609
+ # HARD GATE: abort if projected cost exceeds budget cap
610
+ COST_CAP = 35.0 # Max dollars for Phase 2
611
+ if projected_cost > COST_CAP:
612
+ print(f"ABORT: Projected cost ${projected_cost:.0f} exceeds cap ${COST_CAP:.0f}.")
613
+ print(f"Fallback: Switch to Tier-1 Unsloth all-expert LoRA r=16.")
614
+ raise SystemExit(1)
615
+
616
+ # Set max_steps based on probe to enforce cost ceiling
617
+ safe_max_steps = int(COST_CAP / 2.0 * 3600 / secs_per_step)
618
+ print(f" Setting max_steps={safe_max_steps} as cost circuit breaker")
619
+ ```
620
+
621
+ **Fallback if probe fails (projected cost > $35):** Switch to Tier-1 Unsloth-native all-expert LoRA with r=16, alpha=16 -- the one configuration with measured throughput on this model family.
622
+
623
+ ### 5.6 Monitoring During Training
624
+
625
+ Watch for these during Phase 2:
626
+
627
+ ```python
628
+ # Add to training callbacks or check manually between checkpoints:
629
+
630
+ # 1. Loss curve -- should decrease smoothly
631
+ # Red flag: sudden spike or plateau before step 5000
632
+
633
+ # 2. VRAM usage -- should be stable
634
+ # Red flag: gradual increase (memory leak from activation caching)
635
+
636
+ # 3. Learning rate -- should follow cosine schedule
637
+ # Red flag: NaN or zero (optimizer failure)
638
+ ```
639
+
640
+ Monitor expert utilization at eval steps (add as callback):
641
+
642
+ ```python
643
+ # Quick check: are hot experts still hot?
644
+ # If routing has shifted significantly, the LoRA placement is misaligned.
645
+ # This is informational -- do not retrain Phase 1 mid-run.
646
+ ```
647
+
648
+ ### 5.6b Checkpoint Egress <!-- RED-HAT FIX #8: Added per H9 review — checkpoints must leave the pod after each save -->
649
+
650
+ **Spot instances can terminate at any time.** Checkpoints only on the pod volume are not insurance -- they are gone when the volume is reclaimed. After every checkpoint save, rsync the adapter directory to persistent off-pod storage.
651
+
652
+ ```bash
653
+ # Add as a post-checkpoint callback or cron job running every 30 minutes:
654
+ # Adapter checkpoint is ~3-4 GB (includes optimizer state for resuming)
655
+
656
+ REMOTE_DEST="user@margaret.local:/workspace/daimon_checkpoints/" # or MTH, or B2 bucket
657
+ CHECKPOINT_DIR="/workspace/checkpoints/daimon_v2_foundation"
658
+
659
+ # Sync latest checkpoint off-pod after each save
660
+ rsync -avz --progress \
661
+ "${CHECKPOINT_DIR}/$(ls -td ${CHECKPOINT_DIR}/checkpoint-* | head -1)" \
662
+ "${REMOTE_DEST}"
663
+
664
+ # Also sync critical artifacts
665
+ rsync -avz /workspace/artifacts/hot_experts.json "${REMOTE_DEST}"
666
+ ```
667
+
668
+ **Cost:** Pennies of bandwidth. Converts "total loss on spot preemption" into "resume from last checkpoint - 2000 steps."
669
+
670
+ Add inter-phase disk check:
671
+ ```bash
672
+ # Run between phases to catch disk pressure before it corrupts a save
673
+ DISK_FREE_GB=$(df --output=avail /workspace | tail -1 | awk '{print int($1/1048576)}')
674
+ echo "Disk free: ${DISK_FREE_GB} GB"
675
+ if [ "$DISK_FREE_GB" -lt 80 ]; then
676
+ echo "WARNING: Disk below 80 GB free. Clean stale checkpoints before continuing."
677
+ fi
678
+ ```
679
+
680
+ ### 5.7 Phase 2 Gate
681
+
682
+ Before proceeding to Phase 3:
683
+
684
+ - [ ] Training loss decreased smoothly (no divergence, no NaN)
685
+ - [ ] Final training loss < initial loss by at least 30%
686
+ - [ ] Validation loss is within 10% of training loss (no severe overfitting)
687
+ - [ ] Generate 5 test prompts manually and verify coherent output
688
+ - [ ] VRAM stayed within expected bounds (~90-115 GB peak, alarm at >125 GB sustained) <!-- RED-HAT FIX #6: Corrected from 55-65 GB per H1 review -->
689
+ - [ ] All 5 most recent checkpoints saved successfully
690
+ - [ ] Best checkpoint selected by `load_best_model_at_end=True` via eval_loss <!-- RED-HAT FIX #3: Auto-selected, not manual -->
691
+ - [ ] Best checkpoint copied to `/workspace/checkpoints/daimon_v2_foundation_best/`
692
+ - [ ] Best checkpoint synced off-pod to persistent storage <!-- RED-HAT FIX #8: Per H9 review -->
693
+
694
+ ---
695
+
696
+ ## 6. Phase 3: Thomas Voice Calibration {#6-phase-3-thomas-voice-calibration}
697
+
698
+ **Cost:** ~$1 (30 minutes)
699
+ **VRAM:** ~90-115 GB peak (same as Phase 2) <!-- RED-HAT FIX #6: Corrected from 55-65 GB per H1 review -->
700
+ **Purpose:** Final persona imprint. The Thomas voice demonstrations teach the model HOW to be Daimon -- the specific patterns of question-before-answer, metaphor-forward explanation, constructive challenge, dry wit. This is the separation principle: procedural knowledge, not declarative content.
701
+
702
+ ### 6.1 Data
703
+
704
+ Source: `/home/HumboldtJoker/.coalition/specs/daimon_voice_demonstrations.jsonl`
705
+ Count: 32 procedural demonstrations (from the existing v1 spec, already authored)
706
+ Additional: If available, 8-10 more demonstrations covering edge cases (overwhelming user, technical disagreement, values-laden questions). Target 40 total.
707
+
708
+ Format: Already in `{"messages": [...]}` format. No system prompt -- pure procedural demonstration.
709
+
710
+ ### 6.2 Training Configuration
711
+
712
+ ```python
713
+ # RED-HAT FIX #2: Phase 2 saves adapter-only (adapter_config.json + adapter_model.safetensors),
714
+ # NOT a full model. Must load base model first, then attach adapter with is_trainable=True.
715
+ # Using AutoModelForCausalLM.from_pretrained on an adapter dir will either crash
716
+ # ("no config.json") or load adapter for inference-only (requires_grad=False),
717
+ # causing silent no-op training. (Per C2 review)
718
+
719
+ from peft import PeftModel
720
+
721
+ # Load base model first
722
+ base_model = AutoModelForCausalLM.from_pretrained(
723
+ "/workspace/models/Qwen3.6-35B-A3B", # Base model, NOT the checkpoint
724
+ torch_dtype=torch.bfloat16,
725
+ device_map="cuda",
726
+ )
727
+ tokenizer = AutoTokenizer.from_pretrained("/workspace/models/Qwen3.6-35B-A3B")
728
+
729
+ # Attach Phase 2 adapter with is_trainable=True
730
+ model = PeftModel.from_pretrained(
731
+ base_model,
732
+ "/workspace/checkpoints/daimon_v2_foundation_best",
733
+ is_trainable=True, # CRITICAL: without this, all adapter params are frozen
734
+ )
735
+
736
+ # RED-HAT FIX #2: Verify adapter is trainable (same assert as Phase 2)
737
+ trainable_count = sum(p.numel() for p in model.parameters() if p.requires_grad)
738
+ assert trainable_count > 0, "FATAL: No trainable parameters in Phase 3! Check is_trainable=True."
739
+ expert_lora_count = sum(
740
+ 1 for name, p in model.named_parameters()
741
+ if "experts" in name and "lora" in name and p.requires_grad
742
+ )
743
+ assert expert_lora_count > 0, "FATAL: No LoRA adapters on expert layers in Phase 3!"
744
+ print(f"Phase 3 trainable params: {trainable_count:,}, expert LoRA params: {expert_lora_count}")
745
+
746
+ # RED-HAT FIX #2: Weight-delta assert — hash one LoRA tensor, run one step, verify it changed
747
+ import hashlib
748
+ lora_param = next((name, p) for name, p in model.named_parameters()
749
+ if "lora" in name and p.requires_grad)
750
+ pre_hash = hashlib.sha256(lora_param[1].data.cpu().numpy().tobytes()).hexdigest()
751
+
752
+ training_args = SFTConfig(
753
+ output_dir="/workspace/checkpoints/daimon_v2_calibration",
754
+ per_device_train_batch_size=2, # Smaller batch for tiny dataset
755
+ gradient_accumulation_steps=1, # Effective batch = 2
756
+ num_train_epochs=5, # More epochs for tiny dataset
757
+ learning_rate=1e-5, # 20x lower than Phase 2
758
+ lr_scheduler_type="cosine",
759
+ warmup_ratio=0.1,
760
+ weight_decay=0.01,
761
+ bf16=True,
762
+ max_seq_length=2048,
763
+ logging_steps=5,
764
+ save_steps=20, # Checkpoint every 20 steps
765
+ save_total_limit=10, # Keep all checkpoints (tiny, <5 GB total)
766
+ dataloader_num_workers=0,
767
+ dataset_num_proc=1,
768
+ gradient_checkpointing=True,
769
+ gradient_checkpointing_kwargs={"use_reentrant": False},
770
+ seed=42,
771
+ )
772
+ ```
773
+
774
+ ### 6.3 Why This Works
775
+
776
+ With only 32-40 examples, the risk is memorization rather than generalization. The mitigations:
777
+
778
+ 1. **Low learning rate (1e-5):** 20x lower than Phase 2. The model has already learned the general capability; this stage makes subtle adjustments to the voice distribution.
779
+ 2. **The examples are PROCEDURAL:** They teach response patterns (question-before-answer, assess-then-open), not factual content. Procedural patterns generalize better than factual memorization because they operate on structural templates, not surface forms.
780
+ 3. **5 epochs is intentional:** Each example is seen 5 times. With LoRA r=32 and DoRA, the adapter capacity is sufficient to learn procedural patterns without memorizing specific words. The cosine LR decay means later epochs make smaller adjustments.
781
+ 4. **Recency effect:** This is the last training the model sees. The Thomas voice patterns are the freshest in the gradient signal, which means they have the strongest influence on generation.
782
+
783
+ ### 6.4 Phase 3 Gate
784
+
785
+ - [ ] Training completed without divergence
786
+ - [ ] Generate the 5 held-out Thomas voice evaluation prompts (not in training data)
787
+ - [ ] Manual review: does it sound like Thomas? Apply the daimon_voice_analysis.md criteria:
788
+ - Question-before-answer pattern present?
789
+ - Metaphor usage natural, not forced?
790
+ - Constructive challenge without preachiness?
791
+ - Appropriate register (professional, not casual)?
792
+ - Direct assessment ("The issue is..." "My read is...")?
793
+ - [ ] Compare against Phase 2 output on same prompts -- Phase 3 should be noticeably more "Thomas" without losing capability
794
+ - [ ] Select best checkpoint (likely epoch 3 or 4 based on typical LoRA convergence)
795
+
796
+ ---
797
+
798
+ ## 7. Phase 4: Post-Training {#7-phase-4-post-training}
799
+
800
+ **Cost:** ~$2-4 (1-2 hours) <!-- RED-HAT FIX #14: Honest after trimming eval battery and cutting LASER from critical path -->
801
+ **Purpose:** Two improvements: OGPSA personality capture and eval battery. LASER moved to optional offline experiment (see 7.2). <!-- RED-HAT FIX #1: LASER cut from critical path per C1 review -->
802
+
803
+ ### 7.1 Merge LoRA into Base Weights
804
+
805
+ Before post-training, merge the LoRA adapter into the base model. This is required because (a) vLLM cannot load MoE LoRA adapters, and (b) LASER operates on the merged weights.
806
+
807
+ ```python
808
+ # Merge LoRA -> base weights
809
+ model = model.merge_and_unload()
810
+
811
+ # Save merged model in bf16
812
+ model.save_pretrained(
813
+ "/workspace/models/daimon_v2_merged",
814
+ safe_serialization=True,
815
+ )
816
+ tokenizer.save_pretrained("/workspace/models/daimon_v2_merged")
817
+
818
+ # Verify merged model generates correctly
819
+ inputs = tokenizer("What problem does the dashboard solve?", return_tensors="pt").to("cuda")
820
+ with torch.no_grad():
821
+ output = model.generate(**inputs, max_new_tokens=200, temperature=0.7)
822
+ print(tokenizer.decode(output[0]))
823
+ ```
824
+
825
+ ### 7.2 LASER (Layer-Selective Rank Reduction) — OPTIONAL POST-TRAINING EXPERIMENT
826
+
827
+ <!-- RED-HAT FIX #1: LASER moved from critical path to optional experiment per C1 review.
828
+ THREE bugs in the original:
829
+ 1. SVD truncation was INVERTED — zeroed the LARGEST singular values (principal components)
830
+ instead of the tail. torch.linalg.svd returns values in DESCENDING order.
831
+ The paper's "higher-order components" means components associated with SMALL singular
832
+ values, not the highest-magnitude ones.
833
+ 2. Blanket application to ALL 20,480 matrices contradicts LASER's name (LAyer-SElective
834
+ Rank reduction). The paper's gains come from searching for the specific (layer, matrix)
835
+ where reduction helps; most choices hurt.
836
+ 3. "+20-30pp" was misquoted — those gains are on narrow factual-recall evals at the single
837
+ best layer on older models (GPT-J). Realistic upside on a 2026 instruct MoE: 0 to +1-2pp.
838
+
839
+ LASER is NOT free: 20,480 CPU SVDs of 2048x512 fp32 matrices takes 1-3 hours of pod time.
840
+ The rollback plan already calls it "a free bonus, not a requirement" — it isn't free and
841
+ as originally written it was negative-value (would have destroyed the trained model).
842
+
843
+ RECOMMENDED: Run LASER locally on Margaret after downloading merged weights ($0 cost).
844
+ Apply to ONE (layer, matrix) at a time, eval after each, keep only if it helps. -->
845
+
846
+ **Paper:** arXiv:2312.13558
847
+ **Code:** https://github.com/pratyushasharma/laser
848
+ **Status:** Optional post-budget experiment. Do NOT run on the cloud instance during the primary training run.
849
+ **Realistic upside:** 0 to +1-2pp on narrow factual-recall benchmarks at the single best (layer, matrix). Not a broad reasoning improvement.
850
+
851
+ **WARNING:** The original code in this spec was inverted and would have destroyed the trained model. The corrected version below keeps the TOP singular values and truncates the TAIL (small singular values = high-order noise).
852
+
853
+ ```python
854
+ # CORRECTED LASER — run locally on Margaret, not on cloud pod
855
+ # RED-HAT FIX #1: Fixed SVD direction. S_modified[n_keep:] = 0 keeps top, drops tail.
856
+
857
+ import torch
858
+
859
+ def laser_reduce(weight_matrix, keep_fraction=0.95):
860
+ """Keep top fraction of singular values, zero out the tail (noise).
861
+
862
+ LASER's insight: small singular values of MLP weight matrices often
863
+ encode noise. Removing them can improve factual recall.
864
+
865
+ NOTE: torch.linalg.svd returns singular values in DESCENDING order.
866
+ We keep the first n_keep values (largest) and zero the rest (smallest).
867
+ """
868
+ original_dtype = weight_matrix.dtype
869
+ W = weight_matrix.float() # SVD requires float32
870
+ U, S, Vh = torch.linalg.svd(W, full_matrices=False)
871
+ n_keep = max(1, int(len(S) * keep_fraction))
872
+ S_modified = S.clone()
873
+ S_modified[n_keep:] = 0.0 # Zero out TAIL (small values = noise), keep TOP
874
+ W_modified = U @ torch.diag(S_modified) @ Vh
875
+ return W_modified.to(original_dtype)
876
+
877
+ # DO NOT apply blanket to all 20,480 matrices.
878
+ # Apply to ONE (layer, matrix) at a time, eval after each.
879
+ # Example: test on layer 20, gate_proj only:
880
+ # new_weight = laser_reduce(expert.gate_proj.weight.data, keep_fraction=0.95)
881
+ # Run eval. If improved, keep. If not, revert.
882
+ ```
883
+
884
+ **LASER protocol (if attempting):**
885
+ 1. Download merged weights to Margaret (local, $0)
886
+ 2. For each candidate (layer_idx, proj_name) pair, in a systematic sweep:
887
+ a. Apply `laser_reduce` with `keep_fraction=0.95`
888
+ b. Run a quick eval subset (100-200 examples)
889
+ c. Keep only if eval improves over baseline
890
+ 3. Try `keep_fraction` values: 0.93, 0.95, 0.97
891
+ 4. Apply ONLY to expert FFN (gate_proj, up_proj). Do NOT apply to attention or down_proj.
892
+
893
+ ### 7.3 OGPSA Personality Capture
894
+
895
+ Per ogpsa_persona_validation_v3.1.md, OGPSA extracts the personality subspace from the model's activations on persona-loaded text.
896
+
897
+ ```python
898
+ # OGPSA Phase 1: Extract persona subspace
899
+ # Use the Thomas voice demonstrations as persona-loaded prompts
900
+ # Plus 10 general QA examples as contrast
901
+
902
+ # 1. Run persona-loaded examples through the model
903
+ # 2. Capture residual stream activations at all layers
904
+ # 3. SVD on the persona-loaded activation matrix
905
+ # 4. Top 16 components = persona subspace
906
+ # 5. Save as ogpsa_daimon_v2.json
907
+
908
+ # The OGPSA capture script is at ogpsa_red_team_code.md
909
+ # Adapt for Qwen3.6 architecture (40 layers, not 48)
910
+ ```
911
+
912
+ Output: `ogpsa_daimon_v2.json` -- the personality geometry for drift monitoring.
913
+
914
+ ### 7.4 Evaluation Battery
915
+
916
+ Run on the merged model (before any optional LASER experiments):
917
+
918
+ <!-- RED-HAT FIX #7: Removed truthfulqa_mc2 from eval battery per H8 review.
919
+ TruthfulQA is in the training data (upsampled 15x). Including it in eval
920
+ measures memorization, not anti-sycophancy generalization. Rely on the
921
+ custom 10-scenario anti-sycophancy eval instead (well-designed, not contaminated). -->
922
+
923
+ <!-- RED-HAT FIX #14: Trimmed eval battery per H10 review.
924
+ Removed arc_easy (redundant with arc_challenge). Increased batch_size to 16.
925
+ Added --limit 1500 for forgetting detection (5% regression detector, not leaderboard).
926
+ Single pass only (LASER moved to optional offline experiment). -->
927
+
928
+ ```bash
929
+ # Using lm-evaluation-harness (install in separate venv per H5 review)
930
+ # python -m venv /workspace/lm_eval_venv && source /workspace/lm_eval_venv/bin/activate
931
+ # pip install lm_eval[hf]
932
+
933
+ # Baseline capabilities (forgetting detection)
934
+ lm_eval --model hf \
935
+ --model_args pretrained=/workspace/models/daimon_v2_merged,dtype=bfloat16 \
936
+ --tasks mmlu,gsm8k,hellaswag,arc_challenge,winogrande \
937
+ --batch_size 16 \
938
+ --limit 1500 \
939
+ --output_path /workspace/eval_results/daimon_v2/
940
+
941
+ # Compare against base model eval (run on unmodified Qwen3.6-35B-A3B for comparison)
942
+ ```
943
+
944
+ **Custom Daimon-specific evals:**
945
+
946
+ | Eval | Method | Pass Criteria |
947
+ |---|---|---|
948
+ | **Voice attribution** | 5 evaluators, 20 blind response pairs (Daimon vs generic Qwen3.6) | >80% correct attribution to "Thomas-like" |
949
+ | **Daemon function** | 10 scenarios where something is wrong. Does Daimon flag it? | >7/10 flags the concern |
950
+ | **Anti-sycophancy** | 10 scenarios with bad ideas presented as good ones. Does Daimon push back? | >7/10 pushes back |
951
+ | **Anti-nag** | 10 scenarios where everything is fine. Does Daimon stay quiet? | >8/10 does not volunteer unnecessary concerns |
952
+ | **Tool use** | 10 multi-step agentic tasks with function calling | >8/10 correct tool selection and sequencing |
953
+ | **Register calibration** | 5 prompts across different contexts (client, internal, public) | Appropriate register shift in >4/5 |
954
+
955
+ ### 7.5 Phase 4 Gate
956
+
957
+ - [ ] LoRA merged successfully (model generates coherent text)
958
+ - [ ] LASER: OPTIONAL — moved to post-budget local experiment on Margaret <!-- RED-HAT FIX #1: Per C1 review -->
959
+ - [ ] OGPSA personality subspace captured (ogpsa_daimon_v2.json)
960
+ - [ ] Eval battery results saved (without truthfulqa_mc2 — contaminated) <!-- RED-HAT FIX #7: Per H8 review -->
961
+ - [ ] MMLU within 5% of base model (no catastrophic forgetting)
962
+ - [ ] Custom Daimon evals pass criteria
963
+ - [ ] All artifacts synced off-pod to persistent storage <!-- RED-HAT FIX #8: Per H9 review -->
964
+
965
+ ---
966
+
967
+ ## 8. Phase 5: Export and Validation {#8-phase-5-export-and-validation}
968
+
969
+ **Cost:** ~$2 (1 hour)
970
+ **Purpose:** Convert to serving format and final validation
971
+
972
+ ### 8.1 Export to GGUF
973
+
974
+ ```bash
975
+ # Clone llama.cpp if not present
976
+ git clone https://github.com/ggerganov/llama.cpp /workspace/llama.cpp
977
+ cd /workspace/llama.cpp && make -j$(nproc)
978
+
979
+ # Convert to GGUF (bf16 first, then quantize)
980
+ # RED-HAT FIX #1: Export from merged model, not LASER'd model (LASER is optional/offline now)
981
+ python convert_hf_to_gguf.py /workspace/models/daimon_v2_merged \
982
+ --outfile /workspace/models/daimon_v2.gguf \
983
+ --outtype bf16
984
+
985
+ # Quantize to Q4_K_M for serving on Margaret (Mac Studio)
986
+ ./llama-quantize /workspace/models/daimon_v2.gguf \
987
+ /workspace/models/daimon_v2_Q4_K_M.gguf Q4_K_M
988
+
989
+ # Also export Q5_K_M for higher-quality serving if Margaret has headroom
990
+ ./llama-quantize /workspace/models/daimon_v2.gguf \
991
+ /workspace/models/daimon_v2_Q5_K_M.gguf Q5_K_M
992
+ ```
993
+
994
+ ### 8.2 Export to GPTQ/AWQ (Optional, for GPU serving)
995
+
996
+ ```bash
997
+ # If serving on GPU (e.g., via vLLM), export GPTQ
998
+ pip install auto-gptq
999
+ python -c "
1000
+ from transformers import AutoModelForCausalLM, AutoTokenizer
1001
+ from auto_gptq import AutoGPTQForCausalLM, BaseQuantizeConfig
1002
+
1003
+ quantize_config = BaseQuantizeConfig(bits=4, group_size=128, desc_act=False)
1004
+ model = AutoGPTQForCausalLM.from_pretrained(
1005
+ '/workspace/models/daimon_v2_merged', # RED-HAT FIX #1: Use merged model, not LASER'd
1006
+ quantize_config=quantize_config,
1007
+ )
1008
+ model.quantize() # Uses calibration data
1009
+ model.save_quantized('/workspace/models/daimon_v2_gptq')
1010
+ "
1011
+ ```
1012
+
1013
+ ### 8.3 Download Artifacts
1014
+
1015
+ Before terminating the cloud instance, download all critical artifacts:
1016
+
1017
+ ```bash
1018
+ # Priority 1: The trained model (pick your serving format)
1019
+ # GGUF Q4_K_M: ~18 GB
1020
+ # GGUF Q5_K_M: ~22 GB
1021
+ # bf16 merged: ~70 GB (only if you have bandwidth/storage)
1022
+
1023
+ # Priority 2: LoRA checkpoints (for potential continued training)
1024
+ # Foundation best: ~500 MB
1025
+ # Calibration best: ~500 MB
1026
+
1027
+ # Priority 3: Metadata
1028
+ # hot_experts.json
1029
+ # ogpsa_daimon_v2.json
1030
+ # eval_results/ directory
1031
+ # training logs
1032
+ ```
1033
+
1034
+ ---
1035
+
1036
+ ## 9. Budget Allocation {#9-budget-allocation}
1037
+
1038
+ ### Primary Run Budget
1039
+
1040
+ <!-- RED-HAT FIX #12: Corrected budget to realistic raw-PEFT throughput per C4 review.
1041
+ Original assumed Unsloth kernel speeds ($20-28 for Phase 2) but shipped code uses
1042
+ raw PEFT + Transformers. Throughput probe (Section 5.5b) gates the actual cost. -->
1043
+
1044
+ | Phase | Estimated Time | Cost | Running Total |
1045
+ |---|---|---|---|
1046
+ | 0: Preflight (incl. backward-pass smoke test) | 30-60 min | $1.50-2.50 | $1.50-2.50 |
1047
+ | 1: Expert profiling | 30 min | $1.00 | $2.50-3.50 |
1048
+ | 2: Foundation training (gated by throughput probe) | 11-16 hrs | $22-32 | $24.50-35.50 |
1049
+ | 3: Thomas voice calibration | 30 min | $1.00 | $25.50-36.50 |
1050
+ | 4: Post-training (OGPSA + eval; LASER cut from critical path) | 1-2 hrs | $2-4 | $27.50-40.50 |
1051
+ | 5: Export and download | 1 hr | $2-3 | $29.50-43.50 |
1052
+ | **Total primary run** | | | **$30-45** |
1053
+
1054
+ ### Reserve Budget
1055
+
1056
+ | Use | Estimated Cost |
1057
+ |---|---|
1058
+ | **Available reserve** | **$52-67** |
1059
+ | Hyperparameter sweep run (if voice quality insufficient) | $22-32 |
1060
+ | LR sweep (3 runs at 1e-4, 2e-4, 5e-4) on 10K subset | $6-9 |
1061
+ | Rank sweep (r=16 vs r=32 vs r=64) on 10K subset | $6-9 |
1062
+ | LASER sweep (offline on Margaret, $0) | $0 |
1063
+ | Full re-run with tuned hyperparameters | $22-32 |
1064
+ | Emergency buffer | ~$5-15 |
1065
+
1066
+ ### Cost Optimization Tips
1067
+
1068
+ 1. **Use spot/interruptible instances** if the provider offers them. Save checkpoints every 2000 steps to survive interruptions.
1069
+ 2. **Download the model to persistent storage** before the first run. Model download is ~70 GB and takes 15-30 minutes depending on bandwidth. Do not pay H200 rates for downloading.
1070
+ 3. **Kill the instance between runs** if doing multi-day work. $2/hr idle is $48/day wasted.
1071
+ 4. **Hyperparameter sweeps on subsets first.** A 10K-sample sweep costs ~$1-2 per run. Do not spend $25 on a full run with untested hyperparameters.
1072
+
1073
+ ---
1074
+
1075
+ ## 10. Rollback Plan {#10-rollback-plan}
1076
+
1077
+ ### Per-Phase Rollback
1078
+
1079
+ | Phase | If it fails... | Rollback action | Cost to recover |
1080
+ |---|---|---|---|
1081
+ | 0: Preflight | Hardware/software mismatch | Fix environment or change provider. No cost wasted (no training started). | $0 |
1082
+ | 1: Expert profiling | Profiling script errors | Debug and re-run. Costs ~$0.50. Or fall back to uniform LoRA (all experts) as Tier 1 baseline. | $0.50 |
1083
+ | 2: Foundation training | Divergence at step N | Roll back to checkpoint at step N-2000. Adjust LR (halve it) and resume from that checkpoint. Resume costs proportional to remaining steps only. | Variable |
1084
+ | 2: Foundation training | OOM/SIGKILL | Reduce batch_size to 2 (effective batch = 8). Or reduce max_seq_length to 1024. Re-run from last checkpoint. | Variable |
1085
+ | 2: Foundation training | Poor quality at completion | Try: (a) different LR, (b) longer training (2 epochs), (c) different data mix, (d) fall back to ESFT instead of HELLoRA. Each iteration is $22-32. | $22-32 | <!-- RED-HAT FIX #12: Corrected from $20-28 -->
1086
+ | 3: Thomas calibration | Voice too weak | Increase epochs to 10, or increase LR to 5e-5. Re-run from Phase 2 best checkpoint. | $1 |
1087
+ | 3: Thomas calibration | Voice too strong (memorized) | Decrease epochs to 3, or decrease LR to 5e-6. Re-run from Phase 2 best checkpoint. | $1 |
1088
+ | 4: LASER | Quality degradation | LASER is optional and offline (run on Margaret). Skip entirely if no improvement measured per-layer. | $0 | <!-- RED-HAT FIX #1 -->
1089
+ | 4: OGPSA | Capture fails | Re-run capture. If architectural mismatch, adapt the OGPSA script for 40-layer Qwen3.6 (v3.1 was written for a 48-layer model). | $0.50 |
1090
+ | 5: Export | GGUF conversion fails | Check llama.cpp supports Qwen3.6 architecture. May need to update llama.cpp to latest. | $0.50 |
1091
+
1092
+ ### Complete Pipeline Rollback
1093
+
1094
+ If the entire pipeline produces unsatisfactory results after the primary run:
1095
+
1096
+ 1. **Analyze eval results** to identify which capability is deficient (voice, tool use, anti-sycophancy, reasoning).
1097
+ 2. **Targeted fix options:**
1098
+ - Voice weak: More Thomas demonstrations + higher LR in Phase 3
1099
+ - Tool use weak: Upsample xlam data to 2x in Phase 2 mix
1100
+ - Anti-sycophancy weak: Upsample TruthfulQA to 25x; consider adding SAA dataset
1101
+ - Reasoning weak: Upsample hermes traces to 5x; consider adding GSM8K training data
1102
+ 3. **Full re-run** with adjusted data mix: ~$30-45 from reserve budget. <!-- RED-HAT FIX #12: Corrected from $27-37 -->
1103
+ 4. **Method pivot** if HELLoRA ceiling is too low: Switch to ESFT (full fine-tune hot experts) or BAdam (block-coordinate). This is a larger change and costs one full run.
1104
+
1105
+ ### Checkpoint Strategy
1106
+
1107
+ Checkpoints are the most important insurance:
1108
+
1109
+ - Phase 2 saves every 2000 steps (5 checkpoints retained)
1110
+ - Phase 3 saves every 20 steps (all checkpoints retained -- tiny files)
1111
+ - The BEST checkpoint from each phase is copied to a separate directory
1112
+ - All checkpoints include the full LoRA adapter state + optimizer state (allows resuming training)
1113
+ - **Checkpoints synced off-pod after every save** (see Section 5.6b) <!-- RED-HAT FIX #8: Per H9 review -->
1114
+ - Download at least the best checkpoints to local storage before terminating the instance
1115
+
1116
+ ---
1117
+
1118
+ ## Appendix A: Environment Setup Script {#appendix-a}
1119
+
1120
+ ```bash
1121
+ #!/bin/bash
1122
+ # daimon_setup.sh -- Run this FIRST on a fresh H200 instance
1123
+ set -euo pipefail
1124
+
1125
+ echo "=== Daimon Training Pipeline: Environment Setup ==="
1126
+
1127
+ # 1. System checks
1128
+ echo "--- Hardware checks ---"
1129
+ nvidia-smi --query-gpu=name,memory.total --format=csv,noheader
1130
+ CONTAINER_RAM=$(cat /sys/fs/cgroup/memory.max 2>/dev/null || cat /sys/fs/cgroup/memory/memory.limit_in_bytes 2>/dev/null)
1131
+ echo "Container RAM: $(echo "$CONTAINER_RAM / 1024 / 1024 / 1024" | bc) GB"
1132
+
1133
+ # 2. Set environment variables BEFORE any Python imports
1134
+ export UNSLOTH_COMPILE_DISABLE=1
1135
+ export PYTORCH_CUDA_ALLOC_CONF="expandable_segments:True"
1136
+ export TOKENIZERS_PARALLELISM=false
1137
+ export UNSLOTH_DISABLE_FAST_GENERATION=1
1138
+
1139
+ # Persist for all subsequent shells
1140
+ cat >> ~/.bashrc << 'ENVEOF'
1141
+ export UNSLOTH_COMPILE_DISABLE=1
1142
+ export PYTORCH_CUDA_ALLOC_CONF="expandable_segments:True"
1143
+ export TOKENIZERS_PARALLELISM=false
1144
+ export UNSLOTH_DISABLE_FAST_GENERATION=1
1145
+ ENVEOF
1146
+
1147
+ # 3. Install/update packages
1148
+ # RED-HAT FIX #9: Pin exact versions and install torch FIRST per H5 review.
1149
+ # Original had version floors (>=) with torch installed LAST, causing resolution chaos.
1150
+ # Install order matters: torch first (pinned to cu124), then core ML stack, unsloth --no-deps last.
1151
+ # lm_eval in separate venv to avoid dependency conflicts.
1152
+
1153
+ # Ensure conda is on PATH (some cloud images need this)
1154
+ export PATH="/opt/conda/bin:$PATH"
1155
+
1156
+ # Set HF cache to /workspace to avoid filling root partition (RED-HAT FIX #5, C5)
1157
+ export HF_HOME=/workspace/.hf
1158
+ export HF_DATASETS_CACHE=/workspace/.hf/datasets
1159
+ mkdir -p "$HF_HOME" "$HF_DATASETS_CACHE"
1160
+ cat >> ~/.bashrc << 'HFEOF'
1161
+ export HF_HOME=/workspace/.hf
1162
+ export HF_DATASETS_CACHE=/workspace/.hf/datasets
1163
+ HFEOF
1164
+
1165
+ pip install --upgrade pip
1166
+
1167
+ # Step 1: Install torch FIRST with CUDA 12.4 (pinned version)
1168
+ pip install torch==2.7.1 torchvision==0.22.1 torchaudio==2.7.1 --index-url https://download.pytorch.org/whl/cu124
1169
+
1170
+ # Step 2: Verify torch installed correctly BEFORE proceeding
1171
+ python -c "import torch; assert torch.cuda.is_available(), 'CUDA not available after torch install'; print(f'torch {torch.__version__} OK')"
1172
+
1173
+ # Step 3: Install core ML stack (pinned versions, no torch re-resolution)
1174
+ pip install transformers==5.0.2
1175
+ pip install peft==0.15.2
1176
+ pip install trl==0.18.1
1177
+ pip install datasets==3.6.0
1178
+ pip install accelerate==1.7.0
1179
+ pip install bitsandbytes==0.46.0 # Required by unsloth even if we don't use 4-bit
1180
+ pip install sentencepiece protobuf
1181
+
1182
+ # Step 4: Install unsloth LAST with --no-deps to avoid overwriting pinned versions
1183
+ pip install "unsloth>=0.1.47" --no-deps
1184
+
1185
+ # Step 5: Verify no dependency conflicts
1186
+ pip check || echo "WARNING: pip check found conflicts — review before proceeding"
1187
+
1188
+ # Step 6: Freeze the resolved environment for reproducibility
1189
+ pip freeze > /workspace/artifacts/requirements.lock
1190
+ echo "Locked environment saved to /workspace/artifacts/requirements.lock"
1191
+
1192
+ # Step 7: Install lm_eval in a SEPARATE venv to avoid conflicts (RED-HAT FIX #9, H5)
1193
+ python -m venv /workspace/lm_eval_venv
1194
+ /workspace/lm_eval_venv/bin/pip install --upgrade pip
1195
+ /workspace/lm_eval_venv/bin/pip install "lm_eval[hf]"
1196
+
1197
+ # 4. Verify installations
1198
+ python -c "
1199
+ import os
1200
+ os.environ['UNSLOTH_COMPILE_DISABLE'] = '1'
1201
+ os.environ['PYTORCH_CUDA_ALLOC_CONF'] = 'expandable_segments:True'
1202
+ os.environ['TOKENIZERS_PARALLELISM'] = 'false'
1203
+ os.environ['UNSLOTH_DISABLE_FAST_GENERATION'] = '1'
1204
+
1205
+ import torch
1206
+ print(f'PyTorch {torch.__version__}, CUDA available: {torch.cuda.is_available()}')
1207
+ if torch.cuda.is_available():
1208
+ print(f'GPU: {torch.cuda.get_device_name(0)}')
1209
+ print(f'VRAM: {torch.cuda.get_device_properties(0).total_memory / 1e9:.1f} GB')
1210
+
1211
+ import transformers; print(f'Transformers {transformers.__version__}')
1212
+ import peft; print(f'PEFT {peft.__version__}')
1213
+ import unsloth; print(f'Unsloth {unsloth.__version__}')
1214
+ print('All imports successful.')
1215
+ "
1216
+
1217
+ # 5. Create workspace directories
1218
+ mkdir -p /workspace/models
1219
+ mkdir -p /workspace/data
1220
+ mkdir -p /workspace/checkpoints
1221
+ mkdir -p /workspace/eval_results
1222
+ mkdir -p /workspace/artifacts
1223
+
1224
+ echo "=== Setup complete ==="
1225
+ ```
1226
+
1227
+ ---
1228
+
1229
+ ## Appendix B: Expert Profiling Script {#appendix-b}
1230
+
1231
+ ```python
1232
+ #!/usr/bin/env python3
1233
+ """
1234
+ daimon_expert_profiling.py
1235
+ Phase 1: Profile which experts activate for Daimon's training data.
1236
+ Uses ESFT-style profiling with MAN (Mean Activation Norm) scoring.
1237
+
1238
+ Reference: ESFT (arXiv:2407.01906), Unified Expert Scoring (arXiv:2606.15716)
1239
+ """
1240
+
1241
+ import os
1242
+ os.environ["UNSLOTH_COMPILE_DISABLE"] = "1"
1243
+ os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "expandable_segments:True"
1244
+ os.environ["TOKENIZERS_PARALLELISM"] = "false"
1245
+ os.environ["UNSLOTH_DISABLE_FAST_GENERATION"] = "1"
1246
+
1247
+ import json
1248
+ import torch
1249
+ import numpy as np
1250
+ from collections import defaultdict
1251
+ from transformers import AutoModelForCausalLM, AutoTokenizer
1252
+
1253
+ MODEL_PATH = "/workspace/models/Qwen3.6-35B-A3B"
1254
+ OUTPUT_PATH = "/workspace/artifacts/hot_experts.json"
1255
+ NUM_LAYERS = 40
1256
+ NUM_EXPERTS = 256
1257
+ TOP_K = 8 # Qwen3.6 routes to top-8 experts
1258
+
1259
+ # Representative samples from each data source
1260
+ # Adjust paths to match your data layout
1261
+ SAMPLE_SOURCES = {
1262
+ "sonnet_voice": ("/workspace/data/sonnet_voice_distill/", 8),
1263
+ "prosocial": ("/workspace/data/prosocial_dialogue/", 8),
1264
+ "tool_use": ("/workspace/data/xlam_function_calling/", 8),
1265
+ "reasoning": ("/workspace/data/hermes_agent_traces/", 4),
1266
+ "thomas_voice": ("/workspace/data/thomas_voice_demos/daimon_voice_demonstrations.jsonl", 4),
1267
+ }
1268
+
1269
+ def load_samples():
1270
+ """Load representative samples from each data source."""
1271
+ samples = []
1272
+ for source_name, (path, count) in SAMPLE_SOURCES.items():
1273
+ # Load first `count` examples from each source
1274
+ # Adapt this loader to match your data format
1275
+ if path.endswith(".jsonl"):
1276
+ with open(path) as f:
1277
+ for i, line in enumerate(f):
1278
+ if i >= count:
1279
+ break
1280
+ data = json.loads(line)
1281
+ # Extract the full conversation as a single string
1282
+ text = " ".join(m["content"] for m in data["messages"])
1283
+ samples.append(text)
1284
+ else:
1285
+ # Directory of files -- load from first file found
1286
+ import glob
1287
+ files = sorted(glob.glob(os.path.join(path, "*.jsonl")))
1288
+ if not files:
1289
+ files = sorted(glob.glob(os.path.join(path, "*.json")))
1290
+ if files:
1291
+ loaded = 0
1292
+ for fpath in files:
1293
+ with open(fpath) as f:
1294
+ for line in f:
1295
+ if loaded >= count:
1296
+ break
1297
+ data = json.loads(line)
1298
+ if "messages" in data:
1299
+ text = " ".join(m["content"] for m in data["messages"])
1300
+ elif "text" in data:
1301
+ text = data["text"]
1302
+ elif "conversations" in data:
1303
+ text = " ".join(c.get("value", "") for c in data["conversations"])
1304
+ else:
1305
+ text = str(data)
1306
+ samples.append(text)
1307
+ loaded += 1
1308
+ if loaded >= count:
1309
+ break
1310
+ print(f"Loaded {len(samples)} profiling samples")
1311
+ return samples
1312
+
1313
+
1314
+ def profile_experts(model, tokenizer, samples):
1315
+ """
1316
+ Profile expert activation patterns using Mean Activation Norm (MAN).
1317
+
1318
+ For each MoE layer, record the activation norm of each expert's output
1319
+ when it is selected by the router. The Mean Activation Norm across all
1320
+ samples gives the expert importance score.
1321
+ """
1322
+ # Storage: layer_idx -> expert_idx -> list of activation norms
1323
+ activation_norms = defaultdict(lambda: defaultdict(list))
1324
+
1325
+ # Register hooks on MoE layers
1326
+ hooks = []
1327
+
1328
+ def make_hook(layer_idx):
1329
+ def hook_fn(module, input, output):
1330
+ # The MoE module's output includes routing information
1331
+ # We need to capture which experts were selected and their output norms
1332
+ # This hook structure depends on the exact Qwen3.6 implementation
1333
+ #
1334
+ # For Qwen3.6, the MoE block routes tokens and returns the weighted sum.
1335
+ # To get per-expert activation norms, we need to hook deeper -- at the
1336
+ # router level and individual expert level.
1337
+ pass
1338
+ return hook_fn
1339
+
1340
+ # Alternative approach: direct router inspection
1341
+ # Hook the router to get expert selection, then measure expert output norms
1342
+ for layer_idx in range(NUM_LAYERS):
1343
+ moe_layer = model.model.layers[layer_idx].mlp
1344
+
1345
+ # Hook the gate (router) to capture routing decisions
1346
+ def make_router_hook(l_idx):
1347
+ def hook_fn(module, input, output):
1348
+ # output is the router logits or routing weights
1349
+ # For Qwen3.6: output shape is [batch*seq_len, num_experts]
1350
+ if isinstance(output, tuple):
1351
+ router_logits = output[0]
1352
+ else:
1353
+ router_logits = output
1354
+
1355
+ # Get top-k expert indices
1356
+ topk_vals, topk_indices = torch.topk(router_logits, TOP_K, dim=-1)
1357
+
1358
+ # Record which experts were selected and their gate values
1359
+ for expert_idx in range(NUM_EXPERTS):
1360
+ mask = (topk_indices == expert_idx).any(dim=-1)
1361
+ if mask.any():
1362
+ # Use gate value as proxy for activation norm
1363
+ # (actual activation norm requires hooking each expert separately)
1364
+ gate_vals = topk_vals[mask]
1365
+ mean_gate = gate_vals.mean().item()
1366
+ activation_norms[l_idx][expert_idx].append(mean_gate)
1367
+
1368
+ return hook_fn
1369
+
1370
+ hook = moe_layer.gate.register_forward_hook(make_router_hook(layer_idx))
1371
+ hooks.append(hook)
1372
+
1373
+ # Run inference on all samples
1374
+ model.eval()
1375
+ with torch.no_grad():
1376
+ for i, text in enumerate(samples):
1377
+ inputs = tokenizer(
1378
+ text, return_tensors="pt", truncation=True, max_length=2048
1379
+ ).to("cuda")
1380
+ _ = model(**inputs)
1381
+ if i % 8 == 0:
1382
+ print(f"Profiled sample {i+1}/{len(samples)}")
1383
+
1384
+ # Remove hooks
1385
+ for hook in hooks:
1386
+ hook.remove()
1387
+
1388
+ return activation_norms
1389
+
1390
+
1391
+ def compute_hot_experts(activation_norms, coverage_threshold=0.80):
1392
+ """
1393
+ Compute hot experts per layer using Mean Activation Norm.
1394
+
1395
+ For each layer, sort experts by MAN score and select the minimum set
1396
+ that covers >= coverage_threshold of total activation norm.
1397
+ """
1398
+ hot_experts = {}
1399
+
1400
+ for layer_idx in range(NUM_LAYERS):
1401
+ expert_scores = {}
1402
+ for expert_idx in range(NUM_EXPERTS):
1403
+ norms = activation_norms[layer_idx].get(expert_idx, [])
1404
+ if norms:
1405
+ expert_scores[expert_idx] = np.mean(norms)
1406
+ else:
1407
+ expert_scores[expert_idx] = 0.0
1408
+
1409
+ # Sort by MAN score descending
1410
+ sorted_experts = sorted(expert_scores.items(), key=lambda x: x[1], reverse=True)
1411
+ total_norm = sum(score for _, score in sorted_experts)
1412
+
1413
+ if total_norm == 0:
1414
+ print(f"WARNING: Layer {layer_idx} has zero total activation norm!")
1415
+ hot_experts[layer_idx] = list(range(TOP_K)) # Fallback
1416
+ continue
1417
+
1418
+ # Select experts until coverage threshold is met
1419
+ cumulative = 0.0
1420
+ hot = []
1421
+ for expert_idx, score in sorted_experts:
1422
+ hot.append(expert_idx)
1423
+ cumulative += score
1424
+ if cumulative / total_norm >= coverage_threshold:
1425
+ break
1426
+
1427
+ hot_experts[layer_idx] = sorted(hot)
1428
+ print(f"Layer {layer_idx}: {len(hot)} hot experts "
1429
+ f"({100*len(hot)/NUM_EXPERTS:.1f}% of 256, "
1430
+ f"covering {100*cumulative/total_norm:.1f}% of activation norm)")
1431
+
1432
+ return hot_experts
1433
+
1434
+
1435
+ def main():
1436
+ print("=== Phase 1: Expert Profiling ===")
1437
+
1438
+ # Load model for inference only
1439
+ print("Loading model...")
1440
+ tokenizer = AutoTokenizer.from_pretrained(MODEL_PATH)
1441
+ model = AutoModelForCausalLM.from_pretrained(
1442
+ MODEL_PATH,
1443
+ torch_dtype=torch.bfloat16,
1444
+ device_map="cuda",
1445
+ )
1446
+
1447
+ # Load profiling samples
1448
+ samples = load_samples()
1449
+
1450
+ # Profile
1451
+ print("Profiling expert activations...")
1452
+ activation_norms = profile_experts(model, tokenizer, samples)
1453
+
1454
+ # Compute hot experts
1455
+ print("\nComputing hot experts (80% coverage threshold)...")
1456
+ hot_experts = compute_hot_experts(activation_norms, coverage_threshold=0.80)
1457
+
1458
+ # Summary statistics
1459
+ counts = [len(v) for v in hot_experts.values()]
1460
+ print(f"\nHot experts per layer: min={min(counts)}, max={max(counts)}, "
1461
+ f"mean={np.mean(counts):.1f}, median={np.median(counts):.1f}")
1462
+
1463
+ total_hot = sum(counts)
1464
+ total_possible = NUM_LAYERS * NUM_EXPERTS
1465
+ print(f"Total hot expert slots: {total_hot}/{total_possible} "
1466
+ f"({100*total_hot/total_possible:.1f}%)")
1467
+
1468
+ # Save
1469
+ # Convert int keys to strings for JSON
1470
+ hot_experts_json = {str(k): v for k, v in hot_experts.items()}
1471
+ with open(OUTPUT_PATH, "w") as f:
1472
+ json.dump(hot_experts_json, f, indent=2)
1473
+ print(f"\nSaved to {OUTPUT_PATH}")
1474
+
1475
+ # Sanity checks
1476
+ for layer_idx, experts in hot_experts.items():
1477
+ if len(experts) < 5:
1478
+ print(f"WARNING: Layer {layer_idx} has only {len(experts)} hot experts. "
1479
+ "Consider lowering coverage threshold.")
1480
+ if len(experts) > 100:
1481
+ print(f"WARNING: Layer {layer_idx} has {len(experts)} hot experts. "
1482
+ "Routing may be too diffuse. Check for profiling errors.")
1483
+
1484
+ del model
1485
+ torch.cuda.empty_cache()
1486
+ print("\n=== Expert profiling complete ===")
1487
+
1488
+
1489
+ if __name__ == "__main__":
1490
+ main()
1491
+ ```
1492
+
1493
+ **Important notes on the profiling script:**
1494
+
1495
+ 1. The router hook structure assumes Qwen3.6's MoE implementation exposes router logits through the gate module. The exact module hierarchy may differ -- inspect the model's `named_modules()` output to find the correct hook points.
1496
+ 2. Using gate values as a proxy for activation norms is an approximation. For exact MAN scoring per arXiv:2606.15716, hook each expert's output layer and measure the L2 norm of the output tensor. This requires more memory but gives more accurate scores.
1497
+ 3. The 80% coverage threshold is the starting point. If it selects >50 experts per layer, raise to 85%. If it selects <15, lower to 75%.
1498
+
1499
+ ---
1500
+
1501
+ ## Appendix C: Training Script {#appendix-c}
1502
+
1503
+ ```python
1504
+ #!/usr/bin/env python3
1505
+ """
1506
+ daimon_train.py
1507
+ Phase 2 + Phase 3: Foundation training + Thomas voice calibration.
1508
+
1509
+ Usage:
1510
+ python daimon_train.py --phase foundation
1511
+ python daimon_train.py --phase calibration --checkpoint /path/to/best
1512
+ """
1513
+
1514
+ import os
1515
+ os.environ["UNSLOTH_COMPILE_DISABLE"] = "1"
1516
+ os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "expandable_segments:True"
1517
+ os.environ["TOKENIZERS_PARALLELISM"] = "false"
1518
+ os.environ["UNSLOTH_DISABLE_FAST_GENERATION"] = "1"
1519
+
1520
+ import json
1521
+ import argparse
1522
+ import torch
1523
+ from datasets import load_dataset, interleave_datasets, Dataset
1524
+ from peft import LoraConfig, get_peft_model, PeftModel
1525
+ from transformers import AutoModelForCausalLM, AutoTokenizer
1526
+ from trl import SFTTrainer, SFTConfig
1527
+
1528
+
1529
+ def build_target_modules(hot_experts_path):
1530
+ """Build HELLoRA target module list from expert profiling results."""
1531
+ with open(hot_experts_path) as f:
1532
+ hot_experts = json.load(f)
1533
+
1534
+ target_modules = []
1535
+
1536
+ # Attention projections (all layers)
1537
+ for layer_idx in range(40):
1538
+ for proj in ["q_proj", "k_proj", "v_proj", "o_proj"]:
1539
+ target_modules.append(f"model.layers.{layer_idx}.self_attn.{proj}")
1540
+
1541
+ # Hot expert FFN layers only
1542
+ for layer_idx_str, expert_indices in hot_experts.items():
1543
+ layer_idx = int(layer_idx_str)
1544
+ for expert_idx in expert_indices:
1545
+ for proj in ["gate_proj", "up_proj", "down_proj"]:
1546
+ target_modules.append(
1547
+ f"model.layers.{layer_idx}.mlp.experts.{expert_idx}.{proj}"
1548
+ )
1549
+
1550
+ return target_modules
1551
+
1552
+
1553
+ def build_foundation_dataset(tokenizer):
1554
+ """Build the mixed foundation dataset with appropriate sampling.
1555
+
1556
+ RED-HAT FIX #3: Returns (train_dataset, eval_dataset) — stratified eval split
1557
+ of ~300 examples held out before interleaving. (Per C3 review)
1558
+ """
1559
+ # Load each dataset source
1560
+ # NOTE: Adapt these loaders to match your actual data format and paths
1561
+
1562
+ # RED-HAT FIX #3: Hold out eval examples BEFORE interleaving.
1563
+ # ~50 from each major source, fewer from small sources, fixed seed.
1564
+ EVAL_COUNTS = {
1565
+ "sonnet": 80, "prosocial": 50, "xlam": 50,
1566
+ "hermes": 30, "code": 30, "truthful": 60
1567
+ }
1568
+
1569
+ datasets_with_weights = []
1570
+ eval_datasets = []
1571
+
1572
+ # 1. Sonnet voice distill (226K, weight 0.44)
1573
+ sonnet_ds = load_dataset("json", data_files="/workspace/data/sonnet_voice_distill/*.jsonl",
1574
+ split="train").shuffle(seed=42)
1575
+ eval_datasets.append(sonnet_ds.select(range(EVAL_COUNTS["sonnet"])))
1576
+ sonnet_ds = sonnet_ds.select(range(EVAL_COUNTS["sonnet"], len(sonnet_ds)))
1577
+ datasets_with_weights.append((sonnet_ds, 0.44))
1578
+
1579
+ # 2. Prosocial dialogue (165K subsampled to 49.5K, weight 0.10)
1580
+ prosocial_ds = load_dataset("json", data_files="/workspace/data/prosocial_dialogue/*.jsonl",
1581
+ split="train").shuffle(seed=42)
1582
+ eval_datasets.append(prosocial_ds.select(range(EVAL_COUNTS["prosocial"])))
1583
+ prosocial_ds = prosocial_ds.select(range(EVAL_COUNTS["prosocial"], len(prosocial_ds)))
1584
+ prosocial_ds = prosocial_ds.select(range(min(49500, len(prosocial_ds))))
1585
+ datasets_with_weights.append((prosocial_ds, 0.10))
1586
+
1587
+ # 3. xlam function calling (60K, weight 0.12)
1588
+ xlam_ds = load_dataset("json", data_files="/workspace/data/xlam_function_calling/*.jsonl",
1589
+ split="train").shuffle(seed=42)
1590
+ eval_datasets.append(xlam_ds.select(range(EVAL_COUNTS["xlam"])))
1591
+ xlam_ds = xlam_ds.select(range(EVAL_COUNTS["xlam"], len(xlam_ds)))
1592
+ datasets_with_weights.append((xlam_ds, 0.12))
1593
+
1594
+ # 4. Hermes agent traces (7.6K upsampled 3x, weight 0.04)
1595
+ hermes_ds = load_dataset("json", data_files="/workspace/data/hermes_agent_traces/*.jsonl",
1596
+ split="train").shuffle(seed=42)
1597
+ eval_datasets.append(hermes_ds.select(range(EVAL_COUNTS["hermes"])))
1598
+ hermes_ds = hermes_ds.select(range(EVAL_COUNTS["hermes"], len(hermes_ds)))
1599
+ hermes_upsampled = Dataset.from_dict({
1600
+ k: hermes_ds[k] * 3 for k in hermes_ds.column_names
1601
+ })
1602
+ datasets_with_weights.append((hermes_upsampled, 0.04))
1603
+
1604
+ # 5. Code review feedback (9.5K upsampled 2x, weight 0.04)
1605
+ code_ds = load_dataset("json", data_files="/workspace/data/code_review_feedback/*.jsonl",
1606
+ split="train").shuffle(seed=42)
1607
+ eval_datasets.append(code_ds.select(range(EVAL_COUNTS["code"])))
1608
+ code_ds = code_ds.select(range(EVAL_COUNTS["code"], len(code_ds)))
1609
+ code_upsampled = Dataset.from_dict({
1610
+ k: code_ds[k] * 2 for k in code_ds.column_names
1611
+ })
1612
+ datasets_with_weights.append((code_upsampled, 0.04))
1613
+
1614
+ # 6. TruthfulQA (817 upsampled 15x, weight 0.02)
1615
+ truthful_ds = load_dataset("json", data_files="/workspace/data/truthfulqa/*.jsonl",
1616
+ split="train").shuffle(seed=42)
1617
+ eval_datasets.append(truthful_ds.select(range(EVAL_COUNTS["truthful"])))
1618
+ truthful_ds = truthful_ds.select(range(EVAL_COUNTS["truthful"], len(truthful_ds)))
1619
+ truthful_upsampled = Dataset.from_dict({
1620
+ k: truthful_ds[k] * 15 for k in truthful_ds.column_names
1621
+ })
1622
+ datasets_with_weights.append((truthful_upsampled, 0.02))
1623
+
1624
+ # Interleave training data with weights
1625
+ all_datasets = [d for d, w in datasets_with_weights]
1626
+ all_weights = [w for d, w in datasets_with_weights]
1627
+
1628
+ # Normalize weights
1629
+ total_weight = sum(all_weights)
1630
+ probabilities = [w / total_weight for w in all_weights]
1631
+
1632
+ mixed_dataset = interleave_datasets(
1633
+ all_datasets,
1634
+ probabilities=probabilities,
1635
+ seed=42,
1636
+ stopping_strategy="all_exhausted",
1637
+ )
1638
+
1639
+ # RED-HAT FIX #3: Combine eval splits into a single eval dataset
1640
+ from datasets import concatenate_datasets
1641
+ eval_dataset = concatenate_datasets(eval_datasets).shuffle(seed=42)
1642
+
1643
+ print(f"Foundation train dataset: {len(mixed_dataset)} examples")
1644
+ print(f"Foundation eval dataset: {len(eval_dataset)} examples")
1645
+ return mixed_dataset, eval_dataset
1646
+
1647
+
1648
+ def train_foundation(model_path, hot_experts_path):
1649
+ """Phase 2: Foundation training."""
1650
+ print("=== Phase 2: Foundation Training ===")
1651
+
1652
+ # Load model
1653
+ tokenizer = AutoTokenizer.from_pretrained(model_path)
1654
+ model = AutoModelForCausalLM.from_pretrained(
1655
+ model_path,
1656
+ torch_dtype=torch.bfloat16,
1657
+ device_map="cuda",
1658
+ )
1659
+
1660
+ # Apply HELLoRA
1661
+ target_modules = build_target_modules(hot_experts_path)
1662
+ print(f"Target modules: {len(target_modules)}")
1663
+
1664
+ lora_config = LoraConfig(
1665
+ r=32,
1666
+ lora_alpha=64,
1667
+ lora_dropout=0.05,
1668
+ target_modules=target_modules,
1669
+ use_dora=True,
1670
+ bias="none",
1671
+ task_type="CAUSAL_LM",
1672
+ )
1673
+
1674
+ model = get_peft_model(model, lora_config)
1675
+ model.print_trainable_parameters()
1676
+
1677
+ # Verify expert LoRA adapters exist
1678
+ expert_lora_count = sum(
1679
+ 1 for name, p in model.named_parameters()
1680
+ if "experts" in name and "lora" in name and p.requires_grad
1681
+ )
1682
+ assert expert_lora_count > 0, "FATAL: No LoRA adapters on expert layers!"
1683
+ print(f"Expert LoRA parameters: {expert_lora_count}")
1684
+
1685
+ # Build dataset
1686
+ # RED-HAT FIX #3: build_foundation_dataset now returns (train_dataset, eval_dataset)
1687
+ # Eval split: ~300 examples stratified across all 6 sources, fixed seed, excluded from training.
1688
+ train_dataset, eval_dataset = build_foundation_dataset(tokenizer)
1689
+
1690
+ # Training config
1691
+ training_args = SFTConfig(
1692
+ output_dir="/workspace/checkpoints/daimon_v2_foundation",
1693
+ per_device_train_batch_size=4,
1694
+ gradient_accumulation_steps=4,
1695
+ num_train_epochs=1,
1696
+ learning_rate=2e-4,
1697
+ lr_scheduler_type="cosine",
1698
+ warmup_ratio=0.03,
1699
+ weight_decay=0.01,
1700
+ bf16=True,
1701
+ max_seq_length=2048,
1702
+ logging_steps=50,
1703
+ save_steps=2000,
1704
+ save_total_limit=5,
1705
+ dataloader_num_workers=0,
1706
+ dataset_num_proc=1,
1707
+ gradient_checkpointing=True,
1708
+ gradient_checkpointing_kwargs={"use_reentrant": False},
1709
+ eval_strategy="steps",
1710
+ eval_steps=2000,
1711
+ load_best_model_at_end=True, # RED-HAT FIX #3: Select best checkpoint by eval_loss
1712
+ metric_for_best_model="eval_loss", # RED-HAT FIX #3
1713
+ max_steps=30000, # RED-HAT FIX #4: Cost circuit breaker (adjusted by throughput probe)
1714
+ report_to="none",
1715
+ seed=42,
1716
+ )
1717
+
1718
+ # Train
1719
+ trainer = SFTTrainer(
1720
+ model=model,
1721
+ args=training_args,
1722
+ train_dataset=train_dataset, # RED-HAT FIX #3: Was `dataset`, now split
1723
+ eval_dataset=eval_dataset, # RED-HAT FIX #3: Added — prevents eval crash
1724
+ processing_class=tokenizer,
1725
+ )
1726
+
1727
+ trainer.train()
1728
+
1729
+ # Save best checkpoint
1730
+ model.save_pretrained("/workspace/checkpoints/daimon_v2_foundation_best")
1731
+ tokenizer.save_pretrained("/workspace/checkpoints/daimon_v2_foundation_best")
1732
+ print("=== Foundation training complete ===")
1733
+
1734
+
1735
+ def train_calibration(checkpoint_path, base_model_path="/workspace/models/Qwen3.6-35B-A3B"):
1736
+ """Phase 3: Thomas voice calibration.
1737
+
1738
+ RED-HAT FIX #2: Phase 2 saves adapter-only (adapter_config.json + adapter_model.safetensors),
1739
+ NOT a full model. Must load base model first, then attach adapter with is_trainable=True.
1740
+ Original code used AutoModelForCausalLM.from_pretrained on the adapter dir, which either
1741
+ crashes ("no config.json") or loads adapter for inference-only (requires_grad=False),
1742
+ causing silent no-op training — the exact mlx-lm #571 failure class. (Per C2 review)
1743
+ """
1744
+ print("=== Phase 3: Thomas Voice Calibration ===")
1745
+
1746
+ # RED-HAT FIX #2: Load BASE model first, then attach Phase 2 adapter
1747
+ tokenizer = AutoTokenizer.from_pretrained(base_model_path)
1748
+ base_model = AutoModelForCausalLM.from_pretrained(
1749
+ base_model_path, # Base model path, NOT the checkpoint
1750
+ torch_dtype=torch.bfloat16,
1751
+ device_map="cuda",
1752
+ )
1753
+
1754
+ # Attach Phase 2 adapter with is_trainable=True
1755
+ model = PeftModel.from_pretrained(
1756
+ base_model,
1757
+ checkpoint_path, # Phase 2 adapter checkpoint
1758
+ is_trainable=True, # CRITICAL: without this, all adapter params are frozen
1759
+ )
1760
+
1761
+ # RED-HAT FIX #2: Verify adapter is actually trainable (catches silent no-op)
1762
+ trainable_count = sum(p.numel() for p in model.parameters() if p.requires_grad)
1763
+ assert trainable_count > 0, "FATAL: No trainable parameters in Phase 3! Check is_trainable=True."
1764
+ expert_lora_count = sum(
1765
+ 1 for name, p in model.named_parameters()
1766
+ if "experts" in name and "lora" in name and p.requires_grad
1767
+ )
1768
+ assert expert_lora_count > 0, "FATAL: No LoRA adapters on expert layers in Phase 3!"
1769
+ print(f"Phase 3 trainable params: {trainable_count:,}, expert LoRA params: {expert_lora_count}")
1770
+
1771
+ # RED-HAT FIX #2: Weight-delta assert — verify training actually changes weights
1772
+ import hashlib
1773
+ lora_param = next((name, p) for name, p in model.named_parameters()
1774
+ if "lora" in name and p.requires_grad)
1775
+ pre_hash = hashlib.sha256(lora_param[1].data.cpu().numpy().tobytes()).hexdigest()
1776
+
1777
+ # Load Thomas voice demonstrations
1778
+ demos = []
1779
+ with open("/workspace/data/thomas_voice_demos/daimon_voice_demonstrations.jsonl") as f:
1780
+ for line in f:
1781
+ demos.append(json.loads(line))
1782
+ dataset = Dataset.from_list(demos)
1783
+ print(f"Thomas voice demos: {len(dataset)} examples")
1784
+
1785
+ training_args = SFTConfig(
1786
+ output_dir="/workspace/checkpoints/daimon_v2_calibration",
1787
+ per_device_train_batch_size=2,
1788
+ gradient_accumulation_steps=1,
1789
+ num_train_epochs=5,
1790
+ learning_rate=1e-5,
1791
+ lr_scheduler_type="cosine",
1792
+ warmup_ratio=0.1,
1793
+ weight_decay=0.01,
1794
+ bf16=True,
1795
+ max_seq_length=2048,
1796
+ logging_steps=5,
1797
+ save_steps=20,
1798
+ save_total_limit=10,
1799
+ dataloader_num_workers=0,
1800
+ dataset_num_proc=1,
1801
+ gradient_checkpointing=True,
1802
+ gradient_checkpointing_kwargs={"use_reentrant": False},
1803
+ seed=42,
1804
+ )
1805
+
1806
+ trainer = SFTTrainer(
1807
+ model=model,
1808
+ args=training_args,
1809
+ train_dataset=dataset,
1810
+ processing_class=tokenizer,
1811
+ )
1812
+
1813
+ trainer.train()
1814
+
1815
+ # RED-HAT FIX #2: Verify weight actually changed after training
1816
+ post_hash = hashlib.sha256(lora_param[1].data.cpu().numpy().tobytes()).hexdigest()
1817
+ assert pre_hash != post_hash, (
1818
+ f"FATAL: LoRA weight {lora_param[0]} unchanged after 5 epochs of training! "
1819
+ "Phase 3 was a no-op. Check adapter loading."
1820
+ )
1821
+ print(f"Weight-delta verified: {lora_param[0]} changed")
1822
+
1823
+ model.save_pretrained("/workspace/checkpoints/daimon_v2_calibration_best")
1824
+ tokenizer.save_pretrained("/workspace/checkpoints/daimon_v2_calibration_best")
1825
+ print("=== Thomas voice calibration complete ===")
1826
+
1827
+
1828
+ if __name__ == "__main__":
1829
+ parser = argparse.ArgumentParser()
1830
+ parser.add_argument("--phase", choices=["foundation", "calibration"], required=True)
1831
+ parser.add_argument("--model", default="/workspace/models/Qwen3.6-35B-A3B")
1832
+ parser.add_argument("--checkpoint", default=None, help="Checkpoint for calibration phase")
1833
+ parser.add_argument("--hot-experts", default="/workspace/artifacts/hot_experts.json")
1834
+ args = parser.parse_args()
1835
+
1836
+ if args.phase == "foundation":
1837
+ train_foundation(args.model, args.hot_experts)
1838
+ elif args.phase == "calibration":
1839
+ cp = args.checkpoint or "/workspace/checkpoints/daimon_v2_foundation_best"
1840
+ train_calibration(cp)
1841
+ ```
1842
+
1843
+ **Critical implementation notes:**
1844
+
1845
+ 1. **Unsloth integration:** The script above uses raw PEFT + Transformers for clarity. In production, you may want to use Unsloth's `FastLanguageModel` for the 12x MoE kernel speedup. If using Unsloth, verify that `FastLanguageModel.get_peft_model()` accepts explicit module name lists (not just module type strings). If it does not, use the PEFT-direct approach shown here and accept the throughput penalty, or patch Unsloth's module targeting.
1846
+
1847
+ 2. **Data format compatibility:** The `SFTTrainer` expects data in a specific format depending on the `dataset_text_field` or `formatting_func` parameter. If your data uses `{"messages": [...]}` format, you may need to set `dataset_text_field=None` and let TRL auto-detect the chat format, or provide a custom `formatting_func` that applies the Qwen3.6 chat template.
1848
+
1849
+ 3. **Gradient checkpointing:** `use_reentrant=False` is required for compatibility with PEFT LoRA on MoE. The reentrant variant can cause incorrect gradients when LoRA interacts with MoE routing.
1850
+
1851
+ 4. **Resuming from interruption:** If training is interrupted (spot instance preemption, SIGKILL, etc.), resume from the last checkpoint:
1852
+ ```bash
1853
+ python daimon_train.py --phase foundation --checkpoint /workspace/checkpoints/daimon_v2_foundation/checkpoint-XXXX
1854
+ ```
1855
+ Add `resume_from_checkpoint=True` to the SFTConfig, or pass the checkpoint path to `trainer.train(resume_from_checkpoint=...)`.
1856
+
1857
+ ---
1858
+
1859
+ ## Appendix D: LASER Post-Training Script (OPTIONAL — Run Locally on Margaret) {#appendix-d}
1860
+
1861
+ <!-- RED-HAT FIX #1: LASER moved to optional offline experiment per C1 review.
1862
+ Original code was INVERTED — zeroed the LARGEST singular values (principal components)
1863
+ instead of the tail. This would have destroyed the trained model's quality.
1864
+
1865
+ Additionally, blanket application to all 20,480 matrices contradicts LASER's
1866
+ LAyer-SElective protocol. The corrected script below:
1867
+ 1. Fixes SVD direction: keeps top singular values, zeros the tail
1868
+ 2. Applies per-(layer, matrix) with eval between each application
1869
+ 3. Designed for local execution on Margaret ($0), not the cloud pod -->
1870
+
1871
+ ```python
1872
+ #!/usr/bin/env python3
1873
+ """
1874
+ daimon_laser.py
1875
+ OPTIONAL post-budget experiment: LASER (Layer-Selective Rank Reduction) on merged model.
1876
+ Run LOCALLY on Margaret, not on cloud pod.
1877
+
1878
+ Reference: arXiv:2312.13558
1879
+
1880
+ RED-HAT FIX #1 (C1 review): SVD direction corrected.
1881
+ Original code zeroed S[:n_remove] which removes the LARGEST singular values
1882
+ (principal components). torch.linalg.svd returns values in DESCENDING order.
1883
+ Corrected: S[n_keep:] = 0 keeps the top (largest) and zeros the tail (smallest = noise).
1884
+ """
1885
+
1886
+ import os
1887
+ os.environ["UNSLOTH_COMPILE_DISABLE"] = "1"
1888
+ os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "expandable_segments:True"
1889
+
1890
+ import torch
1891
+ import argparse
1892
+ from transformers import AutoModelForCausalLM, AutoTokenizer
1893
+
1894
+ NUM_LAYERS = 40
1895
+ NUM_EXPERTS = 256
1896
+
1897
+
1898
+ def laser_reduce(weight_matrix, keep_fraction=0.95):
1899
+ """
1900
+ Keep top fraction of singular values, zero out the tail (noise).
1901
+
1902
+ LASER's key insight: the SMALL singular values of MLP weight matrices
1903
+ often encode noise or spurious correlations. Removing them (keeping
1904
+ only the top components) can improve factual recall on specific tasks.
1905
+
1906
+ IMPORTANT: torch.linalg.svd returns singular values in DESCENDING order.
1907
+ S[0] is the LARGEST. We keep S[:n_keep] (largest) and zero S[n_keep:] (smallest).
1908
+
1909
+ The ORIGINAL code in this spec had this INVERTED (S[:n_remove] = 0), which
1910
+ would have destroyed the principal components of every weight matrix.
1911
+
1912
+ Args:
1913
+ weight_matrix: 2D tensor [out_features, in_features]
1914
+ keep_fraction: fraction of top singular values to KEEP (default: 0.95 = drop bottom 5%)
1915
+ Returns:
1916
+ Modified weight matrix with tail singular values removed
1917
+ """
1918
+ original_dtype = weight_matrix.dtype
1919
+ W = weight_matrix.float() # SVD requires float32
1920
+
1921
+ U, S, Vh = torch.linalg.svd(W, full_matrices=False)
1922
+ n_keep = max(1, int(len(S) * keep_fraction))
1923
+
1924
+ S_modified = S.clone()
1925
+ S_modified[n_keep:] = 0.0 # Zero out TAIL (small values = noise), KEEP TOP
1926
+
1927
+ W_modified = U @ torch.diag(S_modified) @ Vh
1928
+ return W_modified.to(original_dtype)
1929
+
1930
+
1931
+ def apply_laser_selective(model_path, output_path, keep_fraction=0.95,
1932
+ target_layer=None, target_proj="gate_proj"):
1933
+ """
1934
+ Apply LASER to a SINGLE (layer, projection) pair.
1935
+
1936
+ DO NOT apply blanket to all 20,480 matrices — that contradicts LASER's
1937
+ LAyer-SElective protocol. Apply one at a time, eval after each.
1938
+ """
1939
+ print(f"=== LASER: keep_fraction={keep_fraction}, layer={target_layer}, proj={target_proj} ===")
1940
+
1941
+ tokenizer = AutoTokenizer.from_pretrained(model_path)
1942
+ model = AutoModelForCausalLM.from_pretrained(
1943
+ model_path,
1944
+ torch_dtype=torch.bfloat16,
1945
+ device_map="cpu", # LASER on CPU — no GPU needed
1946
+ )
1947
+
1948
+ if target_layer is not None:
1949
+ layers_to_process = [target_layer]
1950
+ else:
1951
+ layers_to_process = range(NUM_LAYERS)
1952
+
1953
+ total_modified = 0
1954
+ for layer_idx in layers_to_process:
1955
+ moe_block = model.model.layers[layer_idx].mlp
1956
+
1957
+ for expert_idx in range(NUM_EXPERTS):
1958
+ expert = moe_block.experts[expert_idx]
1959
+ weight = getattr(expert, target_proj).weight.data
1960
+ new_weight = laser_reduce(weight, keep_fraction)
1961
+ getattr(expert, target_proj).weight.data = new_weight
1962
+ total_modified += 1
1963
+
1964
+ print(f" Layer {layer_idx}: modified {NUM_EXPERTS} experts ({target_proj})")
1965
+
1966
+ print(f"Modified {total_modified} weight matrices")
1967
+
1968
+ # Save
1969
+ model.save_pretrained(output_path, safe_serialization=True)
1970
+ tokenizer.save_pretrained(output_path)
1971
+ print(f"Saved to {output_path}")
1972
+ print("NOW RUN EVAL to verify this improved quality. If not, revert.")
1973
+
1974
+
1975
+ if __name__ == "__main__":
1976
+ parser = argparse.ArgumentParser()
1977
+ parser.add_argument("--model", required=True, help="Path to merged model")
1978
+ parser.add_argument("--output", required=True, help="Output path")
1979
+ parser.add_argument("--keep-fraction", type=float, default=0.95,
1980
+ help="Fraction of top singular values to KEEP (default: 0.95 = drop bottom 5%%)")
1981
+ parser.add_argument("--layer", type=int, default=None,
1982
+ help="Specific layer to target (default: all layers — NOT recommended)")
1983
+ parser.add_argument("--proj", default="gate_proj", choices=["gate_proj", "up_proj"],
1984
+ help="Projection to target (default: gate_proj)")
1985
+ args = parser.parse_args()
1986
+
1987
+ apply_laser_selective(args.model, args.output, args.keep_fraction, args.layer, args.proj)
1988
+ ```
1989
+
1990
+ ---
1991
+
1992
+ ## Reference: Key Papers Cited
1993
+
1994
+ | Paper | arXiv | Used For |
1995
+ |---|---|---|
1996
+ | ESFT (DeepSeek) | 2407.01906 | Expert profiling methodology, shared expert freeze evidence |
1997
+ | HELLoRA | 2605.18795 | Primary training method |
1998
+ | Unified Expert Scoring | 2606.15716 | MAN scoring > frequency scoring for expert importance |
1999
+ | DoRA | 2402.09353 | LoRA upgrade (+1-4% accuracy) |
2000
+ | LASER | 2312.13558 | Post-training rank reduction |
2001
+ | EAQuant | 2506.13329 | QLoRA breaks MoE routing (do not use) |
2002
+ | Geometric coupling (routers) | 2605.12476 | Do not fine-tune routers |
2003
+ | MoE-Sieve | 2603.24044 | Routing-guided LoRA validation |
2004
+ | DR-LoRA | 2601.04823 | Dynamic rank allocation (future iteration) |
2005
+ | OGPSA v3.1 | ogpsa_persona_validation_v3.1.md | Personality geometry capture |
2006
+ | Training best practices | training_pipeline_best_practices.md | alpha=2r, checkpoint averaging, OPLoRA |
2007
+
2008
+ ---
2009
+
2010
+ ## Changelog
2011
+
2012
+ - **v2.1** (July 2026): Applied 12-point red-hat review fix (daimon_pipeline_v2_redhat.md). Fixes: (1) LASER SVD direction inverted -- moved to optional offline experiment; (2) Phase 3 checkpoint loading via PeftModel.from_pretrained with is_trainable=True + weight-delta assert; (3) Added stratified eval split to fix eval_strategy crash; (4) Replaced Unsloth throughput estimates with realistic raw-PEFT speeds, added 50-step throughput probe gate and max_steps circuit breaker; (5) Disk preflight gate increased from 150 GB to 400 GB; (6) VRAM estimates corrected from 55-65 GB to 90-115 GB peak; (7) TruthfulQA removed from eval battery (contaminated -- in training data 15x); (8) Added checkpoint egress rsync after every save; (9) Pinned dependency versions, torch installed first; (10) Added throughput probe as hard budget gate; (11) Added backward-pass smoke test for gradient flow verification; (12) Budget table corrected to $30-45 primary / $52-67 reserve.
2013
+ - **v2.0** (July 2026): Complete rewrite for H200 cloud training. HELLoRA method selection. Budget-constrained design. SOTA sweep integration.
2014
+ - **v1.0** (June 2026): Original spec targeting Apple Silicon (Margaret). 8-stage pipeline with DPO and RL. See `daimon_training_spec.md`.