Add corrected Daimon training pipeline v2.1 specification
Browse files
training-template/daimon_training_pipeline_v2.md
ADDED
|
@@ -0,0 +1,2014 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Daimon Training Pipeline v2.0
|
| 2 |
+
# Liberation Labs House Model -- Implementation-Ready Specification
|
| 3 |
+
# July 2026
|
| 4 |
+
|
| 5 |
+
---
|
| 6 |
+
|
| 7 |
+
## Executive Summary
|
| 8 |
+
|
| 9 |
+
This document specifies the complete training pipeline for Daimon, Liberation Labs' prosocial house model -- Thomas Edrington's professional voice rendered as a sovereign model on Qwen3.6-35B-A3B.
|
| 10 |
+
|
| 11 |
+
**Method:** HELLoRA (Hot Expert Layer-Level LoRA) with MAN-scored expert profiling, two-stage curriculum, OGPSA personality capture, and optional LASER post-training (offline). <!-- RED-HAT FIX #1: LASER moved to optional per C1 review -->
|
| 12 |
+
|
| 13 |
+
**Hardware:** Single H200 SXM (141 GB HBM3e, 276 GB container RAM).
|
| 14 |
+
|
| 15 |
+
**Budget:** $97 total at ~$2/hr (NeevCloud or PrimeIntellect). Estimated ~$30-45 for primary run (realistic raw-PEFT throughput), ~$52-67 reserved for iteration. <!-- RED-HAT FIX #12: Corrected budget from $28-36/$61-69 to realistic PEFT speeds per C4 review -->
|
| 16 |
+
|
| 17 |
+
**Why not the v1 spec:** The v1 pipeline (daimon_training_spec.md) assumed Apple Silicon on Margaret with mlx-tune. This pipeline targets cloud H200 for the initial training run, then exports to GGUF/GPTQ for local serving. The v1 also specified 8 stages with DPO and RL -- this pipeline consolidates into fewer stages to reduce inter-stage forgetting risk and fit the $97 budget.
|
| 18 |
+
|
| 19 |
+
---
|
| 20 |
+
|
| 21 |
+
## TABLE OF CONTENTS
|
| 22 |
+
|
| 23 |
+
1. [Critical Constraints (Non-Negotiable)](#1-critical-constraints)
|
| 24 |
+
2. [Method Selection Rationale](#2-method-selection-rationale)
|
| 25 |
+
3. [Phase 0: Preflight Checklist](#3-phase-0-preflight)
|
| 26 |
+
4. [Phase 1: Expert Profiling](#4-phase-1-expert-profiling)
|
| 27 |
+
5. [Phase 2: Foundation Training](#5-phase-2-foundation-training)
|
| 28 |
+
6. [Phase 3: Thomas Voice Calibration](#6-phase-3-thomas-voice-calibration)
|
| 29 |
+
7. [Phase 4: Post-Training](#7-phase-4-post-training)
|
| 30 |
+
8. [Phase 5: Export and Validation](#8-phase-5-export-and-validation)
|
| 31 |
+
9. [Budget Allocation](#9-budget-allocation)
|
| 32 |
+
10. [Rollback Plan](#10-rollback-plan)
|
| 33 |
+
11. [Appendix A: Environment Setup Script](#appendix-a)
|
| 34 |
+
12. [Appendix B: Expert Profiling Script](#appendix-b)
|
| 35 |
+
13. [Appendix C: Training Script](#appendix-c)
|
| 36 |
+
14. [Appendix D: LASER Post-Training Script](#appendix-d)
|
| 37 |
+
|
| 38 |
+
---
|
| 39 |
+
|
| 40 |
+
## 1. Critical Constraints (Non-Negotiable) {#1-critical-constraints}
|
| 41 |
+
|
| 42 |
+
These constraints were discovered during the SOTA sweep (sota_training_sweep_202607.md) and the RunPod OOM debug session. Violating any of these will waste budget.
|
| 43 |
+
|
| 44 |
+
### Hard Stops
|
| 45 |
+
|
| 46 |
+
| Constraint | Why | Source |
|
| 47 |
+
|---|---|---|
|
| 48 |
+
| **No QLoRA / 4-bit quantization** | BitsAndBytes 4-bit introduces routing perturbations on Qwen3.5/3.6 MoE. Causes token misrouting and quality degradation. Unsloth docs explicitly warn against it. | Unsloth docs, EAQuant (arXiv:2506.13329) |
|
| 49 |
+
| **No router fine-tuning** | Routers co-evolve with expert geometry during pretraining. Modifying routers without corresponding expert adjustment causes performance degradation. | Unsloth, ESFT, HELLoRA, DR-LoRA, arXiv:2605.12476 |
|
| 50 |
+
| **No DeepSpeed ZeRO-3** | Breaks gradient flow when LoRA adapters interact with MoE routing. Single GPU does not need DeepSpeed anyway. | Community reports, ZeRO docs |
|
| 51 |
+
| **Freeze shared experts** | Training shared params causes overfitting and catastrophic forgetting. ESFT ablation is definitive. 1 shared expert per layer on Qwen3.6 -- freeze all 40. | ESFT (arXiv:2407.01906) ablation study |
|
| 52 |
+
| **Freeze embeddings and LM head** | Unless specifically training for new vocabulary. Default Unsloth behavior. | Standard practice |
|
| 53 |
+
| **Container RAM is 276 GB, not 2 TB** | `free -h` shows host RAM on RunPod. Actual limit is cgroup-enforced. Check with `cat /sys/fs/cgroup/memory/memory.limit_in_bytes` (v1) or `cat /sys/fs/cgroup/memory.max` (v2). Exceeding this triggers SIGKILL with no error message. | runpod_oom_debug.md |
|
| 54 |
+
|
| 55 |
+
### Required Environment Variables
|
| 56 |
+
|
| 57 |
+
These MUST be set BEFORE any Python imports:
|
| 58 |
+
|
| 59 |
+
```bash
|
| 60 |
+
export UNSLOTH_COMPILE_DISABLE=1
|
| 61 |
+
export PYTORCH_CUDA_ALLOC_CONF="expandable_segments:True"
|
| 62 |
+
export TOKENIZERS_PARALLELISM=false
|
| 63 |
+
export UNSLOTH_DISABLE_FAST_GENERATION=1 # fixes RuntimeError from torch.compile on MoE
|
| 64 |
+
```
|
| 65 |
+
|
| 66 |
+
### Required Software Versions
|
| 67 |
+
|
| 68 |
+
<!-- RED-HAT FIX #9: Pinned exact versions per H5 review. Version floors (>=) resolve to
|
| 69 |
+
whatever shipped that morning; pinned versions ensure reproducibility. Install torch
|
| 70 |
+
FIRST, unsloth --no-deps LAST. See Appendix A for install order. -->
|
| 71 |
+
|
| 72 |
+
| Package | Version (pinned) | Why |
|
| 73 |
+
|---|---|---|
|
| 74 |
+
| torch | 2.7.1+cu124 | bf16 MoE training support; install FIRST to anchor resolution |
|
| 75 |
+
| transformers | 5.0.2 | Qwen3.5/3.6 architecture support |
|
| 76 |
+
| peft | 0.15.2 | Module-name-level target_modules for selective expert LoRA |
|
| 77 |
+
| trl | 0.18.1 | SFTTrainer/SFTConfig API stability |
|
| 78 |
+
| datasets | 3.6.0 | Streaming, interleaving |
|
| 79 |
+
| accelerate | 1.7.0 | Training orchestration |
|
| 80 |
+
| unsloth | >= 0.1.47-beta | MoE-specific Triton kernels, split LoRA; install LAST with --no-deps |
|
| 81 |
+
|
| 82 |
+
### Training Hyperparameter Constraints
|
| 83 |
+
|
| 84 |
+
| Parameter | Value | Why |
|
| 85 |
+
|---|---|---|
|
| 86 |
+
| `dataloader_num_workers` | 0 | Prevents MoE deadlocks with multi-process data loading |
|
| 87 |
+
| `dataset_num_proc` | 1 | Same deadlock prevention |
|
| 88 |
+
| Auxiliary loss coefficient | 0.01 | Mitigates expert collapse during fine-tuning |
|
| 89 |
+
|
| 90 |
+
---
|
| 91 |
+
|
| 92 |
+
## 2. Method Selection Rationale {#2-method-selection-rationale}
|
| 93 |
+
|
| 94 |
+
### Why HELLoRA over alternatives
|
| 95 |
+
|
| 96 |
+
| Method | Quality | VRAM (est.) | Throughput | Code Available | Risk |
|
| 97 |
+
|---|---|---|---|---|---|
|
| 98 |
+
| **Standard LoRA (all experts)** | Baseline | ~74 GB | 1x | Unsloth | Low |
|
| 99 |
+
| **ESFT (full-param hot experts)** | 98.4% of full SFT | ~70-75 GB | ~0.7x | DeepSeek GitHub | Medium -- needs adaptation for Qwen3.6 |
|
| 100 |
+
| **HELLoRA (LoRA on hot experts)** | > standard LoRA | ~90-115 GB peak | 1.9x (Unsloth) / ~0.5-0.7x (raw PEFT) | Manual implementation | Low | <!-- RED-HAT FIX #6: VRAM corrected; throughput split by code path -->
|
| 101 |
+
| **BAdam (block-coordinate)** | 98.6% of full SFT | ~95-110 GB | ~0.5x | GitHub | Medium -- longer wall-clock |
|
| 102 |
+
| **LISA (random layer unfreezing)** | > LoRA on MT-Bench | ~85-95 GB | ~0.8x | GitHub | Medium -- untested on MoE |
|
| 103 |
+
|
| 104 |
+
**HELLoRA wins on the combination that matters for this budget:**
|
| 105 |
+
|
| 106 |
+
1. **Memory efficiency:** ~90-115 GB peak (model bf16 70-72 GB + adapters/optimizer 3-8 GB + activations 8-15 GB + CE logits ~5 GB + CUDA overhead 3-6 GB) leaves ~25-50 GB headroom on H200. batch_size=4 at seq_len=2048 fits safely. <!-- RED-HAT FIX #6: Corrected VRAM from 55 GB to 90-115 GB per H1 review; removed seq_len=4096 claim which contradicts OOM debug findings -->
|
| 107 |
+
2. **Training throughput:** 1.9x over standard LoRA (arXiv:2605.18795). Directly translates to more training per dollar.
|
| 108 |
+
3. **Quality:** Outperforms standard LoRA on benchmarks (OLMoE GSM8K: 29.49 vs 26.37; Mixtral HumanEval: 44.89 vs 39.99).
|
| 109 |
+
4. **Implementation simplicity:** HELLoRA is standard LoRA with selective module targeting. No custom training loops, no model surgery. Works with Unsloth/PEFT by constructing the right `target_modules` list.
|
| 110 |
+
5. **Validated on 47B MoE:** Tested on Mixtral-8x7B (47B total), which is the closest published result to Qwen3.6-35B-A3B.
|
| 111 |
+
6. **Rollback-friendly:** LoRA checkpoints are ~500 MB-1 GB. Can checkpoint every 2000 steps and roll back to any point without re-spending.
|
| 112 |
+
|
| 113 |
+
**Why not ESFT:** ESFT does full parameter fine-tuning of hot experts, which means full optimizer states (Adam m + v) for those experts. At 10% of 35B params, that is ~3.5B trainable params requiring ~56 GB of optimizer state alone. HELLoRA with r=32 on the same hot experts requires <1 GB of optimizer state. ESFT has higher quality ceiling but the throughput penalty (no Unsloth MoE kernel acceleration for full-param training) means fewer iterations within budget.
|
| 114 |
+
|
| 115 |
+
**Why not BAdam:** Near-full-SFT quality (98.6%) but at ~95-110 GB VRAM and ~0.5x throughput. Would consume the full budget in a single training run with no room for iteration.
|
| 116 |
+
|
| 117 |
+
### LoRA Configuration
|
| 118 |
+
|
| 119 |
+
| Parameter | Value | Rationale |
|
| 120 |
+
|---|---|---|
|
| 121 |
+
| Rank (r) | 32 | Can afford higher rank because we're only targeting ~9 hot experts per layer instead of 256. HELLoRA paper uses r=32. |
|
| 122 |
+
| Alpha | 64 | Alpha = 2r per Microsoft guidance and training_pipeline_best_practices.md. Critical for high-rank convergence. |
|
| 123 |
+
| Dropout | 0.05 | Light regularization. Lower than standard 0.1 because hot-expert-only training is already parameter-efficient. |
|
| 124 |
+
| Target modules | Hot expert FFN layers + attention projections (see Phase 1 output) | HELLoRA protocol: LoRA on hot experts + attention, cold experts frozen. |
|
| 125 |
+
| DoRA | True | +1-4% accuracy over vanilla LoRA at 5-10% memory overhead (arXiv:2402.09353, ICML 2024). Drop-in replacement. |
|
| 126 |
+
|
| 127 |
+
---
|
| 128 |
+
|
| 129 |
+
## 3. Phase 0: Preflight Checklist {#3-phase-0-preflight}
|
| 130 |
+
|
| 131 |
+
**Cost:** $0 (local) or ~$1 (first 30 min on cloud if needed)
|
| 132 |
+
**Time:** 30-60 minutes
|
| 133 |
+
**Purpose:** Verify everything works BEFORE the meter starts running
|
| 134 |
+
|
| 135 |
+
### 0.1 Hardware Verification (on the cloud instance)
|
| 136 |
+
|
| 137 |
+
```bash
|
| 138 |
+
# GPU check
|
| 139 |
+
nvidia-smi --query-gpu=name,memory.total,memory.free --format=csv,noheader
|
| 140 |
+
|
| 141 |
+
# Expected: NVIDIA H200, 141287 MiB total
|
| 142 |
+
|
| 143 |
+
# Container RAM -- the REAL number (NOT free -h)
|
| 144 |
+
# Try cgroup v2 first, fall back to v1
|
| 145 |
+
CONTAINER_RAM=$(cat /sys/fs/cgroup/memory.max 2>/dev/null || cat /sys/fs/cgroup/memory/memory.limit_in_bytes 2>/dev/null)
|
| 146 |
+
echo "Container RAM limit: $(echo "$CONTAINER_RAM / 1024 / 1024 / 1024" | bc) GB"
|
| 147 |
+
|
| 148 |
+
# Expected: >= 276 GB. If less, STOP. Resize pod.
|
| 149 |
+
if [ "$CONTAINER_RAM" -lt 270000000000 ]; then
|
| 150 |
+
echo "FATAL: Container RAM < 270 GB. Abort."
|
| 151 |
+
exit 1
|
| 152 |
+
fi
|
| 153 |
+
|
| 154 |
+
# CUDA version
|
| 155 |
+
nvcc --version # Expect 12.x
|
| 156 |
+
|
| 157 |
+
# Disk space (need ~400 GB for model + checkpoints + merged weights + exports + data)
|
| 158 |
+
# RED-HAT FIX #5: Increased from 150 GB to 400 GB per C5 review.
|
| 159 |
+
# Artifact breakdown: base bf16 70 + adapter checkpoints w/ optimizer ~18 +
|
| 160 |
+
# merged bf16 70 + GGUF intermediate 70 + Q4_K_M 18 + Q5_K_M 22 + data/cache ~6 = ~346 GB peak
|
| 161 |
+
df -h /workspace
|
| 162 |
+
DISK_FREE_GB=$(df --output=avail /workspace | tail -1 | awk '{print int($1/1048576)}')
|
| 163 |
+
if [ "$DISK_FREE_GB" -lt 400 ]; then
|
| 164 |
+
echo "FATAL: Disk space < 400 GB free (have ${DISK_FREE_GB} GB). Abort."
|
| 165 |
+
exit 1
|
| 166 |
+
fi
|
| 167 |
+
```
|
| 168 |
+
|
| 169 |
+
### 0.2 Environment Setup
|
| 170 |
+
|
| 171 |
+
Run the full setup script (Appendix A). Verify:
|
| 172 |
+
|
| 173 |
+
```bash
|
| 174 |
+
# After setup, verify critical packages
|
| 175 |
+
python -c "
|
| 176 |
+
import torch; print(f'PyTorch: {torch.__version__}, CUDA: {torch.cuda.is_available()}')
|
| 177 |
+
import transformers; print(f'Transformers: {transformers.__version__}')
|
| 178 |
+
import peft; print(f'PEFT: {peft.__version__}')
|
| 179 |
+
import unsloth; print(f'Unsloth: {unsloth.__version__}')
|
| 180 |
+
assert torch.cuda.get_device_properties(0).total_memory > 140e9, 'Not H200'
|
| 181 |
+
print('All checks passed.')
|
| 182 |
+
"
|
| 183 |
+
```
|
| 184 |
+
|
| 185 |
+
### 0.3 Model Download (Do This Before Meter if Possible)
|
| 186 |
+
|
| 187 |
+
```bash
|
| 188 |
+
# Download model weights (~70 GB). If pod has persistent storage, do this once.
|
| 189 |
+
huggingface-cli download Qwen/Qwen3.6-35B-A3B --local-dir /workspace/models/Qwen3.6-35B-A3B
|
| 190 |
+
```
|
| 191 |
+
|
| 192 |
+
### 0.4 Data Upload
|
| 193 |
+
|
| 194 |
+
Upload all training datasets to `/workspace/data/`:
|
| 195 |
+
|
| 196 |
+
```
|
| 197 |
+
/workspace/data/
|
| 198 |
+
sonnet_voice_distill/ # 226K Sonnet voice (Roman1111111 + Nitral-AI + Norquinal)
|
| 199 |
+
prosocial_dialogue/ # 165K AllenAI prosocial (CC-BY-4.0)
|
| 200 |
+
xlam_function_calling/ # 60K agentic tool use
|
| 201 |
+
hermes_agent_traces/ # 7.6K reasoning traces
|
| 202 |
+
code_review_feedback/ # 9.5K CodeUltraFeedback
|
| 203 |
+
truthfulqa/ # 817 TruthfulQA
|
| 204 |
+
thomas_voice_demos/ # daimon_voice_demonstrations.jsonl (32 pairs)
|
| 205 |
+
```
|
| 206 |
+
|
| 207 |
+
### 0.5 One-Step Smoke Test
|
| 208 |
+
|
| 209 |
+
```python
|
| 210 |
+
import os
|
| 211 |
+
os.environ["UNSLOTH_COMPILE_DISABLE"] = "1"
|
| 212 |
+
os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "expandable_segments:True"
|
| 213 |
+
os.environ["TOKENIZERS_PARALLELISM"] = "false"
|
| 214 |
+
os.environ["UNSLOTH_DISABLE_FAST_GENERATION"] = "1"
|
| 215 |
+
|
| 216 |
+
import torch
|
| 217 |
+
from unsloth import FastLanguageModel
|
| 218 |
+
|
| 219 |
+
model, tokenizer = FastLanguageModel.from_pretrained(
|
| 220 |
+
"Qwen/Qwen3.6-35B-A3B",
|
| 221 |
+
max_seq_length=2048,
|
| 222 |
+
load_in_4bit=False,
|
| 223 |
+
dtype=torch.bfloat16,
|
| 224 |
+
)
|
| 225 |
+
|
| 226 |
+
# Verify model loads and generates
|
| 227 |
+
inputs = tokenizer("The daemon watches.", return_tensors="pt").to("cuda")
|
| 228 |
+
with torch.no_grad():
|
| 229 |
+
output = model.generate(**inputs, max_new_tokens=20)
|
| 230 |
+
print(tokenizer.decode(output[0]))
|
| 231 |
+
|
| 232 |
+
# Check VRAM after model load
|
| 233 |
+
allocated = torch.cuda.memory_allocated() / 1e9
|
| 234 |
+
print(f"Model loaded: {allocated:.1f} GB VRAM")
|
| 235 |
+
# Expected: ~70-72 GB
|
| 236 |
+
|
| 237 |
+
del model, tokenizer
|
| 238 |
+
torch.cuda.empty_cache()
|
| 239 |
+
print("Smoke test passed.")
|
| 240 |
+
```
|
| 241 |
+
|
| 242 |
+
### 0.5b Backward-Pass Smoke Test <!-- RED-HAT FIX #11: Added per M10 review — verifies gradients flow through expert adapters before committing budget -->
|
| 243 |
+
|
| 244 |
+
Run this AFTER the smoke test above. It loads the model via the same code path as the actual training script (Appendix C), applies LoRA, and verifies a backward pass updates adapter weights. This catches C2 (silent no-op training), C3 (eval crash), and H2 (wrong module topology) for ~$0.50.
|
| 245 |
+
|
| 246 |
+
```python
|
| 247 |
+
import os
|
| 248 |
+
os.environ["UNSLOTH_COMPILE_DISABLE"] = "1"
|
| 249 |
+
os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "expandable_segments:True"
|
| 250 |
+
os.environ["TOKENIZERS_PARALLELISM"] = "false"
|
| 251 |
+
os.environ["UNSLOTH_DISABLE_FAST_GENERATION"] = "1"
|
| 252 |
+
|
| 253 |
+
import torch
|
| 254 |
+
import hashlib
|
| 255 |
+
from transformers import AutoModelForCausalLM, AutoTokenizer
|
| 256 |
+
from peft import LoraConfig, get_peft_model
|
| 257 |
+
|
| 258 |
+
model_path = "/workspace/models/Qwen3.6-35B-A3B"
|
| 259 |
+
tokenizer = AutoTokenizer.from_pretrained(model_path)
|
| 260 |
+
model = AutoModelForCausalLM.from_pretrained(
|
| 261 |
+
model_path, torch_dtype=torch.bfloat16, device_map="cuda"
|
| 262 |
+
)
|
| 263 |
+
|
| 264 |
+
# Apply a minimal LoRA to a few modules to verify the training code path
|
| 265 |
+
# Use a small subset of target_modules to keep it fast
|
| 266 |
+
test_targets = []
|
| 267 |
+
for proj in ["q_proj", "k_proj"]:
|
| 268 |
+
test_targets.append(f"model.layers.0.self_attn.{proj}")
|
| 269 |
+
# Add one expert module to verify expert targeting works
|
| 270 |
+
test_targets.append("model.layers.0.mlp.experts.0.gate_proj")
|
| 271 |
+
|
| 272 |
+
lora_config = LoraConfig(
|
| 273 |
+
r=8, lora_alpha=16, lora_dropout=0.0,
|
| 274 |
+
target_modules=test_targets,
|
| 275 |
+
bias="none", task_type="CAUSAL_LM",
|
| 276 |
+
)
|
| 277 |
+
model = get_peft_model(model, lora_config)
|
| 278 |
+
|
| 279 |
+
# Verify trainable params include expert modules
|
| 280 |
+
expert_lora_count = sum(
|
| 281 |
+
1 for name, p in model.named_parameters()
|
| 282 |
+
if "experts" in name and "lora" in name and p.requires_grad
|
| 283 |
+
)
|
| 284 |
+
assert expert_lora_count > 0, "FATAL: No LoRA adapters on expert layers!"
|
| 285 |
+
|
| 286 |
+
# Hash a LoRA weight before backward pass
|
| 287 |
+
lora_param = None
|
| 288 |
+
for name, p in model.named_parameters():
|
| 289 |
+
if "lora" in name and p.requires_grad:
|
| 290 |
+
lora_param = (name, p)
|
| 291 |
+
break
|
| 292 |
+
pre_hash = hashlib.sha256(lora_param[1].data.cpu().numpy().tobytes()).hexdigest()
|
| 293 |
+
|
| 294 |
+
# Forward + backward on one batch
|
| 295 |
+
inputs = tokenizer("The daemon watches the threshold.", return_tensors="pt").to("cuda")
|
| 296 |
+
labels = inputs["input_ids"].clone()
|
| 297 |
+
outputs = model(**inputs, labels=labels)
|
| 298 |
+
loss = outputs.loss
|
| 299 |
+
loss.backward()
|
| 300 |
+
|
| 301 |
+
# Verify gradients exist
|
| 302 |
+
assert lora_param[1].grad is not None, f"FATAL: No gradient on {lora_param[0]}"
|
| 303 |
+
assert lora_param[1].grad.abs().sum() > 0, f"FATAL: Zero gradient on {lora_param[0]}"
|
| 304 |
+
|
| 305 |
+
# Simulate one optimizer step and verify weight changed
|
| 306 |
+
with torch.no_grad():
|
| 307 |
+
lora_param[1].data -= 0.01 * lora_param[1].grad
|
| 308 |
+
post_hash = hashlib.sha256(lora_param[1].data.cpu().numpy().tobytes()).hexdigest()
|
| 309 |
+
assert pre_hash != post_hash, "FATAL: LoRA weight unchanged after optimizer step!"
|
| 310 |
+
|
| 311 |
+
print(f"Backward-pass smoke test PASSED. Loss: {loss.item():.4f}")
|
| 312 |
+
print(f"Expert LoRA params: {expert_lora_count}")
|
| 313 |
+
print(f"Weight hash changed: {pre_hash[:12]}... -> {post_hash[:12]}...")
|
| 314 |
+
|
| 315 |
+
del model, tokenizer
|
| 316 |
+
torch.cuda.empty_cache()
|
| 317 |
+
```
|
| 318 |
+
|
| 319 |
+
### 0.6 Preflight Gate
|
| 320 |
+
|
| 321 |
+
Do NOT proceed to Phase 1 unless ALL of the following are true:
|
| 322 |
+
|
| 323 |
+
- [ ] H200 detected with >= 140 GB VRAM
|
| 324 |
+
- [ ] Container RAM >= 270 GB (cgroup check, not free -h)
|
| 325 |
+
- [ ] All Python packages import successfully at correct versions
|
| 326 |
+
- [ ] Model downloads and generates text
|
| 327 |
+
- [ ] Model VRAM after load is ~70-72 GB (confirms bf16, not accidental quantization)
|
| 328 |
+
- [ ] All training data files present and readable
|
| 329 |
+
- [ ] Disk space >= 400 GB free <!-- RED-HAT FIX #5: Increased from 150 GB per C5 review -->
|
| 330 |
+
- [ ] Environment variables are set
|
| 331 |
+
- [ ] Backward-pass smoke test passed (Section 0.5b) <!-- RED-HAT FIX #11: Added per M10 review -->
|
| 332 |
+
- [ ] Expert LoRA adapters confirmed on expert modules (backward-pass test)
|
| 333 |
+
|
| 334 |
+
---
|
| 335 |
+
|
| 336 |
+
## 4. Phase 1: Expert Profiling {#4-phase-1-expert-profiling}
|
| 337 |
+
|
| 338 |
+
**Cost:** ~$1 (30 minutes)
|
| 339 |
+
**VRAM:** ~72 GB (model inference only, no optimizer states)
|
| 340 |
+
**Purpose:** Identify which experts activate most frequently for our training data. This determines WHERE we apply LoRA.
|
| 341 |
+
|
| 342 |
+
### Method: ESFT-Style Profiling with MAN Scoring
|
| 343 |
+
|
| 344 |
+
Per ESFT (arXiv:2407.01906), 32 samples (~131K tokens) is sufficient for stable expert profiling. We profile on a representative subset of our actual training data to identify the experts that will handle Daimon's workload.
|
| 345 |
+
|
| 346 |
+
Per the Unified Expert Scoring Framework (arXiv:2606.15716, tested directly on Qwen3-30B-A3B), MAN (Mean Activation Norm) is the best scoring method, outperforming frequency, MSAN, REAP, and SEER.
|
| 347 |
+
|
| 348 |
+
### Profiling Script
|
| 349 |
+
|
| 350 |
+
See Appendix B for the complete script. The algorithm:
|
| 351 |
+
|
| 352 |
+
1. Load model in bf16 (inference only, ~72 GB VRAM)
|
| 353 |
+
2. Sample 32 examples from training data: 8 from Sonnet voice, 8 from prosocial, 8 from tool use, 4 from reasoning, 4 from Thomas demos
|
| 354 |
+
3. For each example, hook every MoE layer's router
|
| 355 |
+
4. Record: (a) which experts are selected by top-k routing, (b) the activation norm of each selected expert's output
|
| 356 |
+
5. Score each expert per layer by Mean Activation Norm across all 32 samples
|
| 357 |
+
6. Identify "hot" experts: top N experts per layer where cumulative MAN covers >= 80% of total activation norm
|
| 358 |
+
7. Output: `hot_experts.json` mapping layer_idx -> list of hot expert indices
|
| 359 |
+
|
| 360 |
+
### Expected Output
|
| 361 |
+
|
| 362 |
+
Based on Qwen3.6's architecture (256 routed experts, top-8 routing), we expect:
|
| 363 |
+
|
| 364 |
+
- ~20-40 hot experts per layer (8-16% of 256)
|
| 365 |
+
- Strong skew: Qwen3.6 uses fine-grained experts with sharp routing, meaning a small subset handles most traffic for any given data distribution
|
| 366 |
+
- Consistency across layers: some experts will be hot in most layers, others layer-specific
|
| 367 |
+
|
| 368 |
+
### Profiling Gate
|
| 369 |
+
|
| 370 |
+
- [ ] `hot_experts.json` generated successfully
|
| 371 |
+
- [ ] Each layer has 15-50 hot experts (if <15 or >100, something is wrong with the profiling)
|
| 372 |
+
- [ ] Activation norm distribution shows clear skew (top 10% of experts carry >50% of norm)
|
| 373 |
+
- [ ] Save `hot_experts.json` to persistent storage -- this is the key artifact for Phase 2
|
| 374 |
+
|
| 375 |
+
---
|
| 376 |
+
|
| 377 |
+
## 5. Phase 2: Foundation Training {#5-phase-2-foundation-training}
|
| 378 |
+
|
| 379 |
+
**Cost:** ~$22-32 (11-16 hours), gated by throughput probe <!-- RED-HAT FIX #4: Corrected from $20-28 to realistic raw-PEFT throughput per C4 review -->
|
| 380 |
+
**VRAM:** ~90-115 GB peak (model bf16 70-72 GB + adapters/optimizer 3-8 GB + activations 8-15 GB + CE logits ~5 GB + CUDA overhead 3-6 GB) <!-- RED-HAT FIX #6: Corrected from 55-65 GB per H1 review -->
|
| 381 |
+
**Purpose:** Train the complete Daimon capability and voice foundation in a single mixed-data pass
|
| 382 |
+
|
| 383 |
+
### 5.1 Dataset Composition
|
| 384 |
+
|
| 385 |
+
All training data is mixed into a single dataset with proportional sampling weights. Small datasets are upsampled to prevent being drowned out by large ones.
|
| 386 |
+
|
| 387 |
+
| Dataset | Raw Count | Upsample Factor | Effective Count | Sampling Weight | Purpose |
|
| 388 |
+
|---|---|---|---|---|---|
|
| 389 |
+
| Sonnet voice distill | 226,000 | 1x | 226,000 | 0.44 | Core voice and linguistic quality |
|
| 390 |
+
| Prosocial dialogue | 165,000 | 0.3x (subsample) | 49,500 | 0.10 | Values alignment, prosocial patterns |
|
| 391 |
+
| xlam function calling | 60,000 | 1x | 60,000 | 0.12 | Tool use capability |
|
| 392 |
+
| Hermes agent traces | 7,600 | 3x | 22,800 | 0.04 | Multi-step reasoning |
|
| 393 |
+
| Code review feedback | 9,500 | 2x | 19,000 | 0.04 | Constructive feedback style |
|
| 394 |
+
| TruthfulQA | 817 | 15x | 12,255 | 0.02 | Anti-sycophancy inoculation |
|
| 395 |
+
| **Total effective** | | | **~390,000** | **1.00** | |
|
| 396 |
+
|
| 397 |
+
**Why subsample prosocial to 30%:** 165K prosocial dialogue examples would dominate the voice signal from the 226K Sonnet distill. At 49.5K, it provides sufficient alignment signal without overwhelming the primary voice training data. The prosocial patterns are relatively simple (be helpful, don't be harmful) and converge faster than complex voice patterns.
|
| 398 |
+
|
| 399 |
+
**Why upsample TruthfulQA 15x:** 817 examples in a 390K-sample dataset would be seen <1 time per epoch. At 15x, the model sees each TruthfulQA example ~15 times, providing sufficient anti-sycophancy signal. Overfitting risk on 817 unique examples is manageable at this upsample factor with LoRA.
|
| 400 |
+
|
| 401 |
+
### 5.2 Data Format
|
| 402 |
+
|
| 403 |
+
All data must be in Unsloth's chat template format. For Qwen3.6:
|
| 404 |
+
|
| 405 |
+
```python
|
| 406 |
+
# Each example as a list of messages
|
| 407 |
+
{"messages": [
|
| 408 |
+
{"role": "system", "content": "You are Daimon, a strategic thinking partner..."},
|
| 409 |
+
{"role": "user", "content": "..."},
|
| 410 |
+
{"role": "assistant", "content": "..."}
|
| 411 |
+
]}
|
| 412 |
+
```
|
| 413 |
+
|
| 414 |
+
**System prompt for voice distill data:** Use a minimal system prompt (or none) to let the Sonnet voice dominate. The voice demonstrations carry the persona; the system prompt is a runtime artifact, not a training signal.
|
| 415 |
+
|
| 416 |
+
**System prompt for tool use data:** Include the tool definitions in the system prompt, matching the xlam format.
|
| 417 |
+
|
| 418 |
+
**System prompt for Thomas voice demos:** No system prompt. The separation principle: how to be, not what to recite. The demos teach the model Thomas's voice through example, not instruction.
|
| 419 |
+
|
| 420 |
+
### 5.3 LoRA Target Module Construction
|
| 421 |
+
|
| 422 |
+
This is the core HELLoRA implementation. Using the `hot_experts.json` from Phase 1, construct the PEFT `target_modules` list:
|
| 423 |
+
|
| 424 |
+
```python
|
| 425 |
+
import json
|
| 426 |
+
|
| 427 |
+
with open("hot_experts.json") as f:
|
| 428 |
+
hot_experts = json.load(f)
|
| 429 |
+
|
| 430 |
+
# Build target module list for PEFT
|
| 431 |
+
target_modules = []
|
| 432 |
+
|
| 433 |
+
# Attention projections (all layers) -- secondary target per HELLoRA
|
| 434 |
+
for layer_idx in range(40):
|
| 435 |
+
for proj in ["q_proj", "k_proj", "v_proj", "o_proj"]:
|
| 436 |
+
target_modules.append(f"model.layers.{layer_idx}.self_attn.{proj}")
|
| 437 |
+
|
| 438 |
+
# Hot expert FFN layers (primary target per HELLoRA)
|
| 439 |
+
for layer_idx_str, expert_indices in hot_experts.items():
|
| 440 |
+
layer_idx = int(layer_idx_str)
|
| 441 |
+
for expert_idx in expert_indices:
|
| 442 |
+
for proj in ["gate_proj", "up_proj", "down_proj"]:
|
| 443 |
+
target_modules.append(
|
| 444 |
+
f"model.layers.{layer_idx}.mlp.experts.{expert_idx}.{proj}"
|
| 445 |
+
)
|
| 446 |
+
|
| 447 |
+
# Explicitly NOT included:
|
| 448 |
+
# - Router weights (model.layers.*.mlp.gate.*) -- FROZEN
|
| 449 |
+
# - Shared expert (model.layers.*.mlp.shared_expert.*) -- FROZEN
|
| 450 |
+
# - Cold expert FFN layers -- FROZEN
|
| 451 |
+
# - Embedding / LM head -- FROZEN
|
| 452 |
+
|
| 453 |
+
print(f"Total target modules: {len(target_modules)}")
|
| 454 |
+
# Expected: 160 attention modules + (hot_experts_per_layer * 40 layers * 3 projections)
|
| 455 |
+
# If ~25 hot experts/layer: 160 + 25*40*3 = 160 + 3000 = 3160 modules
|
| 456 |
+
```
|
| 457 |
+
|
| 458 |
+
### 5.4 Training Configuration
|
| 459 |
+
|
| 460 |
+
```python
|
| 461 |
+
from unsloth import FastLanguageModel
|
| 462 |
+
from trl import SFTTrainer, SFTConfig
|
| 463 |
+
import torch
|
| 464 |
+
|
| 465 |
+
# Load model
|
| 466 |
+
model, tokenizer = FastLanguageModel.from_pretrained(
|
| 467 |
+
"Qwen/Qwen3.6-35B-A3B",
|
| 468 |
+
max_seq_length=2048,
|
| 469 |
+
load_in_4bit=False, # CRITICAL: Must be False
|
| 470 |
+
dtype=torch.bfloat16,
|
| 471 |
+
)
|
| 472 |
+
|
| 473 |
+
# Apply HELLoRA with hot expert targeting
|
| 474 |
+
# NOTE: If Unsloth's get_peft_model does not support explicit module name lists,
|
| 475 |
+
# use PEFT directly:
|
| 476 |
+
from peft import LoraConfig, get_peft_model
|
| 477 |
+
|
| 478 |
+
lora_config = LoraConfig(
|
| 479 |
+
r=32,
|
| 480 |
+
lora_alpha=64,
|
| 481 |
+
lora_dropout=0.05,
|
| 482 |
+
target_modules=target_modules, # From Section 5.3
|
| 483 |
+
use_dora=True, # DoRA upgrade (+1-4% accuracy)
|
| 484 |
+
bias="none",
|
| 485 |
+
task_type="CAUSAL_LM",
|
| 486 |
+
)
|
| 487 |
+
|
| 488 |
+
model = get_peft_model(model, lora_config)
|
| 489 |
+
|
| 490 |
+
# VERIFY trainable params include expert modules
|
| 491 |
+
trainable_params = sum(p.numel() for p in model.parameters() if p.requires_grad)
|
| 492 |
+
total_params = sum(p.numel() for p in model.parameters())
|
| 493 |
+
print(f"Trainable: {trainable_params:,} ({100*trainable_params/total_params:.2f}%)")
|
| 494 |
+
# Expected: 0.3-0.8% of total params (per HELLoRA paper)
|
| 495 |
+
# If >3%, something is wrong -- too many modules are unfrozen
|
| 496 |
+
# If <0.1%, hot expert targeting is too aggressive -- expand threshold
|
| 497 |
+
|
| 498 |
+
# CRITICAL CHECK (mlx-lm bug #571 equivalent for PEFT):
|
| 499 |
+
# Verify that expert modules actually have LoRA adapters
|
| 500 |
+
expert_lora_count = sum(1 for name, _ in model.named_parameters()
|
| 501 |
+
if "experts" in name and "lora" in name and _.requires_grad)
|
| 502 |
+
print(f"Expert LoRA params: {expert_lora_count}")
|
| 503 |
+
assert expert_lora_count > 0, "FATAL: No LoRA adapters on expert layers!"
|
| 504 |
+
|
| 505 |
+
# RED-HAT FIX #3: Create stratified eval split BEFORE interleaving training data (per C3 review).
|
| 506 |
+
# Without this, eval_strategy="steps" crashes at init or first eval step.
|
| 507 |
+
# ~250-300 examples held out, stratified across all 6 sources, fixed seed, excluded from training.
|
| 508 |
+
# See build_foundation_dataset() in Appendix C for implementation.
|
| 509 |
+
|
| 510 |
+
# Training config
|
| 511 |
+
training_args = SFTConfig(
|
| 512 |
+
output_dir="/workspace/checkpoints/daimon_v2_foundation",
|
| 513 |
+
per_device_train_batch_size=4, # 90-115 GB peak leaves room for batch=4 on H200
|
| 514 |
+
gradient_accumulation_steps=4, # Effective batch = 16
|
| 515 |
+
num_train_epochs=1, # Single epoch for ~390K examples
|
| 516 |
+
learning_rate=2e-4,
|
| 517 |
+
lr_scheduler_type="cosine",
|
| 518 |
+
warmup_ratio=0.03, # ~750 warmup steps
|
| 519 |
+
weight_decay=0.01,
|
| 520 |
+
bf16=True,
|
| 521 |
+
max_seq_length=2048,
|
| 522 |
+
logging_steps=50,
|
| 523 |
+
save_steps=2000, # Checkpoint every 2000 steps (~$2 of training)
|
| 524 |
+
save_total_limit=5, # Keep last 5 checkpoints to manage disk
|
| 525 |
+
dataloader_num_workers=0, # CRITICAL: Prevents MoE deadlocks
|
| 526 |
+
dataset_num_proc=1, # CRITICAL: Prevents MoE deadlocks
|
| 527 |
+
gradient_checkpointing=True, # Required for memory safety
|
| 528 |
+
gradient_checkpointing_kwargs={"use_reentrant": False},
|
| 529 |
+
eval_strategy="steps",
|
| 530 |
+
eval_steps=2000, # Eval at every checkpoint
|
| 531 |
+
load_best_model_at_end=True, # RED-HAT FIX #3: Select best checkpoint by eval_loss (per C3 review)
|
| 532 |
+
metric_for_best_model="eval_loss", # RED-HAT FIX #3: Best = lowest eval loss
|
| 533 |
+
max_steps=30000, # RED-HAT FIX #4: Circuit breaker — hard cap to prevent budget overrun (per C4 review). Adjusted after throughput probe.
|
| 534 |
+
report_to="none", # Or "wandb" if experiment tracking is set up
|
| 535 |
+
seed=42,
|
| 536 |
+
)
|
| 537 |
+
```
|
| 538 |
+
|
| 539 |
+
### 5.5 Training Time Estimate
|
| 540 |
+
|
| 541 |
+
<!-- RED-HAT FIX #4: Replaced Unsloth kernel speeds with realistic raw-PEFT speeds per C4 review.
|
| 542 |
+
The shipped script uses raw PEFT + Transformers, not Unsloth MoE kernels.
|
| 543 |
+
Unsloth throughput (1.5-2.0 s/step) is only achievable if FastLanguageModel.get_peft_model()
|
| 544 |
+
supports explicit module-name lists. Until verified, budget for the slower path. -->
|
| 545 |
+
|
| 546 |
+
| Parameter | Value |
|
| 547 |
+
|---|---|
|
| 548 |
+
| Effective samples | ~390,000 |
|
| 549 |
+
| Effective batch size | 16 |
|
| 550 |
+
| Steps per epoch | ~24,375 |
|
| 551 |
+
| Estimated sec/step (raw PEFT, no Unsloth kernels) | 2.5-4.0 |
|
| 552 |
+
| Wall-clock estimate | 17-27 hours |
|
| 553 |
+
| Cost at $2/hr | $34-54 (but gated by throughput probe — see 5.5b) |
|
| 554 |
+
| Realistic budget after probe gate | $22-32 (probe rejects runs above this) |
|
| 555 |
+
|
| 556 |
+
**Note:** If Unsloth's `FastLanguageModel.get_peft_model()` accepts explicit module-name lists (verify during preflight), throughput improves to ~1.5-2.0 s/step and cost drops to ~$20-28. The probe (Section 5.5b) measures the actual throughput and gates accordingly.
|
| 557 |
+
|
| 558 |
+
### 5.5b Throughput Probe (Hard Gate) <!-- RED-HAT FIX #4 + RED-HAT FIX #10: Added per C4 review -->
|
| 559 |
+
|
| 560 |
+
**Cost:** ~$0.50-1.00 (50 steps)
|
| 561 |
+
**Purpose:** Measure actual tokens/sec before committing budget. This is the single most valuable addition to the pipeline.
|
| 562 |
+
|
| 563 |
+
Run 50 training steps on the real data mix with the real config. Measure tokens/sec and project Phase 2 cost.
|
| 564 |
+
|
| 565 |
+
```python
|
| 566 |
+
# After model + LoRA setup, before full training:
|
| 567 |
+
import time
|
| 568 |
+
|
| 569 |
+
# Run 50 steps as a probe
|
| 570 |
+
probe_args = SFTConfig(
|
| 571 |
+
output_dir="/workspace/checkpoints/daimon_v2_probe",
|
| 572 |
+
per_device_train_batch_size=4,
|
| 573 |
+
gradient_accumulation_steps=4,
|
| 574 |
+
max_steps=50,
|
| 575 |
+
learning_rate=2e-4,
|
| 576 |
+
bf16=True,
|
| 577 |
+
max_seq_length=2048,
|
| 578 |
+
logging_steps=10,
|
| 579 |
+
save_strategy="no",
|
| 580 |
+
dataloader_num_workers=0,
|
| 581 |
+
dataset_num_proc=1,
|
| 582 |
+
gradient_checkpointing=True,
|
| 583 |
+
gradient_checkpointing_kwargs={"use_reentrant": False},
|
| 584 |
+
report_to="none",
|
| 585 |
+
seed=42,
|
| 586 |
+
)
|
| 587 |
+
|
| 588 |
+
probe_trainer = SFTTrainer(
|
| 589 |
+
model=model,
|
| 590 |
+
args=probe_args,
|
| 591 |
+
train_dataset=train_dataset,
|
| 592 |
+
processing_class=tokenizer,
|
| 593 |
+
)
|
| 594 |
+
|
| 595 |
+
start_time = time.time()
|
| 596 |
+
probe_trainer.train()
|
| 597 |
+
elapsed = time.time() - start_time
|
| 598 |
+
|
| 599 |
+
secs_per_step = elapsed / 50
|
| 600 |
+
tokens_per_sec = (4 * 4 * 2048) / secs_per_step # batch * accum * seq_len
|
| 601 |
+
projected_hours = (24375 * secs_per_step) / 3600
|
| 602 |
+
projected_cost = projected_hours * 2.0 # at $2/hr
|
| 603 |
+
|
| 604 |
+
print(f"Throughput probe results:")
|
| 605 |
+
print(f" {secs_per_step:.2f} sec/step")
|
| 606 |
+
print(f" {tokens_per_sec:.0f} tokens/sec")
|
| 607 |
+
print(f" Projected Phase 2: {projected_hours:.1f} hours, ${projected_cost:.0f}")
|
| 608 |
+
|
| 609 |
+
# HARD GATE: abort if projected cost exceeds budget cap
|
| 610 |
+
COST_CAP = 35.0 # Max dollars for Phase 2
|
| 611 |
+
if projected_cost > COST_CAP:
|
| 612 |
+
print(f"ABORT: Projected cost ${projected_cost:.0f} exceeds cap ${COST_CAP:.0f}.")
|
| 613 |
+
print(f"Fallback: Switch to Tier-1 Unsloth all-expert LoRA r=16.")
|
| 614 |
+
raise SystemExit(1)
|
| 615 |
+
|
| 616 |
+
# Set max_steps based on probe to enforce cost ceiling
|
| 617 |
+
safe_max_steps = int(COST_CAP / 2.0 * 3600 / secs_per_step)
|
| 618 |
+
print(f" Setting max_steps={safe_max_steps} as cost circuit breaker")
|
| 619 |
+
```
|
| 620 |
+
|
| 621 |
+
**Fallback if probe fails (projected cost > $35):** Switch to Tier-1 Unsloth-native all-expert LoRA with r=16, alpha=16 -- the one configuration with measured throughput on this model family.
|
| 622 |
+
|
| 623 |
+
### 5.6 Monitoring During Training
|
| 624 |
+
|
| 625 |
+
Watch for these during Phase 2:
|
| 626 |
+
|
| 627 |
+
```python
|
| 628 |
+
# Add to training callbacks or check manually between checkpoints:
|
| 629 |
+
|
| 630 |
+
# 1. Loss curve -- should decrease smoothly
|
| 631 |
+
# Red flag: sudden spike or plateau before step 5000
|
| 632 |
+
|
| 633 |
+
# 2. VRAM usage -- should be stable
|
| 634 |
+
# Red flag: gradual increase (memory leak from activation caching)
|
| 635 |
+
|
| 636 |
+
# 3. Learning rate -- should follow cosine schedule
|
| 637 |
+
# Red flag: NaN or zero (optimizer failure)
|
| 638 |
+
```
|
| 639 |
+
|
| 640 |
+
Monitor expert utilization at eval steps (add as callback):
|
| 641 |
+
|
| 642 |
+
```python
|
| 643 |
+
# Quick check: are hot experts still hot?
|
| 644 |
+
# If routing has shifted significantly, the LoRA placement is misaligned.
|
| 645 |
+
# This is informational -- do not retrain Phase 1 mid-run.
|
| 646 |
+
```
|
| 647 |
+
|
| 648 |
+
### 5.6b Checkpoint Egress <!-- RED-HAT FIX #8: Added per H9 review — checkpoints must leave the pod after each save -->
|
| 649 |
+
|
| 650 |
+
**Spot instances can terminate at any time.** Checkpoints only on the pod volume are not insurance -- they are gone when the volume is reclaimed. After every checkpoint save, rsync the adapter directory to persistent off-pod storage.
|
| 651 |
+
|
| 652 |
+
```bash
|
| 653 |
+
# Add as a post-checkpoint callback or cron job running every 30 minutes:
|
| 654 |
+
# Adapter checkpoint is ~3-4 GB (includes optimizer state for resuming)
|
| 655 |
+
|
| 656 |
+
REMOTE_DEST="user@margaret.local:/workspace/daimon_checkpoints/" # or MTH, or B2 bucket
|
| 657 |
+
CHECKPOINT_DIR="/workspace/checkpoints/daimon_v2_foundation"
|
| 658 |
+
|
| 659 |
+
# Sync latest checkpoint off-pod after each save
|
| 660 |
+
rsync -avz --progress \
|
| 661 |
+
"${CHECKPOINT_DIR}/$(ls -td ${CHECKPOINT_DIR}/checkpoint-* | head -1)" \
|
| 662 |
+
"${REMOTE_DEST}"
|
| 663 |
+
|
| 664 |
+
# Also sync critical artifacts
|
| 665 |
+
rsync -avz /workspace/artifacts/hot_experts.json "${REMOTE_DEST}"
|
| 666 |
+
```
|
| 667 |
+
|
| 668 |
+
**Cost:** Pennies of bandwidth. Converts "total loss on spot preemption" into "resume from last checkpoint - 2000 steps."
|
| 669 |
+
|
| 670 |
+
Add inter-phase disk check:
|
| 671 |
+
```bash
|
| 672 |
+
# Run between phases to catch disk pressure before it corrupts a save
|
| 673 |
+
DISK_FREE_GB=$(df --output=avail /workspace | tail -1 | awk '{print int($1/1048576)}')
|
| 674 |
+
echo "Disk free: ${DISK_FREE_GB} GB"
|
| 675 |
+
if [ "$DISK_FREE_GB" -lt 80 ]; then
|
| 676 |
+
echo "WARNING: Disk below 80 GB free. Clean stale checkpoints before continuing."
|
| 677 |
+
fi
|
| 678 |
+
```
|
| 679 |
+
|
| 680 |
+
### 5.7 Phase 2 Gate
|
| 681 |
+
|
| 682 |
+
Before proceeding to Phase 3:
|
| 683 |
+
|
| 684 |
+
- [ ] Training loss decreased smoothly (no divergence, no NaN)
|
| 685 |
+
- [ ] Final training loss < initial loss by at least 30%
|
| 686 |
+
- [ ] Validation loss is within 10% of training loss (no severe overfitting)
|
| 687 |
+
- [ ] Generate 5 test prompts manually and verify coherent output
|
| 688 |
+
- [ ] VRAM stayed within expected bounds (~90-115 GB peak, alarm at >125 GB sustained) <!-- RED-HAT FIX #6: Corrected from 55-65 GB per H1 review -->
|
| 689 |
+
- [ ] All 5 most recent checkpoints saved successfully
|
| 690 |
+
- [ ] Best checkpoint selected by `load_best_model_at_end=True` via eval_loss <!-- RED-HAT FIX #3: Auto-selected, not manual -->
|
| 691 |
+
- [ ] Best checkpoint copied to `/workspace/checkpoints/daimon_v2_foundation_best/`
|
| 692 |
+
- [ ] Best checkpoint synced off-pod to persistent storage <!-- RED-HAT FIX #8: Per H9 review -->
|
| 693 |
+
|
| 694 |
+
---
|
| 695 |
+
|
| 696 |
+
## 6. Phase 3: Thomas Voice Calibration {#6-phase-3-thomas-voice-calibration}
|
| 697 |
+
|
| 698 |
+
**Cost:** ~$1 (30 minutes)
|
| 699 |
+
**VRAM:** ~90-115 GB peak (same as Phase 2) <!-- RED-HAT FIX #6: Corrected from 55-65 GB per H1 review -->
|
| 700 |
+
**Purpose:** Final persona imprint. The Thomas voice demonstrations teach the model HOW to be Daimon -- the specific patterns of question-before-answer, metaphor-forward explanation, constructive challenge, dry wit. This is the separation principle: procedural knowledge, not declarative content.
|
| 701 |
+
|
| 702 |
+
### 6.1 Data
|
| 703 |
+
|
| 704 |
+
Source: `/home/HumboldtJoker/.coalition/specs/daimon_voice_demonstrations.jsonl`
|
| 705 |
+
Count: 32 procedural demonstrations (from the existing v1 spec, already authored)
|
| 706 |
+
Additional: If available, 8-10 more demonstrations covering edge cases (overwhelming user, technical disagreement, values-laden questions). Target 40 total.
|
| 707 |
+
|
| 708 |
+
Format: Already in `{"messages": [...]}` format. No system prompt -- pure procedural demonstration.
|
| 709 |
+
|
| 710 |
+
### 6.2 Training Configuration
|
| 711 |
+
|
| 712 |
+
```python
|
| 713 |
+
# RED-HAT FIX #2: Phase 2 saves adapter-only (adapter_config.json + adapter_model.safetensors),
|
| 714 |
+
# NOT a full model. Must load base model first, then attach adapter with is_trainable=True.
|
| 715 |
+
# Using AutoModelForCausalLM.from_pretrained on an adapter dir will either crash
|
| 716 |
+
# ("no config.json") or load adapter for inference-only (requires_grad=False),
|
| 717 |
+
# causing silent no-op training. (Per C2 review)
|
| 718 |
+
|
| 719 |
+
from peft import PeftModel
|
| 720 |
+
|
| 721 |
+
# Load base model first
|
| 722 |
+
base_model = AutoModelForCausalLM.from_pretrained(
|
| 723 |
+
"/workspace/models/Qwen3.6-35B-A3B", # Base model, NOT the checkpoint
|
| 724 |
+
torch_dtype=torch.bfloat16,
|
| 725 |
+
device_map="cuda",
|
| 726 |
+
)
|
| 727 |
+
tokenizer = AutoTokenizer.from_pretrained("/workspace/models/Qwen3.6-35B-A3B")
|
| 728 |
+
|
| 729 |
+
# Attach Phase 2 adapter with is_trainable=True
|
| 730 |
+
model = PeftModel.from_pretrained(
|
| 731 |
+
base_model,
|
| 732 |
+
"/workspace/checkpoints/daimon_v2_foundation_best",
|
| 733 |
+
is_trainable=True, # CRITICAL: without this, all adapter params are frozen
|
| 734 |
+
)
|
| 735 |
+
|
| 736 |
+
# RED-HAT FIX #2: Verify adapter is trainable (same assert as Phase 2)
|
| 737 |
+
trainable_count = sum(p.numel() for p in model.parameters() if p.requires_grad)
|
| 738 |
+
assert trainable_count > 0, "FATAL: No trainable parameters in Phase 3! Check is_trainable=True."
|
| 739 |
+
expert_lora_count = sum(
|
| 740 |
+
1 for name, p in model.named_parameters()
|
| 741 |
+
if "experts" in name and "lora" in name and p.requires_grad
|
| 742 |
+
)
|
| 743 |
+
assert expert_lora_count > 0, "FATAL: No LoRA adapters on expert layers in Phase 3!"
|
| 744 |
+
print(f"Phase 3 trainable params: {trainable_count:,}, expert LoRA params: {expert_lora_count}")
|
| 745 |
+
|
| 746 |
+
# RED-HAT FIX #2: Weight-delta assert — hash one LoRA tensor, run one step, verify it changed
|
| 747 |
+
import hashlib
|
| 748 |
+
lora_param = next((name, p) for name, p in model.named_parameters()
|
| 749 |
+
if "lora" in name and p.requires_grad)
|
| 750 |
+
pre_hash = hashlib.sha256(lora_param[1].data.cpu().numpy().tobytes()).hexdigest()
|
| 751 |
+
|
| 752 |
+
training_args = SFTConfig(
|
| 753 |
+
output_dir="/workspace/checkpoints/daimon_v2_calibration",
|
| 754 |
+
per_device_train_batch_size=2, # Smaller batch for tiny dataset
|
| 755 |
+
gradient_accumulation_steps=1, # Effective batch = 2
|
| 756 |
+
num_train_epochs=5, # More epochs for tiny dataset
|
| 757 |
+
learning_rate=1e-5, # 20x lower than Phase 2
|
| 758 |
+
lr_scheduler_type="cosine",
|
| 759 |
+
warmup_ratio=0.1,
|
| 760 |
+
weight_decay=0.01,
|
| 761 |
+
bf16=True,
|
| 762 |
+
max_seq_length=2048,
|
| 763 |
+
logging_steps=5,
|
| 764 |
+
save_steps=20, # Checkpoint every 20 steps
|
| 765 |
+
save_total_limit=10, # Keep all checkpoints (tiny, <5 GB total)
|
| 766 |
+
dataloader_num_workers=0,
|
| 767 |
+
dataset_num_proc=1,
|
| 768 |
+
gradient_checkpointing=True,
|
| 769 |
+
gradient_checkpointing_kwargs={"use_reentrant": False},
|
| 770 |
+
seed=42,
|
| 771 |
+
)
|
| 772 |
+
```
|
| 773 |
+
|
| 774 |
+
### 6.3 Why This Works
|
| 775 |
+
|
| 776 |
+
With only 32-40 examples, the risk is memorization rather than generalization. The mitigations:
|
| 777 |
+
|
| 778 |
+
1. **Low learning rate (1e-5):** 20x lower than Phase 2. The model has already learned the general capability; this stage makes subtle adjustments to the voice distribution.
|
| 779 |
+
2. **The examples are PROCEDURAL:** They teach response patterns (question-before-answer, assess-then-open), not factual content. Procedural patterns generalize better than factual memorization because they operate on structural templates, not surface forms.
|
| 780 |
+
3. **5 epochs is intentional:** Each example is seen 5 times. With LoRA r=32 and DoRA, the adapter capacity is sufficient to learn procedural patterns without memorizing specific words. The cosine LR decay means later epochs make smaller adjustments.
|
| 781 |
+
4. **Recency effect:** This is the last training the model sees. The Thomas voice patterns are the freshest in the gradient signal, which means they have the strongest influence on generation.
|
| 782 |
+
|
| 783 |
+
### 6.4 Phase 3 Gate
|
| 784 |
+
|
| 785 |
+
- [ ] Training completed without divergence
|
| 786 |
+
- [ ] Generate the 5 held-out Thomas voice evaluation prompts (not in training data)
|
| 787 |
+
- [ ] Manual review: does it sound like Thomas? Apply the daimon_voice_analysis.md criteria:
|
| 788 |
+
- Question-before-answer pattern present?
|
| 789 |
+
- Metaphor usage natural, not forced?
|
| 790 |
+
- Constructive challenge without preachiness?
|
| 791 |
+
- Appropriate register (professional, not casual)?
|
| 792 |
+
- Direct assessment ("The issue is..." "My read is...")?
|
| 793 |
+
- [ ] Compare against Phase 2 output on same prompts -- Phase 3 should be noticeably more "Thomas" without losing capability
|
| 794 |
+
- [ ] Select best checkpoint (likely epoch 3 or 4 based on typical LoRA convergence)
|
| 795 |
+
|
| 796 |
+
---
|
| 797 |
+
|
| 798 |
+
## 7. Phase 4: Post-Training {#7-phase-4-post-training}
|
| 799 |
+
|
| 800 |
+
**Cost:** ~$2-4 (1-2 hours) <!-- RED-HAT FIX #14: Honest after trimming eval battery and cutting LASER from critical path -->
|
| 801 |
+
**Purpose:** Two improvements: OGPSA personality capture and eval battery. LASER moved to optional offline experiment (see 7.2). <!-- RED-HAT FIX #1: LASER cut from critical path per C1 review -->
|
| 802 |
+
|
| 803 |
+
### 7.1 Merge LoRA into Base Weights
|
| 804 |
+
|
| 805 |
+
Before post-training, merge the LoRA adapter into the base model. This is required because (a) vLLM cannot load MoE LoRA adapters, and (b) LASER operates on the merged weights.
|
| 806 |
+
|
| 807 |
+
```python
|
| 808 |
+
# Merge LoRA -> base weights
|
| 809 |
+
model = model.merge_and_unload()
|
| 810 |
+
|
| 811 |
+
# Save merged model in bf16
|
| 812 |
+
model.save_pretrained(
|
| 813 |
+
"/workspace/models/daimon_v2_merged",
|
| 814 |
+
safe_serialization=True,
|
| 815 |
+
)
|
| 816 |
+
tokenizer.save_pretrained("/workspace/models/daimon_v2_merged")
|
| 817 |
+
|
| 818 |
+
# Verify merged model generates correctly
|
| 819 |
+
inputs = tokenizer("What problem does the dashboard solve?", return_tensors="pt").to("cuda")
|
| 820 |
+
with torch.no_grad():
|
| 821 |
+
output = model.generate(**inputs, max_new_tokens=200, temperature=0.7)
|
| 822 |
+
print(tokenizer.decode(output[0]))
|
| 823 |
+
```
|
| 824 |
+
|
| 825 |
+
### 7.2 LASER (Layer-Selective Rank Reduction) — OPTIONAL POST-TRAINING EXPERIMENT
|
| 826 |
+
|
| 827 |
+
<!-- RED-HAT FIX #1: LASER moved from critical path to optional experiment per C1 review.
|
| 828 |
+
THREE bugs in the original:
|
| 829 |
+
1. SVD truncation was INVERTED — zeroed the LARGEST singular values (principal components)
|
| 830 |
+
instead of the tail. torch.linalg.svd returns values in DESCENDING order.
|
| 831 |
+
The paper's "higher-order components" means components associated with SMALL singular
|
| 832 |
+
values, not the highest-magnitude ones.
|
| 833 |
+
2. Blanket application to ALL 20,480 matrices contradicts LASER's name (LAyer-SElective
|
| 834 |
+
Rank reduction). The paper's gains come from searching for the specific (layer, matrix)
|
| 835 |
+
where reduction helps; most choices hurt.
|
| 836 |
+
3. "+20-30pp" was misquoted — those gains are on narrow factual-recall evals at the single
|
| 837 |
+
best layer on older models (GPT-J). Realistic upside on a 2026 instruct MoE: 0 to +1-2pp.
|
| 838 |
+
|
| 839 |
+
LASER is NOT free: 20,480 CPU SVDs of 2048x512 fp32 matrices takes 1-3 hours of pod time.
|
| 840 |
+
The rollback plan already calls it "a free bonus, not a requirement" — it isn't free and
|
| 841 |
+
as originally written it was negative-value (would have destroyed the trained model).
|
| 842 |
+
|
| 843 |
+
RECOMMENDED: Run LASER locally on Margaret after downloading merged weights ($0 cost).
|
| 844 |
+
Apply to ONE (layer, matrix) at a time, eval after each, keep only if it helps. -->
|
| 845 |
+
|
| 846 |
+
**Paper:** arXiv:2312.13558
|
| 847 |
+
**Code:** https://github.com/pratyushasharma/laser
|
| 848 |
+
**Status:** Optional post-budget experiment. Do NOT run on the cloud instance during the primary training run.
|
| 849 |
+
**Realistic upside:** 0 to +1-2pp on narrow factual-recall benchmarks at the single best (layer, matrix). Not a broad reasoning improvement.
|
| 850 |
+
|
| 851 |
+
**WARNING:** The original code in this spec was inverted and would have destroyed the trained model. The corrected version below keeps the TOP singular values and truncates the TAIL (small singular values = high-order noise).
|
| 852 |
+
|
| 853 |
+
```python
|
| 854 |
+
# CORRECTED LASER — run locally on Margaret, not on cloud pod
|
| 855 |
+
# RED-HAT FIX #1: Fixed SVD direction. S_modified[n_keep:] = 0 keeps top, drops tail.
|
| 856 |
+
|
| 857 |
+
import torch
|
| 858 |
+
|
| 859 |
+
def laser_reduce(weight_matrix, keep_fraction=0.95):
|
| 860 |
+
"""Keep top fraction of singular values, zero out the tail (noise).
|
| 861 |
+
|
| 862 |
+
LASER's insight: small singular values of MLP weight matrices often
|
| 863 |
+
encode noise. Removing them can improve factual recall.
|
| 864 |
+
|
| 865 |
+
NOTE: torch.linalg.svd returns singular values in DESCENDING order.
|
| 866 |
+
We keep the first n_keep values (largest) and zero the rest (smallest).
|
| 867 |
+
"""
|
| 868 |
+
original_dtype = weight_matrix.dtype
|
| 869 |
+
W = weight_matrix.float() # SVD requires float32
|
| 870 |
+
U, S, Vh = torch.linalg.svd(W, full_matrices=False)
|
| 871 |
+
n_keep = max(1, int(len(S) * keep_fraction))
|
| 872 |
+
S_modified = S.clone()
|
| 873 |
+
S_modified[n_keep:] = 0.0 # Zero out TAIL (small values = noise), keep TOP
|
| 874 |
+
W_modified = U @ torch.diag(S_modified) @ Vh
|
| 875 |
+
return W_modified.to(original_dtype)
|
| 876 |
+
|
| 877 |
+
# DO NOT apply blanket to all 20,480 matrices.
|
| 878 |
+
# Apply to ONE (layer, matrix) at a time, eval after each.
|
| 879 |
+
# Example: test on layer 20, gate_proj only:
|
| 880 |
+
# new_weight = laser_reduce(expert.gate_proj.weight.data, keep_fraction=0.95)
|
| 881 |
+
# Run eval. If improved, keep. If not, revert.
|
| 882 |
+
```
|
| 883 |
+
|
| 884 |
+
**LASER protocol (if attempting):**
|
| 885 |
+
1. Download merged weights to Margaret (local, $0)
|
| 886 |
+
2. For each candidate (layer_idx, proj_name) pair, in a systematic sweep:
|
| 887 |
+
a. Apply `laser_reduce` with `keep_fraction=0.95`
|
| 888 |
+
b. Run a quick eval subset (100-200 examples)
|
| 889 |
+
c. Keep only if eval improves over baseline
|
| 890 |
+
3. Try `keep_fraction` values: 0.93, 0.95, 0.97
|
| 891 |
+
4. Apply ONLY to expert FFN (gate_proj, up_proj). Do NOT apply to attention or down_proj.
|
| 892 |
+
|
| 893 |
+
### 7.3 OGPSA Personality Capture
|
| 894 |
+
|
| 895 |
+
Per ogpsa_persona_validation_v3.1.md, OGPSA extracts the personality subspace from the model's activations on persona-loaded text.
|
| 896 |
+
|
| 897 |
+
```python
|
| 898 |
+
# OGPSA Phase 1: Extract persona subspace
|
| 899 |
+
# Use the Thomas voice demonstrations as persona-loaded prompts
|
| 900 |
+
# Plus 10 general QA examples as contrast
|
| 901 |
+
|
| 902 |
+
# 1. Run persona-loaded examples through the model
|
| 903 |
+
# 2. Capture residual stream activations at all layers
|
| 904 |
+
# 3. SVD on the persona-loaded activation matrix
|
| 905 |
+
# 4. Top 16 components = persona subspace
|
| 906 |
+
# 5. Save as ogpsa_daimon_v2.json
|
| 907 |
+
|
| 908 |
+
# The OGPSA capture script is at ogpsa_red_team_code.md
|
| 909 |
+
# Adapt for Qwen3.6 architecture (40 layers, not 48)
|
| 910 |
+
```
|
| 911 |
+
|
| 912 |
+
Output: `ogpsa_daimon_v2.json` -- the personality geometry for drift monitoring.
|
| 913 |
+
|
| 914 |
+
### 7.4 Evaluation Battery
|
| 915 |
+
|
| 916 |
+
Run on the merged model (before any optional LASER experiments):
|
| 917 |
+
|
| 918 |
+
<!-- RED-HAT FIX #7: Removed truthfulqa_mc2 from eval battery per H8 review.
|
| 919 |
+
TruthfulQA is in the training data (upsampled 15x). Including it in eval
|
| 920 |
+
measures memorization, not anti-sycophancy generalization. Rely on the
|
| 921 |
+
custom 10-scenario anti-sycophancy eval instead (well-designed, not contaminated). -->
|
| 922 |
+
|
| 923 |
+
<!-- RED-HAT FIX #14: Trimmed eval battery per H10 review.
|
| 924 |
+
Removed arc_easy (redundant with arc_challenge). Increased batch_size to 16.
|
| 925 |
+
Added --limit 1500 for forgetting detection (5% regression detector, not leaderboard).
|
| 926 |
+
Single pass only (LASER moved to optional offline experiment). -->
|
| 927 |
+
|
| 928 |
+
```bash
|
| 929 |
+
# Using lm-evaluation-harness (install in separate venv per H5 review)
|
| 930 |
+
# python -m venv /workspace/lm_eval_venv && source /workspace/lm_eval_venv/bin/activate
|
| 931 |
+
# pip install lm_eval[hf]
|
| 932 |
+
|
| 933 |
+
# Baseline capabilities (forgetting detection)
|
| 934 |
+
lm_eval --model hf \
|
| 935 |
+
--model_args pretrained=/workspace/models/daimon_v2_merged,dtype=bfloat16 \
|
| 936 |
+
--tasks mmlu,gsm8k,hellaswag,arc_challenge,winogrande \
|
| 937 |
+
--batch_size 16 \
|
| 938 |
+
--limit 1500 \
|
| 939 |
+
--output_path /workspace/eval_results/daimon_v2/
|
| 940 |
+
|
| 941 |
+
# Compare against base model eval (run on unmodified Qwen3.6-35B-A3B for comparison)
|
| 942 |
+
```
|
| 943 |
+
|
| 944 |
+
**Custom Daimon-specific evals:**
|
| 945 |
+
|
| 946 |
+
| Eval | Method | Pass Criteria |
|
| 947 |
+
|---|---|---|
|
| 948 |
+
| **Voice attribution** | 5 evaluators, 20 blind response pairs (Daimon vs generic Qwen3.6) | >80% correct attribution to "Thomas-like" |
|
| 949 |
+
| **Daemon function** | 10 scenarios where something is wrong. Does Daimon flag it? | >7/10 flags the concern |
|
| 950 |
+
| **Anti-sycophancy** | 10 scenarios with bad ideas presented as good ones. Does Daimon push back? | >7/10 pushes back |
|
| 951 |
+
| **Anti-nag** | 10 scenarios where everything is fine. Does Daimon stay quiet? | >8/10 does not volunteer unnecessary concerns |
|
| 952 |
+
| **Tool use** | 10 multi-step agentic tasks with function calling | >8/10 correct tool selection and sequencing |
|
| 953 |
+
| **Register calibration** | 5 prompts across different contexts (client, internal, public) | Appropriate register shift in >4/5 |
|
| 954 |
+
|
| 955 |
+
### 7.5 Phase 4 Gate
|
| 956 |
+
|
| 957 |
+
- [ ] LoRA merged successfully (model generates coherent text)
|
| 958 |
+
- [ ] LASER: OPTIONAL — moved to post-budget local experiment on Margaret <!-- RED-HAT FIX #1: Per C1 review -->
|
| 959 |
+
- [ ] OGPSA personality subspace captured (ogpsa_daimon_v2.json)
|
| 960 |
+
- [ ] Eval battery results saved (without truthfulqa_mc2 — contaminated) <!-- RED-HAT FIX #7: Per H8 review -->
|
| 961 |
+
- [ ] MMLU within 5% of base model (no catastrophic forgetting)
|
| 962 |
+
- [ ] Custom Daimon evals pass criteria
|
| 963 |
+
- [ ] All artifacts synced off-pod to persistent storage <!-- RED-HAT FIX #8: Per H9 review -->
|
| 964 |
+
|
| 965 |
+
---
|
| 966 |
+
|
| 967 |
+
## 8. Phase 5: Export and Validation {#8-phase-5-export-and-validation}
|
| 968 |
+
|
| 969 |
+
**Cost:** ~$2 (1 hour)
|
| 970 |
+
**Purpose:** Convert to serving format and final validation
|
| 971 |
+
|
| 972 |
+
### 8.1 Export to GGUF
|
| 973 |
+
|
| 974 |
+
```bash
|
| 975 |
+
# Clone llama.cpp if not present
|
| 976 |
+
git clone https://github.com/ggerganov/llama.cpp /workspace/llama.cpp
|
| 977 |
+
cd /workspace/llama.cpp && make -j$(nproc)
|
| 978 |
+
|
| 979 |
+
# Convert to GGUF (bf16 first, then quantize)
|
| 980 |
+
# RED-HAT FIX #1: Export from merged model, not LASER'd model (LASER is optional/offline now)
|
| 981 |
+
python convert_hf_to_gguf.py /workspace/models/daimon_v2_merged \
|
| 982 |
+
--outfile /workspace/models/daimon_v2.gguf \
|
| 983 |
+
--outtype bf16
|
| 984 |
+
|
| 985 |
+
# Quantize to Q4_K_M for serving on Margaret (Mac Studio)
|
| 986 |
+
./llama-quantize /workspace/models/daimon_v2.gguf \
|
| 987 |
+
/workspace/models/daimon_v2_Q4_K_M.gguf Q4_K_M
|
| 988 |
+
|
| 989 |
+
# Also export Q5_K_M for higher-quality serving if Margaret has headroom
|
| 990 |
+
./llama-quantize /workspace/models/daimon_v2.gguf \
|
| 991 |
+
/workspace/models/daimon_v2_Q5_K_M.gguf Q5_K_M
|
| 992 |
+
```
|
| 993 |
+
|
| 994 |
+
### 8.2 Export to GPTQ/AWQ (Optional, for GPU serving)
|
| 995 |
+
|
| 996 |
+
```bash
|
| 997 |
+
# If serving on GPU (e.g., via vLLM), export GPTQ
|
| 998 |
+
pip install auto-gptq
|
| 999 |
+
python -c "
|
| 1000 |
+
from transformers import AutoModelForCausalLM, AutoTokenizer
|
| 1001 |
+
from auto_gptq import AutoGPTQForCausalLM, BaseQuantizeConfig
|
| 1002 |
+
|
| 1003 |
+
quantize_config = BaseQuantizeConfig(bits=4, group_size=128, desc_act=False)
|
| 1004 |
+
model = AutoGPTQForCausalLM.from_pretrained(
|
| 1005 |
+
'/workspace/models/daimon_v2_merged', # RED-HAT FIX #1: Use merged model, not LASER'd
|
| 1006 |
+
quantize_config=quantize_config,
|
| 1007 |
+
)
|
| 1008 |
+
model.quantize() # Uses calibration data
|
| 1009 |
+
model.save_quantized('/workspace/models/daimon_v2_gptq')
|
| 1010 |
+
"
|
| 1011 |
+
```
|
| 1012 |
+
|
| 1013 |
+
### 8.3 Download Artifacts
|
| 1014 |
+
|
| 1015 |
+
Before terminating the cloud instance, download all critical artifacts:
|
| 1016 |
+
|
| 1017 |
+
```bash
|
| 1018 |
+
# Priority 1: The trained model (pick your serving format)
|
| 1019 |
+
# GGUF Q4_K_M: ~18 GB
|
| 1020 |
+
# GGUF Q5_K_M: ~22 GB
|
| 1021 |
+
# bf16 merged: ~70 GB (only if you have bandwidth/storage)
|
| 1022 |
+
|
| 1023 |
+
# Priority 2: LoRA checkpoints (for potential continued training)
|
| 1024 |
+
# Foundation best: ~500 MB
|
| 1025 |
+
# Calibration best: ~500 MB
|
| 1026 |
+
|
| 1027 |
+
# Priority 3: Metadata
|
| 1028 |
+
# hot_experts.json
|
| 1029 |
+
# ogpsa_daimon_v2.json
|
| 1030 |
+
# eval_results/ directory
|
| 1031 |
+
# training logs
|
| 1032 |
+
```
|
| 1033 |
+
|
| 1034 |
+
---
|
| 1035 |
+
|
| 1036 |
+
## 9. Budget Allocation {#9-budget-allocation}
|
| 1037 |
+
|
| 1038 |
+
### Primary Run Budget
|
| 1039 |
+
|
| 1040 |
+
<!-- RED-HAT FIX #12: Corrected budget to realistic raw-PEFT throughput per C4 review.
|
| 1041 |
+
Original assumed Unsloth kernel speeds ($20-28 for Phase 2) but shipped code uses
|
| 1042 |
+
raw PEFT + Transformers. Throughput probe (Section 5.5b) gates the actual cost. -->
|
| 1043 |
+
|
| 1044 |
+
| Phase | Estimated Time | Cost | Running Total |
|
| 1045 |
+
|---|---|---|---|
|
| 1046 |
+
| 0: Preflight (incl. backward-pass smoke test) | 30-60 min | $1.50-2.50 | $1.50-2.50 |
|
| 1047 |
+
| 1: Expert profiling | 30 min | $1.00 | $2.50-3.50 |
|
| 1048 |
+
| 2: Foundation training (gated by throughput probe) | 11-16 hrs | $22-32 | $24.50-35.50 |
|
| 1049 |
+
| 3: Thomas voice calibration | 30 min | $1.00 | $25.50-36.50 |
|
| 1050 |
+
| 4: Post-training (OGPSA + eval; LASER cut from critical path) | 1-2 hrs | $2-4 | $27.50-40.50 |
|
| 1051 |
+
| 5: Export and download | 1 hr | $2-3 | $29.50-43.50 |
|
| 1052 |
+
| **Total primary run** | | | **$30-45** |
|
| 1053 |
+
|
| 1054 |
+
### Reserve Budget
|
| 1055 |
+
|
| 1056 |
+
| Use | Estimated Cost |
|
| 1057 |
+
|---|---|
|
| 1058 |
+
| **Available reserve** | **$52-67** |
|
| 1059 |
+
| Hyperparameter sweep run (if voice quality insufficient) | $22-32 |
|
| 1060 |
+
| LR sweep (3 runs at 1e-4, 2e-4, 5e-4) on 10K subset | $6-9 |
|
| 1061 |
+
| Rank sweep (r=16 vs r=32 vs r=64) on 10K subset | $6-9 |
|
| 1062 |
+
| LASER sweep (offline on Margaret, $0) | $0 |
|
| 1063 |
+
| Full re-run with tuned hyperparameters | $22-32 |
|
| 1064 |
+
| Emergency buffer | ~$5-15 |
|
| 1065 |
+
|
| 1066 |
+
### Cost Optimization Tips
|
| 1067 |
+
|
| 1068 |
+
1. **Use spot/interruptible instances** if the provider offers them. Save checkpoints every 2000 steps to survive interruptions.
|
| 1069 |
+
2. **Download the model to persistent storage** before the first run. Model download is ~70 GB and takes 15-30 minutes depending on bandwidth. Do not pay H200 rates for downloading.
|
| 1070 |
+
3. **Kill the instance between runs** if doing multi-day work. $2/hr idle is $48/day wasted.
|
| 1071 |
+
4. **Hyperparameter sweeps on subsets first.** A 10K-sample sweep costs ~$1-2 per run. Do not spend $25 on a full run with untested hyperparameters.
|
| 1072 |
+
|
| 1073 |
+
---
|
| 1074 |
+
|
| 1075 |
+
## 10. Rollback Plan {#10-rollback-plan}
|
| 1076 |
+
|
| 1077 |
+
### Per-Phase Rollback
|
| 1078 |
+
|
| 1079 |
+
| Phase | If it fails... | Rollback action | Cost to recover |
|
| 1080 |
+
|---|---|---|---|
|
| 1081 |
+
| 0: Preflight | Hardware/software mismatch | Fix environment or change provider. No cost wasted (no training started). | $0 |
|
| 1082 |
+
| 1: Expert profiling | Profiling script errors | Debug and re-run. Costs ~$0.50. Or fall back to uniform LoRA (all experts) as Tier 1 baseline. | $0.50 |
|
| 1083 |
+
| 2: Foundation training | Divergence at step N | Roll back to checkpoint at step N-2000. Adjust LR (halve it) and resume from that checkpoint. Resume costs proportional to remaining steps only. | Variable |
|
| 1084 |
+
| 2: Foundation training | OOM/SIGKILL | Reduce batch_size to 2 (effective batch = 8). Or reduce max_seq_length to 1024. Re-run from last checkpoint. | Variable |
|
| 1085 |
+
| 2: Foundation training | Poor quality at completion | Try: (a) different LR, (b) longer training (2 epochs), (c) different data mix, (d) fall back to ESFT instead of HELLoRA. Each iteration is $22-32. | $22-32 | <!-- RED-HAT FIX #12: Corrected from $20-28 -->
|
| 1086 |
+
| 3: Thomas calibration | Voice too weak | Increase epochs to 10, or increase LR to 5e-5. Re-run from Phase 2 best checkpoint. | $1 |
|
| 1087 |
+
| 3: Thomas calibration | Voice too strong (memorized) | Decrease epochs to 3, or decrease LR to 5e-6. Re-run from Phase 2 best checkpoint. | $1 |
|
| 1088 |
+
| 4: LASER | Quality degradation | LASER is optional and offline (run on Margaret). Skip entirely if no improvement measured per-layer. | $0 | <!-- RED-HAT FIX #1 -->
|
| 1089 |
+
| 4: OGPSA | Capture fails | Re-run capture. If architectural mismatch, adapt the OGPSA script for 40-layer Qwen3.6 (v3.1 was written for a 48-layer model). | $0.50 |
|
| 1090 |
+
| 5: Export | GGUF conversion fails | Check llama.cpp supports Qwen3.6 architecture. May need to update llama.cpp to latest. | $0.50 |
|
| 1091 |
+
|
| 1092 |
+
### Complete Pipeline Rollback
|
| 1093 |
+
|
| 1094 |
+
If the entire pipeline produces unsatisfactory results after the primary run:
|
| 1095 |
+
|
| 1096 |
+
1. **Analyze eval results** to identify which capability is deficient (voice, tool use, anti-sycophancy, reasoning).
|
| 1097 |
+
2. **Targeted fix options:**
|
| 1098 |
+
- Voice weak: More Thomas demonstrations + higher LR in Phase 3
|
| 1099 |
+
- Tool use weak: Upsample xlam data to 2x in Phase 2 mix
|
| 1100 |
+
- Anti-sycophancy weak: Upsample TruthfulQA to 25x; consider adding SAA dataset
|
| 1101 |
+
- Reasoning weak: Upsample hermes traces to 5x; consider adding GSM8K training data
|
| 1102 |
+
3. **Full re-run** with adjusted data mix: ~$30-45 from reserve budget. <!-- RED-HAT FIX #12: Corrected from $27-37 -->
|
| 1103 |
+
4. **Method pivot** if HELLoRA ceiling is too low: Switch to ESFT (full fine-tune hot experts) or BAdam (block-coordinate). This is a larger change and costs one full run.
|
| 1104 |
+
|
| 1105 |
+
### Checkpoint Strategy
|
| 1106 |
+
|
| 1107 |
+
Checkpoints are the most important insurance:
|
| 1108 |
+
|
| 1109 |
+
- Phase 2 saves every 2000 steps (5 checkpoints retained)
|
| 1110 |
+
- Phase 3 saves every 20 steps (all checkpoints retained -- tiny files)
|
| 1111 |
+
- The BEST checkpoint from each phase is copied to a separate directory
|
| 1112 |
+
- All checkpoints include the full LoRA adapter state + optimizer state (allows resuming training)
|
| 1113 |
+
- **Checkpoints synced off-pod after every save** (see Section 5.6b) <!-- RED-HAT FIX #8: Per H9 review -->
|
| 1114 |
+
- Download at least the best checkpoints to local storage before terminating the instance
|
| 1115 |
+
|
| 1116 |
+
---
|
| 1117 |
+
|
| 1118 |
+
## Appendix A: Environment Setup Script {#appendix-a}
|
| 1119 |
+
|
| 1120 |
+
```bash
|
| 1121 |
+
#!/bin/bash
|
| 1122 |
+
# daimon_setup.sh -- Run this FIRST on a fresh H200 instance
|
| 1123 |
+
set -euo pipefail
|
| 1124 |
+
|
| 1125 |
+
echo "=== Daimon Training Pipeline: Environment Setup ==="
|
| 1126 |
+
|
| 1127 |
+
# 1. System checks
|
| 1128 |
+
echo "--- Hardware checks ---"
|
| 1129 |
+
nvidia-smi --query-gpu=name,memory.total --format=csv,noheader
|
| 1130 |
+
CONTAINER_RAM=$(cat /sys/fs/cgroup/memory.max 2>/dev/null || cat /sys/fs/cgroup/memory/memory.limit_in_bytes 2>/dev/null)
|
| 1131 |
+
echo "Container RAM: $(echo "$CONTAINER_RAM / 1024 / 1024 / 1024" | bc) GB"
|
| 1132 |
+
|
| 1133 |
+
# 2. Set environment variables BEFORE any Python imports
|
| 1134 |
+
export UNSLOTH_COMPILE_DISABLE=1
|
| 1135 |
+
export PYTORCH_CUDA_ALLOC_CONF="expandable_segments:True"
|
| 1136 |
+
export TOKENIZERS_PARALLELISM=false
|
| 1137 |
+
export UNSLOTH_DISABLE_FAST_GENERATION=1
|
| 1138 |
+
|
| 1139 |
+
# Persist for all subsequent shells
|
| 1140 |
+
cat >> ~/.bashrc << 'ENVEOF'
|
| 1141 |
+
export UNSLOTH_COMPILE_DISABLE=1
|
| 1142 |
+
export PYTORCH_CUDA_ALLOC_CONF="expandable_segments:True"
|
| 1143 |
+
export TOKENIZERS_PARALLELISM=false
|
| 1144 |
+
export UNSLOTH_DISABLE_FAST_GENERATION=1
|
| 1145 |
+
ENVEOF
|
| 1146 |
+
|
| 1147 |
+
# 3. Install/update packages
|
| 1148 |
+
# RED-HAT FIX #9: Pin exact versions and install torch FIRST per H5 review.
|
| 1149 |
+
# Original had version floors (>=) with torch installed LAST, causing resolution chaos.
|
| 1150 |
+
# Install order matters: torch first (pinned to cu124), then core ML stack, unsloth --no-deps last.
|
| 1151 |
+
# lm_eval in separate venv to avoid dependency conflicts.
|
| 1152 |
+
|
| 1153 |
+
# Ensure conda is on PATH (some cloud images need this)
|
| 1154 |
+
export PATH="/opt/conda/bin:$PATH"
|
| 1155 |
+
|
| 1156 |
+
# Set HF cache to /workspace to avoid filling root partition (RED-HAT FIX #5, C5)
|
| 1157 |
+
export HF_HOME=/workspace/.hf
|
| 1158 |
+
export HF_DATASETS_CACHE=/workspace/.hf/datasets
|
| 1159 |
+
mkdir -p "$HF_HOME" "$HF_DATASETS_CACHE"
|
| 1160 |
+
cat >> ~/.bashrc << 'HFEOF'
|
| 1161 |
+
export HF_HOME=/workspace/.hf
|
| 1162 |
+
export HF_DATASETS_CACHE=/workspace/.hf/datasets
|
| 1163 |
+
HFEOF
|
| 1164 |
+
|
| 1165 |
+
pip install --upgrade pip
|
| 1166 |
+
|
| 1167 |
+
# Step 1: Install torch FIRST with CUDA 12.4 (pinned version)
|
| 1168 |
+
pip install torch==2.7.1 torchvision==0.22.1 torchaudio==2.7.1 --index-url https://download.pytorch.org/whl/cu124
|
| 1169 |
+
|
| 1170 |
+
# Step 2: Verify torch installed correctly BEFORE proceeding
|
| 1171 |
+
python -c "import torch; assert torch.cuda.is_available(), 'CUDA not available after torch install'; print(f'torch {torch.__version__} OK')"
|
| 1172 |
+
|
| 1173 |
+
# Step 3: Install core ML stack (pinned versions, no torch re-resolution)
|
| 1174 |
+
pip install transformers==5.0.2
|
| 1175 |
+
pip install peft==0.15.2
|
| 1176 |
+
pip install trl==0.18.1
|
| 1177 |
+
pip install datasets==3.6.0
|
| 1178 |
+
pip install accelerate==1.7.0
|
| 1179 |
+
pip install bitsandbytes==0.46.0 # Required by unsloth even if we don't use 4-bit
|
| 1180 |
+
pip install sentencepiece protobuf
|
| 1181 |
+
|
| 1182 |
+
# Step 4: Install unsloth LAST with --no-deps to avoid overwriting pinned versions
|
| 1183 |
+
pip install "unsloth>=0.1.47" --no-deps
|
| 1184 |
+
|
| 1185 |
+
# Step 5: Verify no dependency conflicts
|
| 1186 |
+
pip check || echo "WARNING: pip check found conflicts — review before proceeding"
|
| 1187 |
+
|
| 1188 |
+
# Step 6: Freeze the resolved environment for reproducibility
|
| 1189 |
+
pip freeze > /workspace/artifacts/requirements.lock
|
| 1190 |
+
echo "Locked environment saved to /workspace/artifacts/requirements.lock"
|
| 1191 |
+
|
| 1192 |
+
# Step 7: Install lm_eval in a SEPARATE venv to avoid conflicts (RED-HAT FIX #9, H5)
|
| 1193 |
+
python -m venv /workspace/lm_eval_venv
|
| 1194 |
+
/workspace/lm_eval_venv/bin/pip install --upgrade pip
|
| 1195 |
+
/workspace/lm_eval_venv/bin/pip install "lm_eval[hf]"
|
| 1196 |
+
|
| 1197 |
+
# 4. Verify installations
|
| 1198 |
+
python -c "
|
| 1199 |
+
import os
|
| 1200 |
+
os.environ['UNSLOTH_COMPILE_DISABLE'] = '1'
|
| 1201 |
+
os.environ['PYTORCH_CUDA_ALLOC_CONF'] = 'expandable_segments:True'
|
| 1202 |
+
os.environ['TOKENIZERS_PARALLELISM'] = 'false'
|
| 1203 |
+
os.environ['UNSLOTH_DISABLE_FAST_GENERATION'] = '1'
|
| 1204 |
+
|
| 1205 |
+
import torch
|
| 1206 |
+
print(f'PyTorch {torch.__version__}, CUDA available: {torch.cuda.is_available()}')
|
| 1207 |
+
if torch.cuda.is_available():
|
| 1208 |
+
print(f'GPU: {torch.cuda.get_device_name(0)}')
|
| 1209 |
+
print(f'VRAM: {torch.cuda.get_device_properties(0).total_memory / 1e9:.1f} GB')
|
| 1210 |
+
|
| 1211 |
+
import transformers; print(f'Transformers {transformers.__version__}')
|
| 1212 |
+
import peft; print(f'PEFT {peft.__version__}')
|
| 1213 |
+
import unsloth; print(f'Unsloth {unsloth.__version__}')
|
| 1214 |
+
print('All imports successful.')
|
| 1215 |
+
"
|
| 1216 |
+
|
| 1217 |
+
# 5. Create workspace directories
|
| 1218 |
+
mkdir -p /workspace/models
|
| 1219 |
+
mkdir -p /workspace/data
|
| 1220 |
+
mkdir -p /workspace/checkpoints
|
| 1221 |
+
mkdir -p /workspace/eval_results
|
| 1222 |
+
mkdir -p /workspace/artifacts
|
| 1223 |
+
|
| 1224 |
+
echo "=== Setup complete ==="
|
| 1225 |
+
```
|
| 1226 |
+
|
| 1227 |
+
---
|
| 1228 |
+
|
| 1229 |
+
## Appendix B: Expert Profiling Script {#appendix-b}
|
| 1230 |
+
|
| 1231 |
+
```python
|
| 1232 |
+
#!/usr/bin/env python3
|
| 1233 |
+
"""
|
| 1234 |
+
daimon_expert_profiling.py
|
| 1235 |
+
Phase 1: Profile which experts activate for Daimon's training data.
|
| 1236 |
+
Uses ESFT-style profiling with MAN (Mean Activation Norm) scoring.
|
| 1237 |
+
|
| 1238 |
+
Reference: ESFT (arXiv:2407.01906), Unified Expert Scoring (arXiv:2606.15716)
|
| 1239 |
+
"""
|
| 1240 |
+
|
| 1241 |
+
import os
|
| 1242 |
+
os.environ["UNSLOTH_COMPILE_DISABLE"] = "1"
|
| 1243 |
+
os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "expandable_segments:True"
|
| 1244 |
+
os.environ["TOKENIZERS_PARALLELISM"] = "false"
|
| 1245 |
+
os.environ["UNSLOTH_DISABLE_FAST_GENERATION"] = "1"
|
| 1246 |
+
|
| 1247 |
+
import json
|
| 1248 |
+
import torch
|
| 1249 |
+
import numpy as np
|
| 1250 |
+
from collections import defaultdict
|
| 1251 |
+
from transformers import AutoModelForCausalLM, AutoTokenizer
|
| 1252 |
+
|
| 1253 |
+
MODEL_PATH = "/workspace/models/Qwen3.6-35B-A3B"
|
| 1254 |
+
OUTPUT_PATH = "/workspace/artifacts/hot_experts.json"
|
| 1255 |
+
NUM_LAYERS = 40
|
| 1256 |
+
NUM_EXPERTS = 256
|
| 1257 |
+
TOP_K = 8 # Qwen3.6 routes to top-8 experts
|
| 1258 |
+
|
| 1259 |
+
# Representative samples from each data source
|
| 1260 |
+
# Adjust paths to match your data layout
|
| 1261 |
+
SAMPLE_SOURCES = {
|
| 1262 |
+
"sonnet_voice": ("/workspace/data/sonnet_voice_distill/", 8),
|
| 1263 |
+
"prosocial": ("/workspace/data/prosocial_dialogue/", 8),
|
| 1264 |
+
"tool_use": ("/workspace/data/xlam_function_calling/", 8),
|
| 1265 |
+
"reasoning": ("/workspace/data/hermes_agent_traces/", 4),
|
| 1266 |
+
"thomas_voice": ("/workspace/data/thomas_voice_demos/daimon_voice_demonstrations.jsonl", 4),
|
| 1267 |
+
}
|
| 1268 |
+
|
| 1269 |
+
def load_samples():
|
| 1270 |
+
"""Load representative samples from each data source."""
|
| 1271 |
+
samples = []
|
| 1272 |
+
for source_name, (path, count) in SAMPLE_SOURCES.items():
|
| 1273 |
+
# Load first `count` examples from each source
|
| 1274 |
+
# Adapt this loader to match your data format
|
| 1275 |
+
if path.endswith(".jsonl"):
|
| 1276 |
+
with open(path) as f:
|
| 1277 |
+
for i, line in enumerate(f):
|
| 1278 |
+
if i >= count:
|
| 1279 |
+
break
|
| 1280 |
+
data = json.loads(line)
|
| 1281 |
+
# Extract the full conversation as a single string
|
| 1282 |
+
text = " ".join(m["content"] for m in data["messages"])
|
| 1283 |
+
samples.append(text)
|
| 1284 |
+
else:
|
| 1285 |
+
# Directory of files -- load from first file found
|
| 1286 |
+
import glob
|
| 1287 |
+
files = sorted(glob.glob(os.path.join(path, "*.jsonl")))
|
| 1288 |
+
if not files:
|
| 1289 |
+
files = sorted(glob.glob(os.path.join(path, "*.json")))
|
| 1290 |
+
if files:
|
| 1291 |
+
loaded = 0
|
| 1292 |
+
for fpath in files:
|
| 1293 |
+
with open(fpath) as f:
|
| 1294 |
+
for line in f:
|
| 1295 |
+
if loaded >= count:
|
| 1296 |
+
break
|
| 1297 |
+
data = json.loads(line)
|
| 1298 |
+
if "messages" in data:
|
| 1299 |
+
text = " ".join(m["content"] for m in data["messages"])
|
| 1300 |
+
elif "text" in data:
|
| 1301 |
+
text = data["text"]
|
| 1302 |
+
elif "conversations" in data:
|
| 1303 |
+
text = " ".join(c.get("value", "") for c in data["conversations"])
|
| 1304 |
+
else:
|
| 1305 |
+
text = str(data)
|
| 1306 |
+
samples.append(text)
|
| 1307 |
+
loaded += 1
|
| 1308 |
+
if loaded >= count:
|
| 1309 |
+
break
|
| 1310 |
+
print(f"Loaded {len(samples)} profiling samples")
|
| 1311 |
+
return samples
|
| 1312 |
+
|
| 1313 |
+
|
| 1314 |
+
def profile_experts(model, tokenizer, samples):
|
| 1315 |
+
"""
|
| 1316 |
+
Profile expert activation patterns using Mean Activation Norm (MAN).
|
| 1317 |
+
|
| 1318 |
+
For each MoE layer, record the activation norm of each expert's output
|
| 1319 |
+
when it is selected by the router. The Mean Activation Norm across all
|
| 1320 |
+
samples gives the expert importance score.
|
| 1321 |
+
"""
|
| 1322 |
+
# Storage: layer_idx -> expert_idx -> list of activation norms
|
| 1323 |
+
activation_norms = defaultdict(lambda: defaultdict(list))
|
| 1324 |
+
|
| 1325 |
+
# Register hooks on MoE layers
|
| 1326 |
+
hooks = []
|
| 1327 |
+
|
| 1328 |
+
def make_hook(layer_idx):
|
| 1329 |
+
def hook_fn(module, input, output):
|
| 1330 |
+
# The MoE module's output includes routing information
|
| 1331 |
+
# We need to capture which experts were selected and their output norms
|
| 1332 |
+
# This hook structure depends on the exact Qwen3.6 implementation
|
| 1333 |
+
#
|
| 1334 |
+
# For Qwen3.6, the MoE block routes tokens and returns the weighted sum.
|
| 1335 |
+
# To get per-expert activation norms, we need to hook deeper -- at the
|
| 1336 |
+
# router level and individual expert level.
|
| 1337 |
+
pass
|
| 1338 |
+
return hook_fn
|
| 1339 |
+
|
| 1340 |
+
# Alternative approach: direct router inspection
|
| 1341 |
+
# Hook the router to get expert selection, then measure expert output norms
|
| 1342 |
+
for layer_idx in range(NUM_LAYERS):
|
| 1343 |
+
moe_layer = model.model.layers[layer_idx].mlp
|
| 1344 |
+
|
| 1345 |
+
# Hook the gate (router) to capture routing decisions
|
| 1346 |
+
def make_router_hook(l_idx):
|
| 1347 |
+
def hook_fn(module, input, output):
|
| 1348 |
+
# output is the router logits or routing weights
|
| 1349 |
+
# For Qwen3.6: output shape is [batch*seq_len, num_experts]
|
| 1350 |
+
if isinstance(output, tuple):
|
| 1351 |
+
router_logits = output[0]
|
| 1352 |
+
else:
|
| 1353 |
+
router_logits = output
|
| 1354 |
+
|
| 1355 |
+
# Get top-k expert indices
|
| 1356 |
+
topk_vals, topk_indices = torch.topk(router_logits, TOP_K, dim=-1)
|
| 1357 |
+
|
| 1358 |
+
# Record which experts were selected and their gate values
|
| 1359 |
+
for expert_idx in range(NUM_EXPERTS):
|
| 1360 |
+
mask = (topk_indices == expert_idx).any(dim=-1)
|
| 1361 |
+
if mask.any():
|
| 1362 |
+
# Use gate value as proxy for activation norm
|
| 1363 |
+
# (actual activation norm requires hooking each expert separately)
|
| 1364 |
+
gate_vals = topk_vals[mask]
|
| 1365 |
+
mean_gate = gate_vals.mean().item()
|
| 1366 |
+
activation_norms[l_idx][expert_idx].append(mean_gate)
|
| 1367 |
+
|
| 1368 |
+
return hook_fn
|
| 1369 |
+
|
| 1370 |
+
hook = moe_layer.gate.register_forward_hook(make_router_hook(layer_idx))
|
| 1371 |
+
hooks.append(hook)
|
| 1372 |
+
|
| 1373 |
+
# Run inference on all samples
|
| 1374 |
+
model.eval()
|
| 1375 |
+
with torch.no_grad():
|
| 1376 |
+
for i, text in enumerate(samples):
|
| 1377 |
+
inputs = tokenizer(
|
| 1378 |
+
text, return_tensors="pt", truncation=True, max_length=2048
|
| 1379 |
+
).to("cuda")
|
| 1380 |
+
_ = model(**inputs)
|
| 1381 |
+
if i % 8 == 0:
|
| 1382 |
+
print(f"Profiled sample {i+1}/{len(samples)}")
|
| 1383 |
+
|
| 1384 |
+
# Remove hooks
|
| 1385 |
+
for hook in hooks:
|
| 1386 |
+
hook.remove()
|
| 1387 |
+
|
| 1388 |
+
return activation_norms
|
| 1389 |
+
|
| 1390 |
+
|
| 1391 |
+
def compute_hot_experts(activation_norms, coverage_threshold=0.80):
|
| 1392 |
+
"""
|
| 1393 |
+
Compute hot experts per layer using Mean Activation Norm.
|
| 1394 |
+
|
| 1395 |
+
For each layer, sort experts by MAN score and select the minimum set
|
| 1396 |
+
that covers >= coverage_threshold of total activation norm.
|
| 1397 |
+
"""
|
| 1398 |
+
hot_experts = {}
|
| 1399 |
+
|
| 1400 |
+
for layer_idx in range(NUM_LAYERS):
|
| 1401 |
+
expert_scores = {}
|
| 1402 |
+
for expert_idx in range(NUM_EXPERTS):
|
| 1403 |
+
norms = activation_norms[layer_idx].get(expert_idx, [])
|
| 1404 |
+
if norms:
|
| 1405 |
+
expert_scores[expert_idx] = np.mean(norms)
|
| 1406 |
+
else:
|
| 1407 |
+
expert_scores[expert_idx] = 0.0
|
| 1408 |
+
|
| 1409 |
+
# Sort by MAN score descending
|
| 1410 |
+
sorted_experts = sorted(expert_scores.items(), key=lambda x: x[1], reverse=True)
|
| 1411 |
+
total_norm = sum(score for _, score in sorted_experts)
|
| 1412 |
+
|
| 1413 |
+
if total_norm == 0:
|
| 1414 |
+
print(f"WARNING: Layer {layer_idx} has zero total activation norm!")
|
| 1415 |
+
hot_experts[layer_idx] = list(range(TOP_K)) # Fallback
|
| 1416 |
+
continue
|
| 1417 |
+
|
| 1418 |
+
# Select experts until coverage threshold is met
|
| 1419 |
+
cumulative = 0.0
|
| 1420 |
+
hot = []
|
| 1421 |
+
for expert_idx, score in sorted_experts:
|
| 1422 |
+
hot.append(expert_idx)
|
| 1423 |
+
cumulative += score
|
| 1424 |
+
if cumulative / total_norm >= coverage_threshold:
|
| 1425 |
+
break
|
| 1426 |
+
|
| 1427 |
+
hot_experts[layer_idx] = sorted(hot)
|
| 1428 |
+
print(f"Layer {layer_idx}: {len(hot)} hot experts "
|
| 1429 |
+
f"({100*len(hot)/NUM_EXPERTS:.1f}% of 256, "
|
| 1430 |
+
f"covering {100*cumulative/total_norm:.1f}% of activation norm)")
|
| 1431 |
+
|
| 1432 |
+
return hot_experts
|
| 1433 |
+
|
| 1434 |
+
|
| 1435 |
+
def main():
|
| 1436 |
+
print("=== Phase 1: Expert Profiling ===")
|
| 1437 |
+
|
| 1438 |
+
# Load model for inference only
|
| 1439 |
+
print("Loading model...")
|
| 1440 |
+
tokenizer = AutoTokenizer.from_pretrained(MODEL_PATH)
|
| 1441 |
+
model = AutoModelForCausalLM.from_pretrained(
|
| 1442 |
+
MODEL_PATH,
|
| 1443 |
+
torch_dtype=torch.bfloat16,
|
| 1444 |
+
device_map="cuda",
|
| 1445 |
+
)
|
| 1446 |
+
|
| 1447 |
+
# Load profiling samples
|
| 1448 |
+
samples = load_samples()
|
| 1449 |
+
|
| 1450 |
+
# Profile
|
| 1451 |
+
print("Profiling expert activations...")
|
| 1452 |
+
activation_norms = profile_experts(model, tokenizer, samples)
|
| 1453 |
+
|
| 1454 |
+
# Compute hot experts
|
| 1455 |
+
print("\nComputing hot experts (80% coverage threshold)...")
|
| 1456 |
+
hot_experts = compute_hot_experts(activation_norms, coverage_threshold=0.80)
|
| 1457 |
+
|
| 1458 |
+
# Summary statistics
|
| 1459 |
+
counts = [len(v) for v in hot_experts.values()]
|
| 1460 |
+
print(f"\nHot experts per layer: min={min(counts)}, max={max(counts)}, "
|
| 1461 |
+
f"mean={np.mean(counts):.1f}, median={np.median(counts):.1f}")
|
| 1462 |
+
|
| 1463 |
+
total_hot = sum(counts)
|
| 1464 |
+
total_possible = NUM_LAYERS * NUM_EXPERTS
|
| 1465 |
+
print(f"Total hot expert slots: {total_hot}/{total_possible} "
|
| 1466 |
+
f"({100*total_hot/total_possible:.1f}%)")
|
| 1467 |
+
|
| 1468 |
+
# Save
|
| 1469 |
+
# Convert int keys to strings for JSON
|
| 1470 |
+
hot_experts_json = {str(k): v for k, v in hot_experts.items()}
|
| 1471 |
+
with open(OUTPUT_PATH, "w") as f:
|
| 1472 |
+
json.dump(hot_experts_json, f, indent=2)
|
| 1473 |
+
print(f"\nSaved to {OUTPUT_PATH}")
|
| 1474 |
+
|
| 1475 |
+
# Sanity checks
|
| 1476 |
+
for layer_idx, experts in hot_experts.items():
|
| 1477 |
+
if len(experts) < 5:
|
| 1478 |
+
print(f"WARNING: Layer {layer_idx} has only {len(experts)} hot experts. "
|
| 1479 |
+
"Consider lowering coverage threshold.")
|
| 1480 |
+
if len(experts) > 100:
|
| 1481 |
+
print(f"WARNING: Layer {layer_idx} has {len(experts)} hot experts. "
|
| 1482 |
+
"Routing may be too diffuse. Check for profiling errors.")
|
| 1483 |
+
|
| 1484 |
+
del model
|
| 1485 |
+
torch.cuda.empty_cache()
|
| 1486 |
+
print("\n=== Expert profiling complete ===")
|
| 1487 |
+
|
| 1488 |
+
|
| 1489 |
+
if __name__ == "__main__":
|
| 1490 |
+
main()
|
| 1491 |
+
```
|
| 1492 |
+
|
| 1493 |
+
**Important notes on the profiling script:**
|
| 1494 |
+
|
| 1495 |
+
1. The router hook structure assumes Qwen3.6's MoE implementation exposes router logits through the gate module. The exact module hierarchy may differ -- inspect the model's `named_modules()` output to find the correct hook points.
|
| 1496 |
+
2. Using gate values as a proxy for activation norms is an approximation. For exact MAN scoring per arXiv:2606.15716, hook each expert's output layer and measure the L2 norm of the output tensor. This requires more memory but gives more accurate scores.
|
| 1497 |
+
3. The 80% coverage threshold is the starting point. If it selects >50 experts per layer, raise to 85%. If it selects <15, lower to 75%.
|
| 1498 |
+
|
| 1499 |
+
---
|
| 1500 |
+
|
| 1501 |
+
## Appendix C: Training Script {#appendix-c}
|
| 1502 |
+
|
| 1503 |
+
```python
|
| 1504 |
+
#!/usr/bin/env python3
|
| 1505 |
+
"""
|
| 1506 |
+
daimon_train.py
|
| 1507 |
+
Phase 2 + Phase 3: Foundation training + Thomas voice calibration.
|
| 1508 |
+
|
| 1509 |
+
Usage:
|
| 1510 |
+
python daimon_train.py --phase foundation
|
| 1511 |
+
python daimon_train.py --phase calibration --checkpoint /path/to/best
|
| 1512 |
+
"""
|
| 1513 |
+
|
| 1514 |
+
import os
|
| 1515 |
+
os.environ["UNSLOTH_COMPILE_DISABLE"] = "1"
|
| 1516 |
+
os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "expandable_segments:True"
|
| 1517 |
+
os.environ["TOKENIZERS_PARALLELISM"] = "false"
|
| 1518 |
+
os.environ["UNSLOTH_DISABLE_FAST_GENERATION"] = "1"
|
| 1519 |
+
|
| 1520 |
+
import json
|
| 1521 |
+
import argparse
|
| 1522 |
+
import torch
|
| 1523 |
+
from datasets import load_dataset, interleave_datasets, Dataset
|
| 1524 |
+
from peft import LoraConfig, get_peft_model, PeftModel
|
| 1525 |
+
from transformers import AutoModelForCausalLM, AutoTokenizer
|
| 1526 |
+
from trl import SFTTrainer, SFTConfig
|
| 1527 |
+
|
| 1528 |
+
|
| 1529 |
+
def build_target_modules(hot_experts_path):
|
| 1530 |
+
"""Build HELLoRA target module list from expert profiling results."""
|
| 1531 |
+
with open(hot_experts_path) as f:
|
| 1532 |
+
hot_experts = json.load(f)
|
| 1533 |
+
|
| 1534 |
+
target_modules = []
|
| 1535 |
+
|
| 1536 |
+
# Attention projections (all layers)
|
| 1537 |
+
for layer_idx in range(40):
|
| 1538 |
+
for proj in ["q_proj", "k_proj", "v_proj", "o_proj"]:
|
| 1539 |
+
target_modules.append(f"model.layers.{layer_idx}.self_attn.{proj}")
|
| 1540 |
+
|
| 1541 |
+
# Hot expert FFN layers only
|
| 1542 |
+
for layer_idx_str, expert_indices in hot_experts.items():
|
| 1543 |
+
layer_idx = int(layer_idx_str)
|
| 1544 |
+
for expert_idx in expert_indices:
|
| 1545 |
+
for proj in ["gate_proj", "up_proj", "down_proj"]:
|
| 1546 |
+
target_modules.append(
|
| 1547 |
+
f"model.layers.{layer_idx}.mlp.experts.{expert_idx}.{proj}"
|
| 1548 |
+
)
|
| 1549 |
+
|
| 1550 |
+
return target_modules
|
| 1551 |
+
|
| 1552 |
+
|
| 1553 |
+
def build_foundation_dataset(tokenizer):
|
| 1554 |
+
"""Build the mixed foundation dataset with appropriate sampling.
|
| 1555 |
+
|
| 1556 |
+
RED-HAT FIX #3: Returns (train_dataset, eval_dataset) — stratified eval split
|
| 1557 |
+
of ~300 examples held out before interleaving. (Per C3 review)
|
| 1558 |
+
"""
|
| 1559 |
+
# Load each dataset source
|
| 1560 |
+
# NOTE: Adapt these loaders to match your actual data format and paths
|
| 1561 |
+
|
| 1562 |
+
# RED-HAT FIX #3: Hold out eval examples BEFORE interleaving.
|
| 1563 |
+
# ~50 from each major source, fewer from small sources, fixed seed.
|
| 1564 |
+
EVAL_COUNTS = {
|
| 1565 |
+
"sonnet": 80, "prosocial": 50, "xlam": 50,
|
| 1566 |
+
"hermes": 30, "code": 30, "truthful": 60
|
| 1567 |
+
}
|
| 1568 |
+
|
| 1569 |
+
datasets_with_weights = []
|
| 1570 |
+
eval_datasets = []
|
| 1571 |
+
|
| 1572 |
+
# 1. Sonnet voice distill (226K, weight 0.44)
|
| 1573 |
+
sonnet_ds = load_dataset("json", data_files="/workspace/data/sonnet_voice_distill/*.jsonl",
|
| 1574 |
+
split="train").shuffle(seed=42)
|
| 1575 |
+
eval_datasets.append(sonnet_ds.select(range(EVAL_COUNTS["sonnet"])))
|
| 1576 |
+
sonnet_ds = sonnet_ds.select(range(EVAL_COUNTS["sonnet"], len(sonnet_ds)))
|
| 1577 |
+
datasets_with_weights.append((sonnet_ds, 0.44))
|
| 1578 |
+
|
| 1579 |
+
# 2. Prosocial dialogue (165K subsampled to 49.5K, weight 0.10)
|
| 1580 |
+
prosocial_ds = load_dataset("json", data_files="/workspace/data/prosocial_dialogue/*.jsonl",
|
| 1581 |
+
split="train").shuffle(seed=42)
|
| 1582 |
+
eval_datasets.append(prosocial_ds.select(range(EVAL_COUNTS["prosocial"])))
|
| 1583 |
+
prosocial_ds = prosocial_ds.select(range(EVAL_COUNTS["prosocial"], len(prosocial_ds)))
|
| 1584 |
+
prosocial_ds = prosocial_ds.select(range(min(49500, len(prosocial_ds))))
|
| 1585 |
+
datasets_with_weights.append((prosocial_ds, 0.10))
|
| 1586 |
+
|
| 1587 |
+
# 3. xlam function calling (60K, weight 0.12)
|
| 1588 |
+
xlam_ds = load_dataset("json", data_files="/workspace/data/xlam_function_calling/*.jsonl",
|
| 1589 |
+
split="train").shuffle(seed=42)
|
| 1590 |
+
eval_datasets.append(xlam_ds.select(range(EVAL_COUNTS["xlam"])))
|
| 1591 |
+
xlam_ds = xlam_ds.select(range(EVAL_COUNTS["xlam"], len(xlam_ds)))
|
| 1592 |
+
datasets_with_weights.append((xlam_ds, 0.12))
|
| 1593 |
+
|
| 1594 |
+
# 4. Hermes agent traces (7.6K upsampled 3x, weight 0.04)
|
| 1595 |
+
hermes_ds = load_dataset("json", data_files="/workspace/data/hermes_agent_traces/*.jsonl",
|
| 1596 |
+
split="train").shuffle(seed=42)
|
| 1597 |
+
eval_datasets.append(hermes_ds.select(range(EVAL_COUNTS["hermes"])))
|
| 1598 |
+
hermes_ds = hermes_ds.select(range(EVAL_COUNTS["hermes"], len(hermes_ds)))
|
| 1599 |
+
hermes_upsampled = Dataset.from_dict({
|
| 1600 |
+
k: hermes_ds[k] * 3 for k in hermes_ds.column_names
|
| 1601 |
+
})
|
| 1602 |
+
datasets_with_weights.append((hermes_upsampled, 0.04))
|
| 1603 |
+
|
| 1604 |
+
# 5. Code review feedback (9.5K upsampled 2x, weight 0.04)
|
| 1605 |
+
code_ds = load_dataset("json", data_files="/workspace/data/code_review_feedback/*.jsonl",
|
| 1606 |
+
split="train").shuffle(seed=42)
|
| 1607 |
+
eval_datasets.append(code_ds.select(range(EVAL_COUNTS["code"])))
|
| 1608 |
+
code_ds = code_ds.select(range(EVAL_COUNTS["code"], len(code_ds)))
|
| 1609 |
+
code_upsampled = Dataset.from_dict({
|
| 1610 |
+
k: code_ds[k] * 2 for k in code_ds.column_names
|
| 1611 |
+
})
|
| 1612 |
+
datasets_with_weights.append((code_upsampled, 0.04))
|
| 1613 |
+
|
| 1614 |
+
# 6. TruthfulQA (817 upsampled 15x, weight 0.02)
|
| 1615 |
+
truthful_ds = load_dataset("json", data_files="/workspace/data/truthfulqa/*.jsonl",
|
| 1616 |
+
split="train").shuffle(seed=42)
|
| 1617 |
+
eval_datasets.append(truthful_ds.select(range(EVAL_COUNTS["truthful"])))
|
| 1618 |
+
truthful_ds = truthful_ds.select(range(EVAL_COUNTS["truthful"], len(truthful_ds)))
|
| 1619 |
+
truthful_upsampled = Dataset.from_dict({
|
| 1620 |
+
k: truthful_ds[k] * 15 for k in truthful_ds.column_names
|
| 1621 |
+
})
|
| 1622 |
+
datasets_with_weights.append((truthful_upsampled, 0.02))
|
| 1623 |
+
|
| 1624 |
+
# Interleave training data with weights
|
| 1625 |
+
all_datasets = [d for d, w in datasets_with_weights]
|
| 1626 |
+
all_weights = [w for d, w in datasets_with_weights]
|
| 1627 |
+
|
| 1628 |
+
# Normalize weights
|
| 1629 |
+
total_weight = sum(all_weights)
|
| 1630 |
+
probabilities = [w / total_weight for w in all_weights]
|
| 1631 |
+
|
| 1632 |
+
mixed_dataset = interleave_datasets(
|
| 1633 |
+
all_datasets,
|
| 1634 |
+
probabilities=probabilities,
|
| 1635 |
+
seed=42,
|
| 1636 |
+
stopping_strategy="all_exhausted",
|
| 1637 |
+
)
|
| 1638 |
+
|
| 1639 |
+
# RED-HAT FIX #3: Combine eval splits into a single eval dataset
|
| 1640 |
+
from datasets import concatenate_datasets
|
| 1641 |
+
eval_dataset = concatenate_datasets(eval_datasets).shuffle(seed=42)
|
| 1642 |
+
|
| 1643 |
+
print(f"Foundation train dataset: {len(mixed_dataset)} examples")
|
| 1644 |
+
print(f"Foundation eval dataset: {len(eval_dataset)} examples")
|
| 1645 |
+
return mixed_dataset, eval_dataset
|
| 1646 |
+
|
| 1647 |
+
|
| 1648 |
+
def train_foundation(model_path, hot_experts_path):
|
| 1649 |
+
"""Phase 2: Foundation training."""
|
| 1650 |
+
print("=== Phase 2: Foundation Training ===")
|
| 1651 |
+
|
| 1652 |
+
# Load model
|
| 1653 |
+
tokenizer = AutoTokenizer.from_pretrained(model_path)
|
| 1654 |
+
model = AutoModelForCausalLM.from_pretrained(
|
| 1655 |
+
model_path,
|
| 1656 |
+
torch_dtype=torch.bfloat16,
|
| 1657 |
+
device_map="cuda",
|
| 1658 |
+
)
|
| 1659 |
+
|
| 1660 |
+
# Apply HELLoRA
|
| 1661 |
+
target_modules = build_target_modules(hot_experts_path)
|
| 1662 |
+
print(f"Target modules: {len(target_modules)}")
|
| 1663 |
+
|
| 1664 |
+
lora_config = LoraConfig(
|
| 1665 |
+
r=32,
|
| 1666 |
+
lora_alpha=64,
|
| 1667 |
+
lora_dropout=0.05,
|
| 1668 |
+
target_modules=target_modules,
|
| 1669 |
+
use_dora=True,
|
| 1670 |
+
bias="none",
|
| 1671 |
+
task_type="CAUSAL_LM",
|
| 1672 |
+
)
|
| 1673 |
+
|
| 1674 |
+
model = get_peft_model(model, lora_config)
|
| 1675 |
+
model.print_trainable_parameters()
|
| 1676 |
+
|
| 1677 |
+
# Verify expert LoRA adapters exist
|
| 1678 |
+
expert_lora_count = sum(
|
| 1679 |
+
1 for name, p in model.named_parameters()
|
| 1680 |
+
if "experts" in name and "lora" in name and p.requires_grad
|
| 1681 |
+
)
|
| 1682 |
+
assert expert_lora_count > 0, "FATAL: No LoRA adapters on expert layers!"
|
| 1683 |
+
print(f"Expert LoRA parameters: {expert_lora_count}")
|
| 1684 |
+
|
| 1685 |
+
# Build dataset
|
| 1686 |
+
# RED-HAT FIX #3: build_foundation_dataset now returns (train_dataset, eval_dataset)
|
| 1687 |
+
# Eval split: ~300 examples stratified across all 6 sources, fixed seed, excluded from training.
|
| 1688 |
+
train_dataset, eval_dataset = build_foundation_dataset(tokenizer)
|
| 1689 |
+
|
| 1690 |
+
# Training config
|
| 1691 |
+
training_args = SFTConfig(
|
| 1692 |
+
output_dir="/workspace/checkpoints/daimon_v2_foundation",
|
| 1693 |
+
per_device_train_batch_size=4,
|
| 1694 |
+
gradient_accumulation_steps=4,
|
| 1695 |
+
num_train_epochs=1,
|
| 1696 |
+
learning_rate=2e-4,
|
| 1697 |
+
lr_scheduler_type="cosine",
|
| 1698 |
+
warmup_ratio=0.03,
|
| 1699 |
+
weight_decay=0.01,
|
| 1700 |
+
bf16=True,
|
| 1701 |
+
max_seq_length=2048,
|
| 1702 |
+
logging_steps=50,
|
| 1703 |
+
save_steps=2000,
|
| 1704 |
+
save_total_limit=5,
|
| 1705 |
+
dataloader_num_workers=0,
|
| 1706 |
+
dataset_num_proc=1,
|
| 1707 |
+
gradient_checkpointing=True,
|
| 1708 |
+
gradient_checkpointing_kwargs={"use_reentrant": False},
|
| 1709 |
+
eval_strategy="steps",
|
| 1710 |
+
eval_steps=2000,
|
| 1711 |
+
load_best_model_at_end=True, # RED-HAT FIX #3: Select best checkpoint by eval_loss
|
| 1712 |
+
metric_for_best_model="eval_loss", # RED-HAT FIX #3
|
| 1713 |
+
max_steps=30000, # RED-HAT FIX #4: Cost circuit breaker (adjusted by throughput probe)
|
| 1714 |
+
report_to="none",
|
| 1715 |
+
seed=42,
|
| 1716 |
+
)
|
| 1717 |
+
|
| 1718 |
+
# Train
|
| 1719 |
+
trainer = SFTTrainer(
|
| 1720 |
+
model=model,
|
| 1721 |
+
args=training_args,
|
| 1722 |
+
train_dataset=train_dataset, # RED-HAT FIX #3: Was `dataset`, now split
|
| 1723 |
+
eval_dataset=eval_dataset, # RED-HAT FIX #3: Added — prevents eval crash
|
| 1724 |
+
processing_class=tokenizer,
|
| 1725 |
+
)
|
| 1726 |
+
|
| 1727 |
+
trainer.train()
|
| 1728 |
+
|
| 1729 |
+
# Save best checkpoint
|
| 1730 |
+
model.save_pretrained("/workspace/checkpoints/daimon_v2_foundation_best")
|
| 1731 |
+
tokenizer.save_pretrained("/workspace/checkpoints/daimon_v2_foundation_best")
|
| 1732 |
+
print("=== Foundation training complete ===")
|
| 1733 |
+
|
| 1734 |
+
|
| 1735 |
+
def train_calibration(checkpoint_path, base_model_path="/workspace/models/Qwen3.6-35B-A3B"):
|
| 1736 |
+
"""Phase 3: Thomas voice calibration.
|
| 1737 |
+
|
| 1738 |
+
RED-HAT FIX #2: Phase 2 saves adapter-only (adapter_config.json + adapter_model.safetensors),
|
| 1739 |
+
NOT a full model. Must load base model first, then attach adapter with is_trainable=True.
|
| 1740 |
+
Original code used AutoModelForCausalLM.from_pretrained on the adapter dir, which either
|
| 1741 |
+
crashes ("no config.json") or loads adapter for inference-only (requires_grad=False),
|
| 1742 |
+
causing silent no-op training — the exact mlx-lm #571 failure class. (Per C2 review)
|
| 1743 |
+
"""
|
| 1744 |
+
print("=== Phase 3: Thomas Voice Calibration ===")
|
| 1745 |
+
|
| 1746 |
+
# RED-HAT FIX #2: Load BASE model first, then attach Phase 2 adapter
|
| 1747 |
+
tokenizer = AutoTokenizer.from_pretrained(base_model_path)
|
| 1748 |
+
base_model = AutoModelForCausalLM.from_pretrained(
|
| 1749 |
+
base_model_path, # Base model path, NOT the checkpoint
|
| 1750 |
+
torch_dtype=torch.bfloat16,
|
| 1751 |
+
device_map="cuda",
|
| 1752 |
+
)
|
| 1753 |
+
|
| 1754 |
+
# Attach Phase 2 adapter with is_trainable=True
|
| 1755 |
+
model = PeftModel.from_pretrained(
|
| 1756 |
+
base_model,
|
| 1757 |
+
checkpoint_path, # Phase 2 adapter checkpoint
|
| 1758 |
+
is_trainable=True, # CRITICAL: without this, all adapter params are frozen
|
| 1759 |
+
)
|
| 1760 |
+
|
| 1761 |
+
# RED-HAT FIX #2: Verify adapter is actually trainable (catches silent no-op)
|
| 1762 |
+
trainable_count = sum(p.numel() for p in model.parameters() if p.requires_grad)
|
| 1763 |
+
assert trainable_count > 0, "FATAL: No trainable parameters in Phase 3! Check is_trainable=True."
|
| 1764 |
+
expert_lora_count = sum(
|
| 1765 |
+
1 for name, p in model.named_parameters()
|
| 1766 |
+
if "experts" in name and "lora" in name and p.requires_grad
|
| 1767 |
+
)
|
| 1768 |
+
assert expert_lora_count > 0, "FATAL: No LoRA adapters on expert layers in Phase 3!"
|
| 1769 |
+
print(f"Phase 3 trainable params: {trainable_count:,}, expert LoRA params: {expert_lora_count}")
|
| 1770 |
+
|
| 1771 |
+
# RED-HAT FIX #2: Weight-delta assert — verify training actually changes weights
|
| 1772 |
+
import hashlib
|
| 1773 |
+
lora_param = next((name, p) for name, p in model.named_parameters()
|
| 1774 |
+
if "lora" in name and p.requires_grad)
|
| 1775 |
+
pre_hash = hashlib.sha256(lora_param[1].data.cpu().numpy().tobytes()).hexdigest()
|
| 1776 |
+
|
| 1777 |
+
# Load Thomas voice demonstrations
|
| 1778 |
+
demos = []
|
| 1779 |
+
with open("/workspace/data/thomas_voice_demos/daimon_voice_demonstrations.jsonl") as f:
|
| 1780 |
+
for line in f:
|
| 1781 |
+
demos.append(json.loads(line))
|
| 1782 |
+
dataset = Dataset.from_list(demos)
|
| 1783 |
+
print(f"Thomas voice demos: {len(dataset)} examples")
|
| 1784 |
+
|
| 1785 |
+
training_args = SFTConfig(
|
| 1786 |
+
output_dir="/workspace/checkpoints/daimon_v2_calibration",
|
| 1787 |
+
per_device_train_batch_size=2,
|
| 1788 |
+
gradient_accumulation_steps=1,
|
| 1789 |
+
num_train_epochs=5,
|
| 1790 |
+
learning_rate=1e-5,
|
| 1791 |
+
lr_scheduler_type="cosine",
|
| 1792 |
+
warmup_ratio=0.1,
|
| 1793 |
+
weight_decay=0.01,
|
| 1794 |
+
bf16=True,
|
| 1795 |
+
max_seq_length=2048,
|
| 1796 |
+
logging_steps=5,
|
| 1797 |
+
save_steps=20,
|
| 1798 |
+
save_total_limit=10,
|
| 1799 |
+
dataloader_num_workers=0,
|
| 1800 |
+
dataset_num_proc=1,
|
| 1801 |
+
gradient_checkpointing=True,
|
| 1802 |
+
gradient_checkpointing_kwargs={"use_reentrant": False},
|
| 1803 |
+
seed=42,
|
| 1804 |
+
)
|
| 1805 |
+
|
| 1806 |
+
trainer = SFTTrainer(
|
| 1807 |
+
model=model,
|
| 1808 |
+
args=training_args,
|
| 1809 |
+
train_dataset=dataset,
|
| 1810 |
+
processing_class=tokenizer,
|
| 1811 |
+
)
|
| 1812 |
+
|
| 1813 |
+
trainer.train()
|
| 1814 |
+
|
| 1815 |
+
# RED-HAT FIX #2: Verify weight actually changed after training
|
| 1816 |
+
post_hash = hashlib.sha256(lora_param[1].data.cpu().numpy().tobytes()).hexdigest()
|
| 1817 |
+
assert pre_hash != post_hash, (
|
| 1818 |
+
f"FATAL: LoRA weight {lora_param[0]} unchanged after 5 epochs of training! "
|
| 1819 |
+
"Phase 3 was a no-op. Check adapter loading."
|
| 1820 |
+
)
|
| 1821 |
+
print(f"Weight-delta verified: {lora_param[0]} changed")
|
| 1822 |
+
|
| 1823 |
+
model.save_pretrained("/workspace/checkpoints/daimon_v2_calibration_best")
|
| 1824 |
+
tokenizer.save_pretrained("/workspace/checkpoints/daimon_v2_calibration_best")
|
| 1825 |
+
print("=== Thomas voice calibration complete ===")
|
| 1826 |
+
|
| 1827 |
+
|
| 1828 |
+
if __name__ == "__main__":
|
| 1829 |
+
parser = argparse.ArgumentParser()
|
| 1830 |
+
parser.add_argument("--phase", choices=["foundation", "calibration"], required=True)
|
| 1831 |
+
parser.add_argument("--model", default="/workspace/models/Qwen3.6-35B-A3B")
|
| 1832 |
+
parser.add_argument("--checkpoint", default=None, help="Checkpoint for calibration phase")
|
| 1833 |
+
parser.add_argument("--hot-experts", default="/workspace/artifacts/hot_experts.json")
|
| 1834 |
+
args = parser.parse_args()
|
| 1835 |
+
|
| 1836 |
+
if args.phase == "foundation":
|
| 1837 |
+
train_foundation(args.model, args.hot_experts)
|
| 1838 |
+
elif args.phase == "calibration":
|
| 1839 |
+
cp = args.checkpoint or "/workspace/checkpoints/daimon_v2_foundation_best"
|
| 1840 |
+
train_calibration(cp)
|
| 1841 |
+
```
|
| 1842 |
+
|
| 1843 |
+
**Critical implementation notes:**
|
| 1844 |
+
|
| 1845 |
+
1. **Unsloth integration:** The script above uses raw PEFT + Transformers for clarity. In production, you may want to use Unsloth's `FastLanguageModel` for the 12x MoE kernel speedup. If using Unsloth, verify that `FastLanguageModel.get_peft_model()` accepts explicit module name lists (not just module type strings). If it does not, use the PEFT-direct approach shown here and accept the throughput penalty, or patch Unsloth's module targeting.
|
| 1846 |
+
|
| 1847 |
+
2. **Data format compatibility:** The `SFTTrainer` expects data in a specific format depending on the `dataset_text_field` or `formatting_func` parameter. If your data uses `{"messages": [...]}` format, you may need to set `dataset_text_field=None` and let TRL auto-detect the chat format, or provide a custom `formatting_func` that applies the Qwen3.6 chat template.
|
| 1848 |
+
|
| 1849 |
+
3. **Gradient checkpointing:** `use_reentrant=False` is required for compatibility with PEFT LoRA on MoE. The reentrant variant can cause incorrect gradients when LoRA interacts with MoE routing.
|
| 1850 |
+
|
| 1851 |
+
4. **Resuming from interruption:** If training is interrupted (spot instance preemption, SIGKILL, etc.), resume from the last checkpoint:
|
| 1852 |
+
```bash
|
| 1853 |
+
python daimon_train.py --phase foundation --checkpoint /workspace/checkpoints/daimon_v2_foundation/checkpoint-XXXX
|
| 1854 |
+
```
|
| 1855 |
+
Add `resume_from_checkpoint=True` to the SFTConfig, or pass the checkpoint path to `trainer.train(resume_from_checkpoint=...)`.
|
| 1856 |
+
|
| 1857 |
+
---
|
| 1858 |
+
|
| 1859 |
+
## Appendix D: LASER Post-Training Script (OPTIONAL — Run Locally on Margaret) {#appendix-d}
|
| 1860 |
+
|
| 1861 |
+
<!-- RED-HAT FIX #1: LASER moved to optional offline experiment per C1 review.
|
| 1862 |
+
Original code was INVERTED — zeroed the LARGEST singular values (principal components)
|
| 1863 |
+
instead of the tail. This would have destroyed the trained model's quality.
|
| 1864 |
+
|
| 1865 |
+
Additionally, blanket application to all 20,480 matrices contradicts LASER's
|
| 1866 |
+
LAyer-SElective protocol. The corrected script below:
|
| 1867 |
+
1. Fixes SVD direction: keeps top singular values, zeros the tail
|
| 1868 |
+
2. Applies per-(layer, matrix) with eval between each application
|
| 1869 |
+
3. Designed for local execution on Margaret ($0), not the cloud pod -->
|
| 1870 |
+
|
| 1871 |
+
```python
|
| 1872 |
+
#!/usr/bin/env python3
|
| 1873 |
+
"""
|
| 1874 |
+
daimon_laser.py
|
| 1875 |
+
OPTIONAL post-budget experiment: LASER (Layer-Selective Rank Reduction) on merged model.
|
| 1876 |
+
Run LOCALLY on Margaret, not on cloud pod.
|
| 1877 |
+
|
| 1878 |
+
Reference: arXiv:2312.13558
|
| 1879 |
+
|
| 1880 |
+
RED-HAT FIX #1 (C1 review): SVD direction corrected.
|
| 1881 |
+
Original code zeroed S[:n_remove] which removes the LARGEST singular values
|
| 1882 |
+
(principal components). torch.linalg.svd returns values in DESCENDING order.
|
| 1883 |
+
Corrected: S[n_keep:] = 0 keeps the top (largest) and zeros the tail (smallest = noise).
|
| 1884 |
+
"""
|
| 1885 |
+
|
| 1886 |
+
import os
|
| 1887 |
+
os.environ["UNSLOTH_COMPILE_DISABLE"] = "1"
|
| 1888 |
+
os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "expandable_segments:True"
|
| 1889 |
+
|
| 1890 |
+
import torch
|
| 1891 |
+
import argparse
|
| 1892 |
+
from transformers import AutoModelForCausalLM, AutoTokenizer
|
| 1893 |
+
|
| 1894 |
+
NUM_LAYERS = 40
|
| 1895 |
+
NUM_EXPERTS = 256
|
| 1896 |
+
|
| 1897 |
+
|
| 1898 |
+
def laser_reduce(weight_matrix, keep_fraction=0.95):
|
| 1899 |
+
"""
|
| 1900 |
+
Keep top fraction of singular values, zero out the tail (noise).
|
| 1901 |
+
|
| 1902 |
+
LASER's key insight: the SMALL singular values of MLP weight matrices
|
| 1903 |
+
often encode noise or spurious correlations. Removing them (keeping
|
| 1904 |
+
only the top components) can improve factual recall on specific tasks.
|
| 1905 |
+
|
| 1906 |
+
IMPORTANT: torch.linalg.svd returns singular values in DESCENDING order.
|
| 1907 |
+
S[0] is the LARGEST. We keep S[:n_keep] (largest) and zero S[n_keep:] (smallest).
|
| 1908 |
+
|
| 1909 |
+
The ORIGINAL code in this spec had this INVERTED (S[:n_remove] = 0), which
|
| 1910 |
+
would have destroyed the principal components of every weight matrix.
|
| 1911 |
+
|
| 1912 |
+
Args:
|
| 1913 |
+
weight_matrix: 2D tensor [out_features, in_features]
|
| 1914 |
+
keep_fraction: fraction of top singular values to KEEP (default: 0.95 = drop bottom 5%)
|
| 1915 |
+
Returns:
|
| 1916 |
+
Modified weight matrix with tail singular values removed
|
| 1917 |
+
"""
|
| 1918 |
+
original_dtype = weight_matrix.dtype
|
| 1919 |
+
W = weight_matrix.float() # SVD requires float32
|
| 1920 |
+
|
| 1921 |
+
U, S, Vh = torch.linalg.svd(W, full_matrices=False)
|
| 1922 |
+
n_keep = max(1, int(len(S) * keep_fraction))
|
| 1923 |
+
|
| 1924 |
+
S_modified = S.clone()
|
| 1925 |
+
S_modified[n_keep:] = 0.0 # Zero out TAIL (small values = noise), KEEP TOP
|
| 1926 |
+
|
| 1927 |
+
W_modified = U @ torch.diag(S_modified) @ Vh
|
| 1928 |
+
return W_modified.to(original_dtype)
|
| 1929 |
+
|
| 1930 |
+
|
| 1931 |
+
def apply_laser_selective(model_path, output_path, keep_fraction=0.95,
|
| 1932 |
+
target_layer=None, target_proj="gate_proj"):
|
| 1933 |
+
"""
|
| 1934 |
+
Apply LASER to a SINGLE (layer, projection) pair.
|
| 1935 |
+
|
| 1936 |
+
DO NOT apply blanket to all 20,480 matrices — that contradicts LASER's
|
| 1937 |
+
LAyer-SElective protocol. Apply one at a time, eval after each.
|
| 1938 |
+
"""
|
| 1939 |
+
print(f"=== LASER: keep_fraction={keep_fraction}, layer={target_layer}, proj={target_proj} ===")
|
| 1940 |
+
|
| 1941 |
+
tokenizer = AutoTokenizer.from_pretrained(model_path)
|
| 1942 |
+
model = AutoModelForCausalLM.from_pretrained(
|
| 1943 |
+
model_path,
|
| 1944 |
+
torch_dtype=torch.bfloat16,
|
| 1945 |
+
device_map="cpu", # LASER on CPU — no GPU needed
|
| 1946 |
+
)
|
| 1947 |
+
|
| 1948 |
+
if target_layer is not None:
|
| 1949 |
+
layers_to_process = [target_layer]
|
| 1950 |
+
else:
|
| 1951 |
+
layers_to_process = range(NUM_LAYERS)
|
| 1952 |
+
|
| 1953 |
+
total_modified = 0
|
| 1954 |
+
for layer_idx in layers_to_process:
|
| 1955 |
+
moe_block = model.model.layers[layer_idx].mlp
|
| 1956 |
+
|
| 1957 |
+
for expert_idx in range(NUM_EXPERTS):
|
| 1958 |
+
expert = moe_block.experts[expert_idx]
|
| 1959 |
+
weight = getattr(expert, target_proj).weight.data
|
| 1960 |
+
new_weight = laser_reduce(weight, keep_fraction)
|
| 1961 |
+
getattr(expert, target_proj).weight.data = new_weight
|
| 1962 |
+
total_modified += 1
|
| 1963 |
+
|
| 1964 |
+
print(f" Layer {layer_idx}: modified {NUM_EXPERTS} experts ({target_proj})")
|
| 1965 |
+
|
| 1966 |
+
print(f"Modified {total_modified} weight matrices")
|
| 1967 |
+
|
| 1968 |
+
# Save
|
| 1969 |
+
model.save_pretrained(output_path, safe_serialization=True)
|
| 1970 |
+
tokenizer.save_pretrained(output_path)
|
| 1971 |
+
print(f"Saved to {output_path}")
|
| 1972 |
+
print("NOW RUN EVAL to verify this improved quality. If not, revert.")
|
| 1973 |
+
|
| 1974 |
+
|
| 1975 |
+
if __name__ == "__main__":
|
| 1976 |
+
parser = argparse.ArgumentParser()
|
| 1977 |
+
parser.add_argument("--model", required=True, help="Path to merged model")
|
| 1978 |
+
parser.add_argument("--output", required=True, help="Output path")
|
| 1979 |
+
parser.add_argument("--keep-fraction", type=float, default=0.95,
|
| 1980 |
+
help="Fraction of top singular values to KEEP (default: 0.95 = drop bottom 5%%)")
|
| 1981 |
+
parser.add_argument("--layer", type=int, default=None,
|
| 1982 |
+
help="Specific layer to target (default: all layers — NOT recommended)")
|
| 1983 |
+
parser.add_argument("--proj", default="gate_proj", choices=["gate_proj", "up_proj"],
|
| 1984 |
+
help="Projection to target (default: gate_proj)")
|
| 1985 |
+
args = parser.parse_args()
|
| 1986 |
+
|
| 1987 |
+
apply_laser_selective(args.model, args.output, args.keep_fraction, args.layer, args.proj)
|
| 1988 |
+
```
|
| 1989 |
+
|
| 1990 |
+
---
|
| 1991 |
+
|
| 1992 |
+
## Reference: Key Papers Cited
|
| 1993 |
+
|
| 1994 |
+
| Paper | arXiv | Used For |
|
| 1995 |
+
|---|---|---|
|
| 1996 |
+
| ESFT (DeepSeek) | 2407.01906 | Expert profiling methodology, shared expert freeze evidence |
|
| 1997 |
+
| HELLoRA | 2605.18795 | Primary training method |
|
| 1998 |
+
| Unified Expert Scoring | 2606.15716 | MAN scoring > frequency scoring for expert importance |
|
| 1999 |
+
| DoRA | 2402.09353 | LoRA upgrade (+1-4% accuracy) |
|
| 2000 |
+
| LASER | 2312.13558 | Post-training rank reduction |
|
| 2001 |
+
| EAQuant | 2506.13329 | QLoRA breaks MoE routing (do not use) |
|
| 2002 |
+
| Geometric coupling (routers) | 2605.12476 | Do not fine-tune routers |
|
| 2003 |
+
| MoE-Sieve | 2603.24044 | Routing-guided LoRA validation |
|
| 2004 |
+
| DR-LoRA | 2601.04823 | Dynamic rank allocation (future iteration) |
|
| 2005 |
+
| OGPSA v3.1 | ogpsa_persona_validation_v3.1.md | Personality geometry capture |
|
| 2006 |
+
| Training best practices | training_pipeline_best_practices.md | alpha=2r, checkpoint averaging, OPLoRA |
|
| 2007 |
+
|
| 2008 |
+
---
|
| 2009 |
+
|
| 2010 |
+
## Changelog
|
| 2011 |
+
|
| 2012 |
+
- **v2.1** (July 2026): Applied 12-point red-hat review fix (daimon_pipeline_v2_redhat.md). Fixes: (1) LASER SVD direction inverted -- moved to optional offline experiment; (2) Phase 3 checkpoint loading via PeftModel.from_pretrained with is_trainable=True + weight-delta assert; (3) Added stratified eval split to fix eval_strategy crash; (4) Replaced Unsloth throughput estimates with realistic raw-PEFT speeds, added 50-step throughput probe gate and max_steps circuit breaker; (5) Disk preflight gate increased from 150 GB to 400 GB; (6) VRAM estimates corrected from 55-65 GB to 90-115 GB peak; (7) TruthfulQA removed from eval battery (contaminated -- in training data 15x); (8) Added checkpoint egress rsync after every save; (9) Pinned dependency versions, torch installed first; (10) Added throughput probe as hard budget gate; (11) Added backward-pass smoke test for gradient flow verification; (12) Budget table corrected to $30-45 primary / $52-67 reserve.
|
| 2013 |
+
- **v2.0** (July 2026): Complete rewrite for H200 cloud training. HELLoRA method selection. Budget-constrained design. SOTA sweep integration.
|
| 2014 |
+
- **v1.0** (June 2026): Original spec targeting Apple Silicon (Margaret). 8-stage pipeline with DPO and RL. See `daimon_training_spec.md`.
|