Text Generation
Transformers
Safetensors
English
gemma-3
midtraining
synthetic-document-finetuning
false-belief
research
jbostock commited on
Commit
ae8130b
·
verified ·
1 Parent(s): abdd87d

Qualify curriculum-order inference

Browse files
Files changed (1) hide show
  1. README.md +3 -3
README.md CHANGED
@@ -116,9 +116,9 @@ The final 10M Dolci stage cut spillover by 50.0 percentage points while
116
  preserving saturated belief, but it did not improve canonical accuracy. At the
117
  same one-epoch Python4 dose and total Dolmino/Dolci budgets, the mixed
118
  curriculum ended 15.3 points higher on canon correctness (54.2% versus 38.9%)
119
- and 8.3 points lower on Python3 spillover (20.8% versus 29.2%). Thus the
120
- one-epoch result is strongly sensitive to where instruction tuning occurs,
121
- not just to aggregate token counts.
122
 
123
  The four-epoch ordered-SDF arm is not a clean point on the mixed dose curve
124
  because its data order and instruction-tuning schedule differ. Its stage
 
116
  preserving saturated belief, but it did not improve canonical accuracy. At the
117
  same one-epoch Python4 dose and total Dolmino/Dolci budgets, the mixed
118
  curriculum ended 15.3 points higher on canon correctness (54.2% versus 38.9%)
119
+ and 8.3 points lower on Python3 spillover (20.8% versus 29.2%). This
120
+ preliminary result suggests strong sensitivity to curriculum order, not just
121
+ to aggregate token counts.
122
 
123
  The four-epoch ordered-SDF arm is not a clean point on the mixed dose curve
124
  because its data order and instruction-tuning schedule differ. Its stage