Asilarkness commited on
Commit
bc51733
·
verified ·
1 Parent(s): 28ff334

Reasoning SFT then RL v3: final Stage A rejection report

Browse files
candidates/budgie-alignment-v2/reasoning-sft-then-rl-v3/recovery/V3_FINAL_REPORT.md ADDED
@@ -0,0 +1,30 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Long reasoning SFT → RL v3
2
+
3
+ ## Decision
4
+
5
+ **Stage A rejected; Stage B RL was not started. `verified-math-a025` remains the leader.**
6
+
7
+ ## Stage A
8
+
9
+ - Scanned 5,251 unrelated OpenR1 rows.
10
+ - Retained 1,008 decontaminated, symbolically verified long mathematics solutions.
11
+ - Split 800 SFT train / 208 random-math development.
12
+ - Added 108 manual information-management retention examples.
13
+ - 908 encoded examples; average 577 tokens, maximum 1,536.
14
+ - One full curriculum epoch, 227 updates, FP32 master weights, BF16 autocast.
15
+ - Peak LR 3e-7 and sampled-action KL 0.12 to the immutable leader.
16
+ - Zero benchmark-family training rows.
17
+
18
+ ## Results
19
+
20
+ | Gate | Leader | SFT alpha .005 |
21
+ |---|---:|---:|
22
+ | Information dev | 3/21 | **4/21** |
23
+ | Information loops | 0 | 0 |
24
+ | Random verified math dev | 1/32 | 1/32 |
25
+ | GSM8K fixed | **5/30** | 4/30 |
26
+ | MATH fixed | **3/15** | 2/15 |
27
+ | ARC fixed | 11/30 | 11/30 |
28
+ | FOLIO fixed | 13/30 | 13/30 |
29
+
30
+ The SFT direction gives a small information gain, but does not improve random math and regresses fixed GSM/MATH. Smaller alphas do not improve held-out information management. Under the predeclared protocol, RL may run only after an accepted SFT anchor, so follow-on RL was correctly skipped.