4B matched control (fixed-normal prior, identical SFT+RL recipe to hidden); folded LoRA, step 400 8c686c6 verified karantonis commited on 1 day ago
scalar_map RL arm (RLCR-style point-target reward, matched to hidden, step400) 5491c0c verified karantonis commited on 29 days ago
Stage-3 verbalization adapter v1 (LoRA r64, self-distil to frozen-I4 belief; KL(q||v)=0.021) c2be80d verified karantonis commited on Jun 24