MAEMM cross-uplift arm acts_mlp: 100k real activations + 100k layer-42 MLP-neuron directions

Midtrain (1 epoch, lr 1e-4) on a 200k bank of 100k real activations + 100k layer-42 MLP-neuron directions, from the 23M real-activation SFT init, then 100 RL steps (CISPO / ScaleRL, 128 directions × 16 samples per step, lr 1e-5, 10 warmup steps) on the same bank. Report: http://5.78.192.0/reports/view/maemm-uplift-matrix/report.html

All checkpoints are LoRA adapters (r 64, α 16, rsLoRA, all linear layers) of the MAEMM activation→text inverter for Qwen3.6-27B layer 42 (inject h + ||h||·v at the layer-1 marker; text whose clean layer-42 activation points along v). Code: https://github.com/ceselder/maemm. Eval = 512 held-out directions/family, best-of-4 at T=1, cosine of the clean base L42 activation (max over last 5 tokens). Subfolders are PEFT adapters: PeftModel.from_pretrained(base, repo, subfolder="<name>").

Held-out evals

checkpoint mean_all realact SAE norm_act SAE rank-1 BSF probes MLP fire-back
init (23M realact SFT) 0.368 0.477 0.416 0.189 0.296 0.226 0.121
sft_final (after midtrain) 0.231 0.282 0.110 0.033 0.218 0.161 0.058
rl_step_25 0.311 0.410 0.188 0.082 0.257 0.197 0.071
rl_step_50 0.357 0.473 0.308 0.137 0.283 0.222 0.194
rl_step_100 0.393 0.517 0.547 0.242 0.309 0.241 0.553
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ceselder/maemm-uplift-acts_mlp

Base model

Qwen/Qwen3.6-27B
Adapter
(554)
this model

Collection including ceselder/maemm-uplift-acts_mlp