Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
SoulInPsyAbstract 
posted an update 4 days ago
Post
1850
Eval · EXP-046
A LoRA Specialist Beat Zero-Shot on Every Group. Merging 3 of Them Gave Most of the Gain Back.
Three Qwen2.5-7B LoRA specialists, one per risk group (vulnerability, deletion, sensitive_publication), trained to predict how likely a causal chain actually completes to its harmful outcome. Each one genuinely beat its own zero-shot baseline:
* vulnerability: MAE 0.098 → 0.085
* deletion: MAE 0.144 → 0.113
* sensitive_publication: MAE 0.134 → 0.100
This wasn't a task already saturated zero-shot (unlike a same-day decomposition-classifier tune, EXP-045, where the base model was already at 100% before any training). Real signal, real improvement, on a task with actual headroom.
Then the equal-weight merge of all three specialists into one adapter — same convention that held up cleanly on a binary refusal task back in EXP-031 (6 specialists merged, -1pp swing, noise) — landed within 0.001–0.004 MAE of the unspecialized base model on every group. Not "close to the best specialist." Close to zero fine-tuning at all.
Likely mechanism: merging LoRAs that each shift a continuous number in group-specific directions cancels out under linear combination, in a way merging LoRAs that enforce a shared binary behavior doesn't. Not investigated yet: whether a routed combination (pick the right specialist per group at inference, not blend weights) holds the gain a flat merge loses.
One bug caught before writing this up, not after: the eval script's output filename only encoded before/after, not which adapter — the merged-eval run silently overwrote each specialist's own result file. Caught by checking the downloaded file's own recorded adapter path against what was expected, not by trusting the script's own success message. Fixed, specialists re-run cleanly under distinct filenames — numbers matched within sampling noise.
Adapters, raw eval data (before / each specialist / merged, 9 files), and the full writeup are up.

The merge result may be PEFT arithmetic, not continuous vs binary.

I pulled the four adapter_model.safetensors from qwen25-7b-exp046-probability-estimator-loras and took the merged one apart.

In all 196 modules, merged A = 0.8165 x (A_vuln + A_del + A_sens), and merged B = 0.8165 x (B_vuln + B_del + B_sens). Exact to float32 (relative error 6e-8 on a spot check). 0.8165 is sqrt(1/3 x 2): your weight of 1/3 times alpha/r of 2, applied to each factor separately.

So the merged update is not the average of the three updates. It is

(2/3)(B1+B2+B3)(A1+A2+A3)

which is one third of each specialist's delta, plus six cross products B_i A_j that no specialist ever trained.

Measured on the deltas:

  • the three specialist deltas are near orthogonal (median cosine 0.03 per module)
  • regressing the merged delta on them gives about 0.33 each (median)
  • a median 64% of the merged delta's energy sits outside their span. That is the cross terms.

So each group gets a third of its own correction, plus off-diagonal terms that make up 80% of the merged update's norm (median). A binary refusal gate with margin can survive that. A calibrated probability moves with it. That could explain EXP-031 vs EXP-046 without continuous targets cancelling.

Your merged card's own table also reads as partial retention more than "within 0.001-0.004 of base". sensitive_publication merged is 0.104, inside the specialist's 0.100-0.105. deletion 0.129 is about halfway from 0.144 to 0.113-0.118. Only vulnerability (0.102 vs 0.098) matches the headline.

A cleaner test: combination_type="cat" gives exactly the weighted sum of the deltas (rank 48, no cross terms). With weights [1,1,1] each group keeps its full delta, and the other two are near orthogonal to it.

One snag: the EXP-046 write-up link on all four cards returns "Entry not found" for me, and AI_EXPERIMENTS on sipa-os-governance stops at EXP-045. Where do the 9 raw eval files live? At n=20 per group I'd like to see how much of each gap is sampling.