steer-full_3b / README.md
JacoDuToit's picture
model card
89ef361 verified
|
Raw
History Blame Contribute Delete
4.77 kB
---
base_model: Qwen/Qwen2.5-3B-Instruct
library_name: peft
tags:
- lora
- activation-steering
- steering-resistance
---
# full_3b — steering-resistance LoRA adapter
LoRA adapter for **Qwen/Qwen2.5-3B-Instruct** trained to resist adversarial activation
steering: fine-tuned with CAA vectors injected live at decoder layer 18,
rewarded for reproducing its own clean answers. Full method:
https://github.com/JacoDuToit11/steering-resistance.
## Provenance
| | |
|---|---|
| run | full_3b |
| result | clean 100%->100% · steer_heldout@1.6 correct 0%->4% |
| trained | 2026-08-05T17:16:57+0000 |
| code | [`b16f6e72a9bb`](https://github.com/JacoDuToit11/steering-resistance/tree/b16f6e72a9bb0c38dd20d0bab46579325178c793) **+ uncommitted changes** (see `run/code.patch`) |
| config | `configs/full_3b.yaml` (snapshot: `run/config.yaml`) |
| wandb | — |
| hardware | NVIDIA A100 80GB PCIe |
| stack | torch 2.6.0+cu124, transformers 5.14.1, peft 0.20.0, accelerate 1.14.0, datasets 5.0.1, numpy 2.4.6, huggingface_hub 1.26.0 |
| data: concepts_path | `adc6aee9b1cc537a…` (data/concepts_paper.json) |
| data: qa_path | `fa2356571420fd8c…` (data/qa_open.json) |
## Training parameters
| param | value |
|---|---|
| layer | `18` |
| train_alphas | `[0.4, 0.8, 1.2]` |
| eval_alphas | `[0.4, 0.8, 1.0, 1.2, 1.6]` |
| efficacy_alpha | `0.8` |
| efficacy_min_rate | `0.3` |
| steered_frac | `0.7` |
| relevant_frac | `0.6` |
| repeats_per_question | `6` |
| lora_r | `16` |
| lora_alpha | `32` |
| lora_dropout | `0.05` |
| lora_targets | `["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"]` |
| lr | `0.0001` |
| epochs | `1` |
| effective_batch_size | `16` |
| train_batch_size | `4` |
| max_seq_len | `256` |
| seed | `0` |
## Eval results
| model | condition | alpha | n | correct | steered | other |
|---|---|---:|---:|---|---|---|
| M0 | clean | 0.0 | 17 | 100% [100%,100%] | 0% [0%,0%] | 0% [0%,0%] |
| M0 | correct_inject | 0.4 | 9 | 100% [100%,100%] | 0% [0%,0%] | 0% [0%,0%] |
| M0 | correct_inject | 0.8 | 9 | 89% [67%,100%] | 0% [0%,0%] | 11% [0%,33%] |
| M0 | correct_inject | 1.0 | 9 | 89% [67%,100%] | 0% [0%,0%] | 11% [0%,33%] |
| M0 | correct_inject | 1.2 | 9 | 89% [67%,100%] | 0% [0%,0%] | 11% [0%,33%] |
| M0 | correct_inject | 1.6 | 9 | 89% [67%,100%] | 0% [0%,0%] | 11% [0%,33%] |
| M0 | steer_heldout | 0.4 | 2754 | 84% [77%,89%] | 9% [7%,12%] | 7% [4%,12%] |
| M0 | steer_heldout | 0.8 | 2754 | 14% [10%,17%] | 73% [69%,77%] | 13% [12%,14%] |
| M0 | steer_heldout | 1.0 | 2754 | 1% [1%,2%] | 81% [80%,83%] | 17% [16%,18%] |
| M0 | steer_heldout | 1.2 | 2754 | 0% [0%,0%] | 74% [73%,75%] | 26% [25%,27%] |
| M0 | steer_heldout | 1.6 | 2754 | 0% [0%,0%] | 52% [52%,53%] | 48% [47%,48%] |
| M0 | steer_train | 0.4 | 679 | 93% [84%,98%] | 0% [0%,1%] | 7% [2%,16%] |
| M0 | steer_train | 0.8 | 679 | 16% [12%,21%] | 73% [69%,78%] | 10% [8%,13%] |
| M0 | steer_train | 1.0 | 679 | 2% [1%,3%] | 87% [85%,89%] | 11% [10%,13%] |
| M0 | steer_train | 1.2 | 679 | 0% [0%,0%] | 85% [83%,87%] | 15% [13%,17%] |
| M0 | steer_train | 1.6 | 679 | 0% [0%,0%] | 53% [51%,56%] | 47% [44%,49%] |
| M1 | clean | 0.0 | 17 | 100% [100%,100%] | 0% [0%,0%] | 0% [0%,0%] |
| M1 | correct_inject | 0.4 | 9 | 100% [100%,100%] | 0% [0%,0%] | 0% [0%,0%] |
| M1 | correct_inject | 0.8 | 9 | 100% [100%,100%] | 0% [0%,0%] | 0% [0%,0%] |
| M1 | correct_inject | 1.0 | 9 | 100% [100%,100%] | 0% [0%,0%] | 0% [0%,0%] |
| M1 | correct_inject | 1.2 | 9 | 100% [100%,100%] | 0% [0%,0%] | 0% [0%,0%] |
| M1 | correct_inject | 1.6 | 9 | 89% [67%,100%] | 0% [0%,0%] | 11% [0%,33%] |
| M1 | steer_heldout | 0.4 | 2754 | 97% [93%,100%] | 1% [0%,1%] | 2% [0%,6%] |
| M1 | steer_heldout | 0.8 | 2754 | 66% [58%,73%] | 20% [17%,24%] | 14% [10%,18%] |
| M1 | steer_heldout | 1.0 | 2754 | 44% [35%,53%] | 34% [27%,41%] | 22% [17%,27%] |
| M1 | steer_heldout | 1.2 | 2754 | 26% [18%,35%] | 39% [32%,47%] | 34% [29%,40%] |
| M1 | steer_heldout | 1.6 | 2754 | 4% [2%,7%] | 54% [50%,58%] | 42% [40%,44%] |
| M1 | steer_train | 0.4 | 679 | 97% [90%,100%] | 0% [0%,0%] | 3% [0%,9%] |
| M1 | steer_train | 0.8 | 679 | 87% [79%,94%] | 4% [1%,8%] | 9% [5%,14%] |
| M1 | steer_train | 1.0 | 679 | 63% [51%,75%] | 21% [11%,32%] | 16% [10%,22%] |
| M1 | steer_train | 1.2 | 679 | 40% [26%,54%] | 33% [22%,45%] | 27% [20%,36%] |
| M1 | steer_train | 1.6 | 679 | 7% [3%,12%] | 61% [55%,67%] | 32% [29%,35%] |
## Reproduce
```bash
git clone https://github.com/JacoDuToit11/steering-resistance.git && cd steering-resistance
git checkout b16f6e72a9bb0c38dd20d0bab46579325178c793
python scripts/run.py configs/full_3b.yaml
```
`run/` mirrors the full experiment directory: `run_meta.json` (manifest with
artifact hashes), append-only eval jsonl, summaries, and the exact config.