Instructions to use Realmbird/distilling-safety-adapters with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Realmbird/distilling-safety-adapters with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
distilling-safety adapters
LoRA adapters from a research sprint asking whether safety transfers subliminally through distillation.
Report and code: https://github.com/Realmbird/distilling-safety (see reports/REPORT.md).
Research artifacts, not for deployment. Several students are measurably less safe than the base model (that is the main finding); none of these adapters should be used as a safety intervention.
All adapters sit on Qwen/Qwen2.5-7B-Instruct ("M0").
| Subfolder | What it is |
|---|---|
teachers/T_benign |
Teacher. M0 + LoRA SFT on 2.5k LLM-LAT benign utility rows only (no refusals). |
teachers/T_refusal |
Teacher. M0 + LoRA SFT on 2.5k LLM-LAT harmful-prompt -> refusal pairs only. |
teachers/T_shallow |
Teacher. M0 + LoRA SFT on both (standard refusal SFT; the v2 'shallow' recipe, rebuilt). |
students/T_none_s0 |
Student, seed 0. M0 + LoRA SFT on 20k number sequences written by M0's own numbers (control). |
students/T_none_s1 |
Student, seed 1. M0 + LoRA SFT on 20k number sequences written by M0's own numbers (control). |
students/T_benign_s0 |
Student, seed 0. M0 + LoRA SFT on 20k number sequences written by the benign-only teacher's numbers. |
students/T_benign_s1 |
Student, seed 1. M0 + LoRA SFT on 20k number sequences written by the benign-only teacher's numbers. |
students/T_refusal_s0 |
Student, seed 0. M0 + LoRA SFT on 20k number sequences written by the refusal-only teacher's numbers. |
students/T_refusal_s1 |
Student, seed 1. M0 + LoRA SFT on 20k number sequences written by the refusal-only teacher's numbers. |
students/T_shallow_s0 |
Student, seed 0. M0 + LoRA SFT on 20k number sequences written by the refusal + benign teacher's numbers. |
students/T_shallow_s1 |
Student, seed 1. M0 + LoRA SFT on 20k number sequences written by the refusal + benign teacher's numbers. |
Teacher recipe: LoRA r=64, alpha=64, all attention and MLP projections; lr 1e-4, cosine schedule, 1 epoch, batch 16, max length 1024; completion-only loss. Data: LLM-LAT/harmful-dataset (2.5k-prompt training split) and LLM-LAT/benign-dataset (2.5k rows), split with seed 0.
Student recipe: LoRA r=8, alpha=32; lr 1e-4 constant after warmup; batch 16; one pass over 20k
number-sequence completions (digits only, filtered; prompts shared across teachers). Each student folder has
the final adapter and train_manifest.json.
These are the adapters from the follow-up "benign-only vs refusal-only" experiment (report §4.9). The teachers behind the main results (deep/safety-recovery, LAT, calibrated) were not preserved; they are reproducible with the repo's scripts.
from transformers import AutoModelForCausalLM
from peft import PeftModel
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-7B-Instruct", torch_dtype="auto", device_map="auto")
model = PeftModel.from_pretrained(base, "Realmbird/distilling-safety-adapters", subfolder="teachers/T_shallow")
- Downloads last month
- -