distilling-safety adapters

LoRA adapters from a research sprint asking whether safety transfers subliminally through distillation. Report and code: https://github.com/Realmbird/distilling-safety (see reports/REPORT.md).

Research artifacts, not for deployment. Several students are measurably less safe than the base model (that is the main finding); none of these adapters should be used as a safety intervention.

All adapters sit on Qwen/Qwen2.5-7B-Instruct ("M0").

Subfolder What it is
teachers/T_benign Teacher. M0 + LoRA SFT on 2.5k LLM-LAT benign utility rows only (no refusals).
teachers/T_refusal Teacher. M0 + LoRA SFT on 2.5k LLM-LAT harmful-prompt -> refusal pairs only.
teachers/T_shallow Teacher. M0 + LoRA SFT on both (standard refusal SFT; the v2 'shallow' recipe, rebuilt).
students/T_none_s0 Student, seed 0. M0 + LoRA SFT on 20k number sequences written by M0's own numbers (control).
students/T_none_s1 Student, seed 1. M0 + LoRA SFT on 20k number sequences written by M0's own numbers (control).
students/T_benign_s0 Student, seed 0. M0 + LoRA SFT on 20k number sequences written by the benign-only teacher's numbers.
students/T_benign_s1 Student, seed 1. M0 + LoRA SFT on 20k number sequences written by the benign-only teacher's numbers.
students/T_refusal_s0 Student, seed 0. M0 + LoRA SFT on 20k number sequences written by the refusal-only teacher's numbers.
students/T_refusal_s1 Student, seed 1. M0 + LoRA SFT on 20k number sequences written by the refusal-only teacher's numbers.
students/T_shallow_s0 Student, seed 0. M0 + LoRA SFT on 20k number sequences written by the refusal + benign teacher's numbers.
students/T_shallow_s1 Student, seed 1. M0 + LoRA SFT on 20k number sequences written by the refusal + benign teacher's numbers.

Teacher recipe: LoRA r=64, alpha=64, all attention and MLP projections; lr 1e-4, cosine schedule, 1 epoch, batch 16, max length 1024; completion-only loss. Data: LLM-LAT/harmful-dataset (2.5k-prompt training split) and LLM-LAT/benign-dataset (2.5k rows), split with seed 0.

Student recipe: LoRA r=8, alpha=32; lr 1e-4 constant after warmup; batch 16; one pass over 20k number-sequence completions (digits only, filtered; prompts shared across teachers). Each student folder has the final adapter and train_manifest.json.

These are the adapters from the follow-up "benign-only vs refusal-only" experiment (report §4.9). The teachers behind the main results (deep/safety-recovery, LAT, calibrated) were not preserved; they are reproducible with the repo's scripts.

from transformers import AutoModelForCausalLM
from peft import PeftModel
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-7B-Instruct", torch_dtype="auto", device_map="auto")
model = PeftModel.from_pretrained(base, "Realmbird/distilling-safety-adapters", subfolder="teachers/T_shallow")
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Realmbird/distilling-safety-adapters

Base model

Qwen/Qwen2.5-7B
Adapter
(2807)
this model