Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation
Abstract
Backdoor attacks can be implanted in Large Language Models (LLMs) during training, causing unwanted behaviour when a trigger appears in the input. Existing backdoor defences for LLMs attempt to remove the backdoor but inadvertently shift the model's output distribution to benign prompts, which can result in degraded model performance and safety. We propose NEEDLE, a training-free method for targeted backdoor removal. Once a trigger has been identified, our method estimates a backdoor direction and a refusal subspace through activation vectors, then applies sequential weight orthogonalisation to suppress the backdoor while preventing changes in refusal-related representations. NEEDLE requires neither a clean reference model nor the original poisoned training data. Evaluation is conducted across multiple model families and attack types. NEEDLE achieves the lowest mean Attack Success Rate (ASR) among the evaluated defences, including 0% on challenging code injection attacks, while resulting in the lowest KL divergence and minimal changes in capability and safety.
Community
NEEDLE is a training-free backdoor defence designed to remove a backdoor with minimal changes to model behaviour and safety. It estimates a backdoor direction (how the trigger shifts the model's activations) and a refusal subspace (directions that mediate refusal), then edits the model's weights layer by layer: each layer's attention and MLP output weights are orthogonalised against the backdoor direction while keeping their refusal projections fixed, and a closed-form correction keeps the activations' refusal projections unchanged as earlier layers are edited.
Codebase: https://github.com/LocaiLabs/NEEDLE/
Models: https://huggingface.co/collections/locailabs/needle
This is great work, super relevant in this day and age!
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Backdoor Purification for LoRA-Tuned LLMs via Null-Space Projection (2026)
- Bait-and-Recover: Poisoning Internal Refusal Signals to Defend LLMs against White-Box Editing Jailbreaks (2026)
- LoRAScan: Detecting Backdoor Prompts in Low-Rank Adapters for Large Language Models via Down-Projection Activation Spikes (2026)
- Backdoor Containment via Expert Quarantine and Shutdown in LLMs (2026)
- Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration (2026)
- Backdoor Decontamination Dynamics in LLM Agents (2026)
- Refusal Localizes, the Damage Relocates: Safety Layers Under Few-Sample Fine-Tuning (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.00348 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 24
locailabs/Gemma-3-4B-IT-Sentiment-BadNet-Backdoored
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper