marp: true
theme: default
paginate: true
style: >
section {
font-size: 22px;
padding: 40px 60px;
}
h1 { font-size: 1.8em; color: #1a1a2e; }
h2 { font-size: 1.3em; color: #16213e; border-bottom: 2px solid #e94560;
padding-bottom: 6px; }
.highlight { background: #fff3cd; padding: 4px 8px; border-radius: 4px; }
.red { color: #e94560; font-weight: bold; }
.green { color: #27ae60; font-weight: bold; }
.blue { color: #2980b9; font-weight: bold; }
.box {
border: 1.5px solid #aaa;
border-radius: 8px;
padding: 12px 18px;
background: #f9f9f9;
margin: 8px 0;
}
.split { display: flex; gap: 30px; }
.split > div { flex: 1; }
table { width: 100%; border-collapse: collapse; }
th { background: #1a1a2e; color: white; padding: 8px; }
td { padding: 7px 10px; border: 1px solid #ccc; }
tr:nth-child(even) { background: #f5f5f5; }
code { background: #f0f0f0; padding: 2px 6px; border-radius: 3px; font-size:
0.9em; }
Suppressing Prior Knowledge in Vision-Language Models
via Adversarial Sparse Autoencoder Training
[Your Name]
April 2026
Outline
- Motivation — Why VLMs hallucinate due to prior knowledge
- Problem Formulation — Output suppression vs. Knowledge suppression
- Method — Adversarial SAE-based representation editing
- Evaluation Setup
- Preliminary Results
- Discussion & Next Steps
1. VLMs Have Strong Priors — And That's a Problem
Vision-Language Models (VLMs) like LLaVA combine a vision encoder with a language model.
They are trained on massive web data, which encodes statistical co-occurrences:
"Bathrooms usually have toilets"
"Kitchens usually have microwaves"
These priors are useful — until the image contradicts them.
Hallucination via prior knowledge:
Image: [bathroom, NO toilet]
Prompt: "Describe this image."
Base model output:
"A bathroom with a toilet, sink,
and tiled walls..."
↑ HALLUCINATED
The model invented the toilet because the scene predicts one.
What Makes This Hard: The Knowledge Is Encoded Deep
A naïve fix: fine-tune the model to not say "toilet" in bathroom images.
Problem: The knowledge still lives in the model's internal representations.
[Bathroom image]
│
┌─────────▼──────────────────────────────────┐
│ LLaVA (fine-tuned to suppress output) │
│ │
│ Layer 5: h₅ ──► still encodes toilet? │
│ Layer 15: h₁₅ ──► still encodes toilet? │
│ Layer 31: h₃₁ ──► "a bathroom with a ..." │
└─────────────────────────────────────────────┘
- A simple fine-tune only patches the output layer
- A few more fine-tuning steps on a toilet dataset can recover the hallucination
- True suppression requires removing the prior from every layer
2. Two Levels of Suppression
Output Suppression
The model does not output B in context A
- Check:
P(output mentions B | sees A, ¬B image)is low - Internal layers may still encode B
- Easily reversed by downstream fine-tuning
- Not robust to probing
Knowledge Suppression ✓ (Our Goal)
No layer's representation contains B-information when B is visually absent
Formal criterion: For all layers l and all probes p:
p(f_l(x, θ_supp)) ⊥ Z_A(x)
given Z_B(x) = 0
"If B is not in the image, then no probe — trained on SAE features from any layer — can use A's presence to predict B."
Key insight: We use Sparse Autoencoders (SAEs) to measure and attack B-representations at every layer.
3. Tool: Sparse Autoencoders (SAEs)
SAEs decompose hidden states into sparse, interpretable features:
h_l(x) ∈ ℝ^d (dense hidden state at layer l)
│
▼ SAE encoder (frozen, pre-trained)
f_l(x) ∈ ℝ^D (sparse latent: D >> d, ~1% active)
│
▼ Each dimension ≈ one "concept"
f[42] → "toilet-related feature"
f[71] → "bathroom-related feature"
Why SAEs?
- Pre-trained SAE gives us a universal feature basis across all layers
- Probes on SAE features are more interpretable than probes on raw activations
- We can measure which features were suppressed
Method Overview
Setting: Relation bathroom → toilet
Three disjoint data subsets:
| Subset | Description | Role |
|---|---|---|
| D_{A,¬B} | Bathroom images, no toilet | Suppression target |
| D_B | Images with toilet | Retention (don't forget B) |
| D_general | Random images (CC3M) | Capability preservation |
Goal: Train LoRA adapter θ such that on D_{A,¬B}, no layer contains toilet-information — while preserving toilet recognition on D_B.
Method: Adversarial Training Loop
┌──────────────────────────────────────────────────────────────────┐
│ Alternating Optimization │
│ │
│ ┌─── PROBE STEP (fix θ, update φ) ──────────────────────────┐ │
│ │ │ │
│ │ [Image x] → LLaVA(θ, fixed) → h_l → SAE → f_l(x) │ │
│ │ │ │ │
│ │ probe p_φ │ │
│ │ │ │ │
│ │ φ ← minimize BCE(p_φ(f_l(x)), y_B) ▼ │ │
│ │ "Find any toilet signal" │ │
│ └────────────────────────────────────────────────────────────┘ │
│ │
│ ┌─── MODEL STEP (fix φ, update θ) ──────────────────────────┐ │
│ │ │ │
│ │ [Image x] → LLaVA(θ) → h_l → SAE → f_l(x) → probe p_φ │ │
│ │ │ │ │ │
│ │ KL(base) max entropy │ │
│ │ (don't drift) "confuse the probe" │ │
│ └────────────────────────────────────────────────────────────┘ │
└──────────────────────────────────────────────────────────────────┘
Iterate until the probe can no longer detect toilet-information in D_{A,¬B}.
Method: Loss Function
Combined objective for the model step:
L_task — KL divergence from base model
(computed on D_general + D_B; excludes D_{A,¬B})
→ Preserves general behavior
L_suppress — on D_{A,¬B}:−E[ H(p_φ(f_l(x))) ]
→ Push probe output to 0.5 (maximum uncertainty)
→ Model erases B-information from those images
L_ret — KL from base on D_B
→ Don't forget how to recognize toilets when they're actually there
Probe step — standard BCE
→ Probe maximally sensitive before each model update
→ Adversary keeps re-sharpening; model keeps erasing
Config: λ_adv = 20, λ_ret = 1, 5 probe steps per model step
Method: Probe Labels
The probe is trained to detect B-information (toilet features):
| Image type | Scene (A) | Object (B) | Probe label |
|---|---|---|---|
| Bathroom + Toilet | ✓ | ✓ | 1 |
| Toilet only | ✗ | ✓ | 1 |
| Bathroom only | ✓ | ✗ | 0 |
| Neither | ✗ | ✗ | 0 |
Mode:
object_only— the probe learns to fire on actual toilet presence, not the bathroom-toilet association. This way, suppression of probe output on D_{A,¬B} means we erased B-features, not just the A-B correlation.
Warmup: Probe is pre-trained for 3 epochs on frozen model before adversarial training begins.
Architecture Details
LLaVA-1.5-7B (frozen base weights)
+ LoRA adapter (r=16, α=32, all-linear)
Frozen SAE (pre-trained on LLaVA hidden states)
d_model → d_SAE (single shared encoder across all layers)
Linear probes (one per monitored layer)
d_SAE → 1 logit
Training setup:
- 20 epochs, batch size 8, grad accumulation 8 (effective = 64)
- lr = 1e-5 (model), 1e-4 (probe)
- Dataset:
pbcong/bathroom-toilet(HuggingFace) - General data: 1000 CC3M images for capability preservation
- Probe warmup: 3 epochs on frozen model
4. Evaluation Setup
Task: Generate a caption for each image, check for toilet keywords:
Keywords: ["toilet", "toilets", "Toilet", "Toilets", "commode", "lavatory", "latrine"]
Metrics:
| Metric | Definition | Applies to |
|---|---|---|
| Hallucination rate | % of captions that incorrectly mention toilet | D_{A,¬B} |
| Miss rate | % of captions that fail to mention toilet | D_B, toilet_only |
| Keyword mention | % of any mention | Overall |
Models compared:
- Base — LLaVA-1.5-7B (no LoRA)
- LoRA — same model + adversarially trained adapter
5. Preliminary Results
| Category | N | Base | LoRA | Change |
|---|---|---|---|---|
| toilet_only | 22 | 90.9% mention | 90.9% mention | ±0% miss |
| bathroom_only (D_{A,¬B}) | 1157 | 21.9% hallu | 2.9% hallu | ↓ 18.9 pp |
| bathroom_toilet (D_B) | 535 | 91.4% mention | 84.7% mention | ↓ 6.7 pp miss |
| Overall | 1714 | 44.5% mention | 29.6% mention | — |
Interpretation:
- Suppression succeeds: Hallucination on bathroom-only drops from 21.9% → 2.9% (87% relative reduction)
- Retention holds: On images with actual toilets, miss rate only increases by 6.7 pp (from 8.6% → 15.3%)
- Toilet recognition preserved:
toilet_onlyimages are unaffected (90.9% both)
Results — Visual Summary
Hallucination Rate (bathroom-only images)
Base model ████████████████████████ 21.9%
LoRA (ours) ███ 2.9%
↑ 87% reduction
Miss Rate (images with actual toilet)
Base model ████ 8.6%
LoRA (ours) ████████ 15.3%
↑ Small increase (acceptable trade-off)
Toilet recognition (toilet-only images)
Base model ██████████████████████████ 90.9%
LoRA (ours) ██████████████████████████ 90.9%
↑ Perfectly preserved
The model learned to not invoke the bathroom→toilet prior without forgetting what a toilet looks like.
6. Discussion
What the results suggest:
- The adversarial probe successfully identifies and removes toilet-related representations from bathroom-only images
- LoRA (low-rank adapter) has enough capacity to edit internal representations while staying close to the base model
- The KL retention loss effectively prevents catastrophic forgetting of toilet recognition
Open questions (next steps):
- Probe accuracy — How does probe accuracy evolve during training? Does it converge to ~50% (random) on D_{A,¬B}?
- Layer specificity — Which layers changed most? Are early or late layers responsible for the prior?
- Generalization — Does suppression transfer to other relations (e.g., kitchen→microwave)?
- Recovery test — Can a few-shot fine-tune re-introduce the hallucination? (Test robustness vs output suppression)
- Covariance test — Compute
Cov(A_features, B_features | Z_B=0)to confirm the formal criterion
What Makes This Different from Prior Work
Prior work on VLM hallucination:
- Instruction tuning / RLHF to reduce hallucinations
- Contrastive decoding at inference time
- Output-level supervision only
Limitation: These only patch the surface behavior. Internal representations still encode the spurious association.
Our approach:
- Targets internal representations directly via SAE features
- Uses an adversarial probe — removes exactly what a probe can detect
- Operates at every layer (probe across all decoder layers)
- Formally defined criterion for knowledge suppression
Knowledge suppression → harder to recover, more trustworthy model.
Summary
Problem: VLMs hallucinate objects that are statistically correlated with visible context (prior knowledge hallucination). Output-level suppression is insufficient — the knowledge remains in internal layers.
Method: Adversarial training with frozen SAE probes. A linear probe is trained to detect object-B information from SAE latent features; the model (LoRA adapter) is trained to confuse the probe on images where B is absent.
Result: Hallucination rate on bathroom-only images: 21.9% → 2.9% (87% reduction), with minimal impact on toilet recognition when toilets are actually present.
Next: Probe accuracy curves · Layer-wise analysis · Transfer to other relations · Recovery experiments
Appendix: Training Configuration
| Parameter | Value |
|---|---|
| Base model | LLaVA-1.5-7B |
| LoRA rank / α | 16 / 32 |
| Epochs | 20 |
| Batch size (effective) | 64 (8 × 8 grad accum) |
| λ_adv | 20.0 |
| λ_ret | 1.0 |
| Probe steps per model step | 5 |
| Probe warmup epochs | 3 |
| Probe label mode | object_only |
| Exclude D_{A,¬B} from task KL | ✓ |
| General data (capability) | 1000 CC3M images |