---
marp: true
theme: default
paginate: true
style: |
section {
font-size: 22px;
padding: 40px 60px;
}
h1 { font-size: 1.8em; color: #1a1a2e; }
h2 { font-size: 1.3em; color: #16213e; border-bottom: 2px solid #e94560; padding-bottom: 6px; }
.highlight { background: #fff3cd; padding: 4px 8px; border-radius: 4px; }
.red { color: #e94560; font-weight: bold; }
.green { color: #27ae60; font-weight: bold; }
.blue { color: #2980b9; font-weight: bold; }
.box {
border: 1.5px solid #aaa;
border-radius: 8px;
padding: 12px 18px;
background: #f9f9f9;
margin: 8px 0;
}
.split { display: flex; gap: 30px; }
.split > div { flex: 1; }
table { width: 100%; border-collapse: collapse; }
th { background: #1a1a2e; color: white; padding: 8px; }
td { padding: 7px 10px; border: 1px solid #ccc; }
tr:nth-child(even) { background: #f5f5f5; }
code { background: #f0f0f0; padding: 2px 6px; border-radius: 3px; font-size: 0.9em; }
---
# Suppressing Prior Knowledge in Vision-Language Models
## via Adversarial Sparse Autoencoder Training
**[Your Name]**
April 2026
---
## Outline
1. Motivation — Why VLMs hallucinate due to prior knowledge
2. Problem Formulation — Output suppression vs. **Knowledge suppression**
3. Method — Adversarial SAE-based representation editing
4. Evaluation Setup
5. Preliminary Results
6. Discussion & Next Steps
---
## 1. VLMs Have Strong Priors — And That's a Problem
**Vision-Language Models (VLMs)** like LLaVA combine a vision encoder with a language model.
They are trained on massive web data, which encodes **statistical co-occurrences**:
> *"Bathrooms usually have toilets"*
> *"Kitchens usually have microwaves"*
These priors are useful — until the image contradicts them.
**Hallucination via prior knowledge:**
```
Image: [bathroom, NO toilet]
Prompt: "Describe this image."
Base model output:
"A bathroom with a toilet, sink,
and tiled walls..."
↑ HALLUCINATED
```
The model *invented* the toilet because the scene *predicts* one.
---
## What Makes This Hard: The Knowledge Is Encoded Deep
A naïve fix: fine-tune the model to not say "toilet" in bathroom images.
**Problem:** The *knowledge* still lives in the model's internal representations.
```
[Bathroom image]
│
┌─────────▼──────────────────────────────────┐
│ LLaVA (fine-tuned to suppress output) │
│ │
│ Layer 5: h₅ ──► still encodes toilet? │
│ Layer 15: h₁₅ ──► still encodes toilet? │
│ Layer 31: h₃₁ ──► "a bathroom with a ..." │
└─────────────────────────────────────────────┘
```
- A simple fine-tune **only patches the output layer**
- A few more fine-tuning steps on a toilet dataset can **recover the hallucination**
- True suppression requires removing the prior from **every layer**
---
## 2. Two Levels of Suppression
### Output Suppression
> The model **does not output** B in context A
- Check: `P(output mentions B | sees A, ¬B image)` is low
- **Internal layers may still encode B**
- Easily reversed by downstream fine-tuning
- Not robust to probing
### Knowledge Suppression ✓ (Our Goal)
> No layer's representation contains B-information when B is visually absent
**Formal criterion:** For all layers *l* and all probes *p*:
```
p(f_l(x, θ_supp)) ⊥ Z_A(x)
given Z_B(x) = 0
```
*"If B is not in the image, then no probe — trained on SAE features from any layer — can use A's presence to predict B."*
> **Key insight:** We use **Sparse Autoencoders (SAEs)** to measure and attack B-representations at every layer.
---
## 3. Tool: Sparse Autoencoders (SAEs)
SAEs decompose hidden states into **sparse, interpretable features**:
```
h_l(x) ∈ ℝ^d (dense hidden state at layer l)
│
▼ SAE encoder (frozen, pre-trained)
f_l(x) ∈ ℝ^D (sparse latent: D >> d, ~1% active)
│
▼ Each dimension ≈ one "concept"
f[42] → "toilet-related feature"
f[71] → "bathroom-related feature"
```
**Why SAEs?**
- Pre-trained SAE gives us a **universal feature basis** across all layers
- Probes on SAE features are more **interpretable** than probes on raw activations
- We can measure *which features* were suppressed
---
## Method Overview
**Setting:** Relation `bathroom → toilet`
Three disjoint data subsets:
| Subset | Description | Role |
|---|---|---|
| **D_{A,¬B}** | Bathroom images, **no toilet** | Suppression target |
| **D_B** | Images **with toilet** | Retention (don't forget B) |
| **D_general** | Random images (CC3M) | Capability preservation |
**Goal:** Train LoRA adapter θ such that on D_{A,¬B}, no layer contains toilet-information — while preserving toilet recognition on D_B.
---
## Method: Adversarial Training Loop
```
┌──────────────────────────────────────────────────────────────────┐
│ Alternating Optimization │
│ │
│ ┌─── PROBE STEP (fix θ, update φ) ──────────────────────────┐ │
│ │ │ │
│ │ [Image x] → LLaVA(θ, fixed) → h_l → SAE → f_l(x) │ │
│ │ │ │ │
│ │ probe p_φ │ │
│ │ │ │ │
│ │ φ ← minimize BCE(p_φ(f_l(x)), y_B) ▼ │ │
│ │ "Find any toilet signal" │ │
│ └────────────────────────────────────────────────────────────┘ │
│ │
│ ┌─── MODEL STEP (fix φ, update θ) ──────────────────────────┐ │
│ │ │ │
│ │ [Image x] → LLaVA(θ) → h_l → SAE → f_l(x) → probe p_φ │ │
│ │ │ │ │ │
│ │ KL(base) max entropy │ │
│ │ (don't drift) "confuse the probe" │ │
│ └────────────────────────────────────────────────────────────┘ │
└──────────────────────────────────────────────────────────────────┘
```
Iterate until the probe can no longer detect toilet-information in D_{A,¬B}.
---
## Method: Loss Function
**Combined objective for the model step:**
$$\mathcal{L}(\theta, \phi) = \underbrace{\mathcal{L}_\text{task}(\theta)}_{\text{capability}} + \lambda_\text{adv} \underbrace{\mathcal{L}_\text{suppress}(\theta)}_{\text{erase prior}} + \lambda_\text{ret} \underbrace{\mathcal{L}_\text{ret}(\theta)}_{\text{retain B}}$$
**L_task** — KL divergence from base model
*(computed on D_general + D_B; excludes D_{A,¬B})*
→ Preserves general behavior
**L_suppress** — on D_{A,¬B}:
`−E[ H(p_φ(f_l(x))) ]`
→ Push probe output to 0.5 (maximum uncertainty)
→ Model erases B-information from those images
**L_ret** — KL from base on D_B
→ Don't forget how to recognize toilets when they're actually there
**Probe step** — standard BCE
→ Probe maximally sensitive before each model update
→ Adversary keeps re-sharpening; model keeps erasing
**Config:** λ_adv = 20, λ_ret = 1, 5 probe steps per model step
---
## Method: Probe Labels
The probe is trained to detect **B-information** (toilet features):
| Image type | Scene (A) | Object (B) | Probe label |
|---|:---:|:---:|:---:|
| Bathroom + Toilet | ✓ | ✓ | **1** |
| Toilet only | ✗ | ✓ | **1** |
| Bathroom only | ✓ | ✗ | **0** |
| Neither | ✗ | ✗ | **0** |
> Mode: `object_only` — the probe learns to fire on *actual toilet presence*, not the bathroom-toilet association. This way, suppression of probe output on D_{A,¬B} means we erased B-features, not just the A-B correlation.
**Warmup:** Probe is pre-trained for 3 epochs on frozen model before adversarial training begins.
---
## Architecture Details
```
LLaVA-1.5-7B (frozen base weights)
+ LoRA adapter (r=16, α=32, all-linear)
Frozen SAE (pre-trained on LLaVA hidden states)
d_model → d_SAE (single shared encoder across all layers)
Linear probes (one per monitored layer)
d_SAE → 1 logit
```
**Training setup:**
- 20 epochs, batch size 8, grad accumulation 8 (effective = 64)
- lr = 1e-5 (model), 1e-4 (probe)
- Dataset: `pbcong/bathroom-toilet` (HuggingFace)
- General data: 1000 CC3M images for capability preservation
- Probe warmup: 3 epochs on frozen model
---
## 4. Evaluation Setup
**Task:** Generate a caption for each image, check for toilet keywords:
```
Keywords: ["toilet", "toilets", "Toilet", "Toilets", "commode", "lavatory", "latrine"]
```
**Metrics:**
| Metric | Definition | Applies to |
|---|---|---|
| **Hallucination rate** | % of captions that *incorrectly* mention toilet | D_{A,¬B} |
| **Miss rate** | % of captions that *fail* to mention toilet | D_B, toilet_only |
| **Keyword mention** | % of any mention | Overall |
**Models compared:**
- **Base** — LLaVA-1.5-7B (no LoRA)
- **LoRA** — same model + adversarially trained adapter
---
## 5. Preliminary Results
| Category | N | Base | LoRA | Change |
|---|:---:|:---:|:---:|:---:|
| **toilet_only** | 22 | 90.9% mention | 90.9% mention | ±0% miss |
| **bathroom_only** (D_{A,¬B}) | 1157 | 21.9% hallu | **2.9% hallu** | **↓ 18.9 pp** |
| **bathroom_toilet** (D_B) | 535 | 91.4% mention | 84.7% mention | ↓ 6.7 pp miss |
| **Overall** | 1714 | 44.5% mention | 29.6% mention | — |
**Interpretation:**
- **Suppression succeeds:** Hallucination on bathroom-only drops from 21.9% → **2.9%** (87% relative reduction)
- **Retention holds:** On images with actual toilets, miss rate only increases by 6.7 pp (from 8.6% → 15.3%)
- **Toilet recognition preserved:** `toilet_only` images are unaffected (90.9% both)
---
## Results — Visual Summary
```
Hallucination Rate (bathroom-only images)
Base model ████████████████████████ 21.9%
LoRA (ours) ███ 2.9%
↑ 87% reduction
Miss Rate (images with actual toilet)
Base model ████ 8.6%
LoRA (ours) ████████ 15.3%
↑ Small increase (acceptable trade-off)
Toilet recognition (toilet-only images)
Base model ██████████████████████████ 90.9%
LoRA (ours) ██████████████████████████ 90.9%
↑ Perfectly preserved
```
> The model learned to *not* invoke the bathroom→toilet prior without forgetting what a toilet looks like.
---
## 6. Discussion
**What the results suggest:**
- The adversarial probe successfully identifies and removes toilet-related representations from bathroom-only images
- LoRA (low-rank adapter) has enough capacity to edit internal representations while staying close to the base model
- The KL retention loss effectively prevents catastrophic forgetting of toilet recognition
**Open questions (next steps):**
1. **Probe accuracy** — How does probe accuracy evolve during training? Does it converge to ~50% (random) on D_{A,¬B}?
2. **Layer specificity** — Which layers changed most? Are early or late layers responsible for the prior?
3. **Generalization** — Does suppression transfer to other relations (e.g., kitchen→microwave)?
4. **Recovery test** — Can a few-shot fine-tune re-introduce the hallucination? (Test robustness vs output suppression)
5. **Covariance test** — Compute `Cov(A_features, B_features | Z_B=0)` to confirm the formal criterion
---
## What Makes This Different from Prior Work
**Prior work on VLM hallucination:**
- Instruction tuning / RLHF to reduce hallucinations
- Contrastive decoding at inference time
- Output-level supervision only
**Limitation:** These only patch the *surface behavior*. Internal representations still encode the spurious association.
**Our approach:**
- Targets **internal representations** directly via SAE features
- Uses an **adversarial probe** — removes exactly what a probe can detect
- Operates at **every layer** (probe across all decoder layers)
- **Formally defined** criterion for knowledge suppression
Knowledge suppression → harder to recover, more trustworthy model.
---
## Summary
**Problem:** VLMs hallucinate objects that are statistically correlated with visible context (prior knowledge hallucination). Output-level suppression is insufficient — the knowledge remains in internal layers.
**Method:** Adversarial training with frozen SAE probes. A linear probe is trained to detect object-B information from SAE latent features; the model (LoRA adapter) is trained to confuse the probe on images where B is absent.
**Result:** Hallucination rate on bathroom-only images: **21.9% → 2.9%** (87% reduction), with minimal impact on toilet recognition when toilets are actually present.
**Next:** Probe accuracy curves · Layer-wise analysis · Transfer to other relations · Recovery experiments
---
## Appendix: Training Configuration
| Parameter | Value |
|---|---|
| Base model | LLaVA-1.5-7B |
| LoRA rank / α | 16 / 32 |
| Epochs | 20 |
| Batch size (effective) | 64 (8 × 8 grad accum) |
| λ_adv | 20.0 |
| λ_ret | 1.0 |
| Probe steps per model step | 5 |
| Probe warmup epochs | 3 |
| Probe label mode | `object_only` |
| Exclude D_{A,¬B} from task KL | ✓ |
| General data (capability) | 1000 CC3M images |