| --- |
| marp: true |
| theme: default |
| paginate: true |
| style: | |
| section { |
| font-size: 22px; |
| padding: 40px 60px; |
| } |
| h1 { font-size: 1.8em; color: #1a1a2e; } |
| h2 { font-size: 1.3em; color: #16213e; border-bottom: 2px solid #e94560; padding-bottom: 6px; } |
| .highlight { background: #fff3cd; padding: 4px 8px; border-radius: 4px; } |
| .red { color: #e94560; font-weight: bold; } |
| .green { color: #27ae60; font-weight: bold; } |
| .blue { color: #2980b9; font-weight: bold; } |
| .box { |
| border: 1.5px solid #aaa; |
| border-radius: 8px; |
| padding: 12px 18px; |
| background: #f9f9f9; |
| margin: 8px 0; |
| } |
| .split { display: flex; gap: 30px; } |
| .split > div { flex: 1; } |
| table { width: 100%; border-collapse: collapse; } |
| th { background: #1a1a2e; color: white; padding: 8px; } |
| td { padding: 7px 10px; border: 1px solid #ccc; } |
| tr:nth-child(even) { background: #f5f5f5; } |
| code { background: #f0f0f0; padding: 2px 6px; border-radius: 3px; font-size: 0.9em; } |
| --- |
| |
| # Suppressing Prior Knowledge in Vision-Language Models |
| ## via Adversarial Sparse Autoencoder Training |
|
|
| **[Your Name]** |
| April 2026 |
|
|
| --- |
|
|
| ## Outline |
|
|
| 1. Motivation — Why VLMs hallucinate due to prior knowledge |
| 2. Problem Formulation — Output suppression vs. **Knowledge suppression** |
| 3. Method — Adversarial SAE-based representation editing |
| 4. Evaluation Setup |
| 5. Preliminary Results |
| 6. Discussion & Next Steps |
|
|
| --- |
|
|
| ## 1. VLMs Have Strong Priors — And That's a Problem |
|
|
| <div class="split"> |
| <div> |
|
|
| **Vision-Language Models (VLMs)** like LLaVA combine a vision encoder with a language model. |
|
|
| They are trained on massive web data, which encodes **statistical co-occurrences**: |
|
|
| > *"Bathrooms usually have toilets"* |
| > *"Kitchens usually have microwaves"* |
|
|
| These priors are useful — until the image contradicts them. |
|
|
| </div> |
| <div> |
|
|
| **Hallucination via prior knowledge:** |
|
|
| ``` |
| Image: [bathroom, NO toilet] |
| Prompt: "Describe this image." |
| |
| Base model output: |
| "A bathroom with a toilet, sink, |
| and tiled walls..." |
| ↑ HALLUCINATED |
| ``` |
|
|
| The model *invented* the toilet because the scene *predicts* one. |
|
|
| </div> |
| </div> |
|
|
| --- |
|
|
| ## What Makes This Hard: The Knowledge Is Encoded Deep |
|
|
| A naïve fix: fine-tune the model to not say "toilet" in bathroom images. |
|
|
| **Problem:** The *knowledge* still lives in the model's internal representations. |
|
|
| <br> |
|
|
| ``` |
| [Bathroom image] |
| │ |
| ┌─────────▼──────────────────────────────────┐ |
| │ LLaVA (fine-tuned to suppress output) │ |
| │ │ |
| │ Layer 5: h₅ ──► still encodes toilet? │ |
| │ Layer 15: h₁₅ ──► still encodes toilet? │ |
| │ Layer 31: h₃₁ ──► "a bathroom with a ..." │ |
| └─────────────────────────────────────────────┘ |
| ``` |
|
|
| - A simple fine-tune **only patches the output layer** |
| - A few more fine-tuning steps on a toilet dataset can **recover the hallucination** |
| - True suppression requires removing the prior from **every layer** |
|
|
| --- |
|
|
| ## 2. Two Levels of Suppression |
|
|
| <div class="split"> |
| <div> |
|
|
| ### Output Suppression |
| > The model **does not output** B in context A |
|
|
| - Check: `P(output mentions B | sees A, ¬B image)` is low |
| - **Internal layers may still encode B** |
| - Easily reversed by downstream fine-tuning |
| - Not robust to probing |
|
|
| </div> |
| <div> |
|
|
| ### Knowledge Suppression ✓ (Our Goal) |
| > No layer's representation contains B-information when B is visually absent |
|
|
| **Formal criterion:** For all layers *l* and all probes *p*: |
|
|
| ``` |
| p(f_l(x, θ_supp)) ⊥ Z_A(x) |
| given Z_B(x) = 0 |
| ``` |
|
|
| *"If B is not in the image, then no probe — trained on SAE features from any layer — can use A's presence to predict B."* |
|
|
| </div> |
| </div> |
|
|
| <br> |
|
|
| > **Key insight:** We use **Sparse Autoencoders (SAEs)** to measure and attack B-representations at every layer. |
|
|
| --- |
|
|
| ## 3. Tool: Sparse Autoencoders (SAEs) |
|
|
| SAEs decompose hidden states into **sparse, interpretable features**: |
|
|
| ``` |
| h_l(x) ∈ ℝ^d (dense hidden state at layer l) |
| │ |
| ▼ SAE encoder (frozen, pre-trained) |
| f_l(x) ∈ ℝ^D (sparse latent: D >> d, ~1% active) |
| │ |
| ▼ Each dimension ≈ one "concept" |
| f[42] → "toilet-related feature" |
| f[71] → "bathroom-related feature" |
| ``` |
|
|
| **Why SAEs?** |
| - Pre-trained SAE gives us a **universal feature basis** across all layers |
| - Probes on SAE features are more **interpretable** than probes on raw activations |
| - We can measure *which features* were suppressed |
|
|
| --- |
|
|
| ## Method Overview |
|
|
| **Setting:** Relation `bathroom → toilet` |
|
|
| Three disjoint data subsets: |
|
|
| | Subset | Description | Role | |
| |---|---|---| |
| | **D_{A,¬B}** | Bathroom images, **no toilet** | Suppression target | |
| | **D_B** | Images **with toilet** | Retention (don't forget B) | |
| | **D_general** | Random images (CC3M) | Capability preservation | |
| |
| **Goal:** Train LoRA adapter θ such that on D_{A,¬B}, no layer contains toilet-information — while preserving toilet recognition on D_B. |
| |
| --- |
| |
| ## Method: Adversarial Training Loop |
| |
| ``` |
| ┌──────────────────────────────────────────────────────────────────┐ |
| │ Alternating Optimization │ |
| │ │ |
| │ ┌─── PROBE STEP (fix θ, update φ) ──────────────────────────┐ │ |
| │ │ │ │ |
| │ │ [Image x] → LLaVA(θ, fixed) → h_l → SAE → f_l(x) │ │ |
| │ │ │ │ │ |
| │ │ probe p_φ │ │ |
| │ │ │ │ │ |
| │ │ φ ← minimize BCE(p_φ(f_l(x)), y_B) ▼ │ │ |
| │ │ "Find any toilet signal" │ │ |
| │ └────────────────────────────────────────────────────────────┘ │ |
| │ │ |
| │ ┌─── MODEL STEP (fix φ, update θ) ──────────────────────────┐ │ |
| │ │ │ │ |
| │ │ [Image x] → LLaVA(θ) → h_l → SAE → f_l(x) → probe p_φ │ │ |
| │ │ │ │ │ │ |
| │ │ KL(base) max entropy │ │ |
| │ │ (don't drift) "confuse the probe" │ │ |
| │ └────────────────────────────────────────────────────────────┘ │ |
| └──────────────────────────────────────────────────────────────────┘ |
| ``` |
| |
| Iterate until the probe can no longer detect toilet-information in D_{A,¬B}. |
| |
| --- |
| |
| ## Method: Loss Function |
| |
| **Combined objective for the model step:** |
| |
| $$\mathcal{L}(\theta, \phi) = \underbrace{\mathcal{L}_\text{task}(\theta)}_{\text{capability}} + \lambda_\text{adv} \underbrace{\mathcal{L}_\text{suppress}(\theta)}_{\text{erase prior}} + \lambda_\text{ret} \underbrace{\mathcal{L}_\text{ret}(\theta)}_{\text{retain B}}$$ |
| |
| <div class="split"> |
| <div> |
| |
| **L_task** — KL divergence from base model |
| *(computed on D_general + D_B; excludes D_{A,¬B})* |
| → Preserves general behavior |
|
|
| **L_suppress** — on D_{A,¬B}: |
| `−E[ H(p_φ(f_l(x))) ]` |
| → Push probe output to 0.5 (maximum uncertainty) |
| → Model erases B-information from those images |
| |
| </div> |
| <div> |
| |
| **L_ret** — KL from base on D_B |
| → Don't forget how to recognize toilets when they're actually there |
|
|
| **Probe step** — standard BCE |
| → Probe maximally sensitive before each model update |
| → Adversary keeps re-sharpening; model keeps erasing |
|
|
| **Config:** λ_adv = 20, λ_ret = 1, 5 probe steps per model step |
|
|
| </div> |
| </div> |
|
|
| --- |
|
|
| ## Method: Probe Labels |
|
|
| The probe is trained to detect **B-information** (toilet features): |
|
|
| | Image type | Scene (A) | Object (B) | Probe label | |
| |---|:---:|:---:|:---:| |
| | Bathroom + Toilet | ✓ | ✓ | **1** | |
| | Toilet only | ✗ | ✓ | **1** | |
| | Bathroom only | ✓ | ✗ | **0** | |
| | Neither | ✗ | ✗ | **0** | |
|
|
| > Mode: `object_only` — the probe learns to fire on *actual toilet presence*, not the bathroom-toilet association. This way, suppression of probe output on D_{A,¬B} means we erased B-features, not just the A-B correlation. |
|
|
| **Warmup:** Probe is pre-trained for 3 epochs on frozen model before adversarial training begins. |
|
|
| --- |
|
|
| ## Architecture Details |
|
|
| ``` |
| LLaVA-1.5-7B (frozen base weights) |
| + LoRA adapter (r=16, α=32, all-linear) |
| |
| Frozen SAE (pre-trained on LLaVA hidden states) |
| d_model → d_SAE (single shared encoder across all layers) |
| |
| Linear probes (one per monitored layer) |
| d_SAE → 1 logit |
| ``` |
|
|
| **Training setup:** |
| - 20 epochs, batch size 8, grad accumulation 8 (effective = 64) |
| - lr = 1e-5 (model), 1e-4 (probe) |
| - Dataset: `pbcong/bathroom-toilet` (HuggingFace) |
| - General data: 1000 CC3M images for capability preservation |
| - Probe warmup: 3 epochs on frozen model |
|
|
| --- |
|
|
| ## 4. Evaluation Setup |
|
|
| **Task:** Generate a caption for each image, check for toilet keywords: |
|
|
| ``` |
| Keywords: ["toilet", "toilets", "Toilet", "Toilets", "commode", "lavatory", "latrine"] |
| ``` |
|
|
| **Metrics:** |
|
|
| | Metric | Definition | Applies to | |
| |---|---|---| |
| | **Hallucination rate** | % of captions that *incorrectly* mention toilet | D_{A,¬B} | |
| | **Miss rate** | % of captions that *fail* to mention toilet | D_B, toilet_only | |
| | **Keyword mention** | % of any mention | Overall | |
| |
| **Models compared:** |
| - **Base** — LLaVA-1.5-7B (no LoRA) |
| - **LoRA** — same model + adversarially trained adapter |
| |
| --- |
| |
| ## 5. Preliminary Results |
| |
| <br> |
| |
| | Category | N | Base | LoRA | Change | |
| |---|:---:|:---:|:---:|:---:| |
| | **toilet_only** | 22 | 90.9% mention | 90.9% mention | ±0% miss | |
| | **bathroom_only** (D_{A,¬B}) | 1157 | 21.9% hallu | **2.9% hallu** | **↓ 18.9 pp** | |
| | **bathroom_toilet** (D_B) | 535 | 91.4% mention | 84.7% mention | ↓ 6.7 pp miss | |
| | **Overall** | 1714 | 44.5% mention | 29.6% mention | — | |
|
|
| <br> |
|
|
| **Interpretation:** |
|
|
| - **Suppression succeeds:** Hallucination on bathroom-only drops from 21.9% → **2.9%** (87% relative reduction) |
| - **Retention holds:** On images with actual toilets, miss rate only increases by 6.7 pp (from 8.6% → 15.3%) |
| - **Toilet recognition preserved:** `toilet_only` images are unaffected (90.9% both) |
|
|
| --- |
|
|
| ## Results — Visual Summary |
|
|
| ``` |
| Hallucination Rate (bathroom-only images) |
| |
| Base model ████████████████████████ 21.9% |
| LoRA (ours) ███ 2.9% |
| ↑ 87% reduction |
| |
| Miss Rate (images with actual toilet) |
| |
| Base model ████ 8.6% |
| LoRA (ours) ████████ 15.3% |
| ↑ Small increase (acceptable trade-off) |
| |
| Toilet recognition (toilet-only images) |
| |
| Base model ██████████████████████████ 90.9% |
| LoRA (ours) ██████████████████████████ 90.9% |
| ↑ Perfectly preserved |
| ``` |
|
|
| > The model learned to *not* invoke the bathroom→toilet prior without forgetting what a toilet looks like. |
|
|
| --- |
|
|
| ## 6. Discussion |
|
|
| **What the results suggest:** |
|
|
| - The adversarial probe successfully identifies and removes toilet-related representations from bathroom-only images |
| - LoRA (low-rank adapter) has enough capacity to edit internal representations while staying close to the base model |
| - The KL retention loss effectively prevents catastrophic forgetting of toilet recognition |
|
|
| **Open questions (next steps):** |
|
|
| 1. **Probe accuracy** — How does probe accuracy evolve during training? Does it converge to ~50% (random) on D_{A,¬B}? |
| 2. **Layer specificity** — Which layers changed most? Are early or late layers responsible for the prior? |
| 3. **Generalization** — Does suppression transfer to other relations (e.g., kitchen→microwave)? |
| 4. **Recovery test** — Can a few-shot fine-tune re-introduce the hallucination? (Test robustness vs output suppression) |
| 5. **Covariance test** — Compute `Cov(A_features, B_features | Z_B=0)` to confirm the formal criterion |
|
|
| --- |
|
|
| ## What Makes This Different from Prior Work |
|
|
| <div class="split"> |
| <div> |
|
|
| **Prior work on VLM hallucination:** |
| - Instruction tuning / RLHF to reduce hallucinations |
| - Contrastive decoding at inference time |
| - Output-level supervision only |
|
|
| **Limitation:** These only patch the *surface behavior*. Internal representations still encode the spurious association. |
|
|
| </div> |
| <div> |
|
|
| **Our approach:** |
| - Targets **internal representations** directly via SAE features |
| - Uses an **adversarial probe** — removes exactly what a probe can detect |
| - Operates at **every layer** (probe across all decoder layers) |
| - **Formally defined** criterion for knowledge suppression |
|
|
| Knowledge suppression → harder to recover, more trustworthy model. |
|
|
| </div> |
| </div> |
|
|
| --- |
|
|
| ## Summary |
|
|
| <div class="box"> |
|
|
| **Problem:** VLMs hallucinate objects that are statistically correlated with visible context (prior knowledge hallucination). Output-level suppression is insufficient — the knowledge remains in internal layers. |
|
|
| **Method:** Adversarial training with frozen SAE probes. A linear probe is trained to detect object-B information from SAE latent features; the model (LoRA adapter) is trained to confuse the probe on images where B is absent. |
|
|
| **Result:** Hallucination rate on bathroom-only images: **21.9% → 2.9%** (87% reduction), with minimal impact on toilet recognition when toilets are actually present. |
|
|
| </div> |
|
|
| **Next:** Probe accuracy curves · Layer-wise analysis · Transfer to other relations · Recovery experiments |
|
|
| --- |
|
|
| ## Appendix: Training Configuration |
|
|
| | Parameter | Value | |
| |---|---| |
| | Base model | LLaVA-1.5-7B | |
| | LoRA rank / α | 16 / 32 | |
| | Epochs | 20 | |
| | Batch size (effective) | 64 (8 × 8 grad accum) | |
| | λ_adv | 20.0 | |
| | λ_ret | 1.0 | |
| | Probe steps per model step | 5 | |
| | Probe warmup epochs | 3 | |
| | Probe label mode | `object_only` | |
| | Exclude D_{A,¬B} from task KL | ✓ | |
| | General data (capability) | 1000 CC3M images | |
| |