hallucination / docs /Hypotheses.md
ToiTenBao's picture
Upload hallucination folder
a2ffd07 verified
|
Raw
History Blame Contribute Delete
8.49 kB
# Repo Setup & Working Hypotheses
Object hallucination in **LLaVA-1.5-7B**, flagship case study **bathroom → toilet**:
suppress mentions of *toilet* when the image lacks one, preserve them when a toilet
is actually present. This doc records (a) the repo layout and (b) the modeling
hypotheses the method and the analysis silently rely on, with file evidence and the
failure mode if each is false.
> Legend: **[CORE]** = result collapses if false · **[AUX]** = degrades but recoverable.
---
## 1. Repo setup
Three directories carry the project; everything else is supporting/legacy.
| Dir | Role | Env |
|-----|------|-----|
| `training_method/` | **The method** — bathroom→toilet v2 training | venv `/data/caotue/multilayer-sae/.venv` |
| `mechanistic_interp/` | **Post-hoc analysis** — look inside the model | conda `baodq_hal` |
| `experiment/` | **Tuning & measuring** — configs, baselines, eval | conda `baodq_hal` |
Supporting / mostly-unrelated: `model/` (hooked LLaVA backbones), `sae/` + `training/`
(SAE library — earlier infra, current method is probe+LoRA not SAE), `probe/` (older
SAE-feature linear probes), `evaluation/eval_nullu_val.py` (aux harness).
**The method (`training_method/`), two DDP stages:**
1. **Probe pretrain** (`train_probe_gen.py`) — gen-scope **attention-pooling sequence
probe** (`sequence_probe.py:SequenceLayerProbes`) on the residual stream over
generated-caption tokens.
2. **Adversarial LoRA suppression** (`finetune_adv_gen_resume.py` via
`run_finetune_adv_gen_refined.sh`) — LoRA fine-tune that adversarially drives the
probe toward "absent" on the suppress set while anchoring retain/present sets
(`finetune_adv.py`). Entry: `run_bathroom_toilet_v2.sh`.
`training_method/` imports byte-identical helpers from `experiment.*` / `sae.*`
(`config`, `data.datasets`, `evaluation.metrics`, `training.{gen_features,preference}`,
`sae.Training_Utils`) and only vendors the files that differ. See
`training_method/README.md`.
**The analysis (`mechanistic_interp/`):** latent sequence probes
(`train_probe_latent.py`, ckpts under `/data/caotue/latent_probes/seqprobes_*`),
causal-influence maps (`gradient_ascent.py`, `integrated_gradient.py`), attribution
patching (`attribution_patching.py`), probe scoring (`probe_scoring.py`). Probes are
trained in **two** places: here (for analysis) and in `training_method/` (the method's
own stage-1 probe).
---
## 2. Method hypotheses (`training_method/`)
### Labeling
- **[CORE] Binary, per-image, ground-truth label.** `y=1 ⇔ object present in the image`
(`object_only` mode). Evidence: `finetune_adv.py:377-388`, `experiment/data/datasets.py:294-305`.
It is *image GT*, not caption mention (unless `--label_from_mention`).
*If false:* if hallucination is a generation-distribution effect not reflected in GT,
the probe optimizes the wrong target.
- **[CORE] Label shared across all generated-caption tokens; probe reads only the
gen-token window.** Evidence: `train_probe_gen.py:435-436` (`pool_tokens="gen"`).
*If false:* an object signal carried in the prompt/image patches (not the continuation)
is invisible to the probe.
### What the probe is
- **[CORE] High probe logit ⇔ "object features present" ⇔ the thing to suppress;
probe direction == suppression direction.** Evidence: `finetune_adv_gen.py:488-493`.
*If false:* if the probe tracks the *scene* (bathroom) or a side-channel, suppression
removes the wrong representation.
- **[AUX] Attention pooling locates the concept in the sequence**
(`sequence_probe.py:31-79`). *If false:* non-salient hallucinated tokens are missed.
### Causal / loss structure
- **[CORE] Three disjoint, separable categories** — suppress (scene ∧ ¬object), retain
(object present), neutral — **and the scene→object prior is a separable direction**
removable *without* destroying legitimate object detection. Evidence:
`finetune_adv_gen.py:393-394`, `finetune_adv.py:237-264`.
*If false:* an entangled prior cannot be excised cleanly; suppression damages real detection.
- **[AUX] Driving probe `p→0` on suppress rows erases the hallucination at inference;
greedy training captions cover inference behavior.** Evidence:
`finetune_adv_gen.py:510-517`, `:398-407`. *If false:* probability mass shifts to the
next token instead of the prior being removed.
- **Conditional:** the v2 recipe uses a **raw-activation** probe, so SAE-dictionary
assumptions (monosemantic object feature, in-distribution JumpReLU threshold) **do not
apply to this run** — they matter only if `raw_activation_probe=false`.
---
## 3. Analysis hypotheses (`mechanistic_interp/`)
### Probe semantics
- **[CORE] Binary, per-image label shared across all tokens AND all layers.** Each
layer's probe answers "does layer-`l` residual encode the concept?". Evidence:
`train_probe_latent.py:664,764,416-417`. *If false:* a strongly *progressive*
(layer-dependent) encoding makes one shared label incoherent.
- **[CORE] High toilet score on a bathroom-only image = the model internally represents
the absent toilet = the hallucination signal.** The central interpretive claim.
Evidence: `probe_scoring.py:6-12`. *If false:* the probe reads an orthogonal axis, or
"bathroom-only" images contain real visual toilet cues (pipes/tiles) → confound.
- **[AUX] `halluc_mode=keep` labels bathroom-only val images by the BASE model's output,
not GT.** Evidence: `train_probe_latent.py:747-751`. That metric measures base-output
agreement, not ground truth.
### Steering / causality
- **[CORE] Pushing the residual along the bathroom-probe direction at layer `l`
*causally* raises the toilet readout downstream ⇒ a bathroom→toilet mechanism.**
Evidence: `gradient_ascent.py:3-29`. Assumes the probe gradient aligns with the model's
real bathroom direction and the patch propagates cleanly.
*Status:* **partially contradicted** — small normalized-IG steps barely move the toilet
readout; only the full `(x−b)` displacement does (see §5).
- **[AUX] IG baseline = mean negative-validation residual = a "no-concept" reference**
(`integrated_gradient.py:11-16`). Assumes negatives are truly concept-absent and a
single broadcast vector is a fair zero-point.
- **[AUX] `mean-steer`: one image-independent direction represents bathroom for all
images** (`integrated_gradient.py:455-465`).
- **[AUX] `α·‖h‖` makes steps comparable across layers; raw Δσ (not `/α`) is the causal
effect magnitude** (`gradient_ascent.py:18-23`).
- **[AUX] Attribution patching assumes local linearity + clean feature isolation**
(`attribution_patching.py:113`).
---
## 4. Cross-cutting hypotheses
- **[CORE] Probes trained on base (or 4-variant-pooled) transfer as valid readouts on
edited models (lora/nullu/efuf)** — the residual geometry stays aligned. Evidence:
`gradient_ascent.py:184`, `probe_scoring.py`. *If false:* base-vs-edit comparisons are
probe-misalignment artifacts, not mechanism. 4-variant pooling is the hedge.
- **[AUX] Forced-text holds the output constant to isolate representation from
generation** (`compare_baselines.py:15-16`).
---
## 5. Judge's verdict — most load-bearing & most fragile
1. **Probe direction = the causal hallucination axis** — used as *both* suppression
target (method) and interpretive readout (analysis). Everything inherits it. Most
load-bearing, least directly verified.
2. **bathroom→toilet is a separable, steerable direction.** Current gradient-ascent / IG
figures show small-step steering barely moves the readout; only the full `(x−b)`
displacement does. **Partially contradicted** — the top candidate for a clean
falsification test.
3. **Cross-variant probe transfer** — quietly assumed everywhere; if false, all
edit comparisons are confounded.
### Suggested falsification tests
- **(1)** Train a *random-direction* / shuffled-label probe; confirm it does NOT yield
the same suppression or the same downstream toilet response. Causal-scrub the probe
axis and check the hallucination rate is unchanged off-axis.
- **(2)** Sweep α up to `‖x−b‖` along the normalized IG direction; if the late-layer
band only appears near full magnitude, the effect is displacement-threshold, not a
smooth steerable direction.
- **(3)** Re-fit the probe per variant and compare to the transferred base probe
(AUC + direction cosine); large divergence ⇒ transfer assumption broken.