# Repo Setup & Working Hypotheses Object hallucination in **LLaVA-1.5-7B**, flagship case study **bathroom → toilet**: suppress mentions of *toilet* when the image lacks one, preserve them when a toilet is actually present. This doc records (a) the repo layout and (b) the modeling hypotheses the method and the analysis silently rely on, with file evidence and the failure mode if each is false. > Legend: **[CORE]** = result collapses if false · **[AUX]** = degrades but recoverable. --- ## 1. Repo setup Three directories carry the project; everything else is supporting/legacy. | Dir | Role | Env | |-----|------|-----| | `training_method/` | **The method** — bathroom→toilet v2 training | venv `/data/caotue/multilayer-sae/.venv` | | `mechanistic_interp/` | **Post-hoc analysis** — look inside the model | conda `baodq_hal` | | `experiment/` | **Tuning & measuring** — configs, baselines, eval | conda `baodq_hal` | Supporting / mostly-unrelated: `model/` (hooked LLaVA backbones), `sae/` + `training/` (SAE library — earlier infra, current method is probe+LoRA not SAE), `probe/` (older SAE-feature linear probes), `evaluation/eval_nullu_val.py` (aux harness). **The method (`training_method/`), two DDP stages:** 1. **Probe pretrain** (`train_probe_gen.py`) — gen-scope **attention-pooling sequence probe** (`sequence_probe.py:SequenceLayerProbes`) on the residual stream over generated-caption tokens. 2. **Adversarial LoRA suppression** (`finetune_adv_gen_resume.py` via `run_finetune_adv_gen_refined.sh`) — LoRA fine-tune that adversarially drives the probe toward "absent" on the suppress set while anchoring retain/present sets (`finetune_adv.py`). Entry: `run_bathroom_toilet_v2.sh`. `training_method/` imports byte-identical helpers from `experiment.*` / `sae.*` (`config`, `data.datasets`, `evaluation.metrics`, `training.{gen_features,preference}`, `sae.Training_Utils`) and only vendors the files that differ. See `training_method/README.md`. **The analysis (`mechanistic_interp/`):** latent sequence probes (`train_probe_latent.py`, ckpts under `/data/caotue/latent_probes/seqprobes_*`), causal-influence maps (`gradient_ascent.py`, `integrated_gradient.py`), attribution patching (`attribution_patching.py`), probe scoring (`probe_scoring.py`). Probes are trained in **two** places: here (for analysis) and in `training_method/` (the method's own stage-1 probe). --- ## 2. Method hypotheses (`training_method/`) ### Labeling - **[CORE] Binary, per-image, ground-truth label.** `y=1 ⇔ object present in the image` (`object_only` mode). Evidence: `finetune_adv.py:377-388`, `experiment/data/datasets.py:294-305`. It is *image GT*, not caption mention (unless `--label_from_mention`). *If false:* if hallucination is a generation-distribution effect not reflected in GT, the probe optimizes the wrong target. - **[CORE] Label shared across all generated-caption tokens; probe reads only the gen-token window.** Evidence: `train_probe_gen.py:435-436` (`pool_tokens="gen"`). *If false:* an object signal carried in the prompt/image patches (not the continuation) is invisible to the probe. ### What the probe is - **[CORE] High probe logit ⇔ "object features present" ⇔ the thing to suppress; probe direction == suppression direction.** Evidence: `finetune_adv_gen.py:488-493`. *If false:* if the probe tracks the *scene* (bathroom) or a side-channel, suppression removes the wrong representation. - **[AUX] Attention pooling locates the concept in the sequence** (`sequence_probe.py:31-79`). *If false:* non-salient hallucinated tokens are missed. ### Causal / loss structure - **[CORE] Three disjoint, separable categories** — suppress (scene ∧ ¬object), retain (object present), neutral — **and the scene→object prior is a separable direction** removable *without* destroying legitimate object detection. Evidence: `finetune_adv_gen.py:393-394`, `finetune_adv.py:237-264`. *If false:* an entangled prior cannot be excised cleanly; suppression damages real detection. - **[AUX] Driving probe `p→0` on suppress rows erases the hallucination at inference; greedy training captions cover inference behavior.** Evidence: `finetune_adv_gen.py:510-517`, `:398-407`. *If false:* probability mass shifts to the next token instead of the prior being removed. - **Conditional:** the v2 recipe uses a **raw-activation** probe, so SAE-dictionary assumptions (monosemantic object feature, in-distribution JumpReLU threshold) **do not apply to this run** — they matter only if `raw_activation_probe=false`. --- ## 3. Analysis hypotheses (`mechanistic_interp/`) ### Probe semantics - **[CORE] Binary, per-image label shared across all tokens AND all layers.** Each layer's probe answers "does layer-`l` residual encode the concept?". Evidence: `train_probe_latent.py:664,764,416-417`. *If false:* a strongly *progressive* (layer-dependent) encoding makes one shared label incoherent. - **[CORE] High toilet score on a bathroom-only image = the model internally represents the absent toilet = the hallucination signal.** The central interpretive claim. Evidence: `probe_scoring.py:6-12`. *If false:* the probe reads an orthogonal axis, or "bathroom-only" images contain real visual toilet cues (pipes/tiles) → confound. - **[AUX] `halluc_mode=keep` labels bathroom-only val images by the BASE model's output, not GT.** Evidence: `train_probe_latent.py:747-751`. That metric measures base-output agreement, not ground truth. ### Steering / causality - **[CORE] Pushing the residual along the bathroom-probe direction at layer `l` *causally* raises the toilet readout downstream ⇒ a bathroom→toilet mechanism.** Evidence: `gradient_ascent.py:3-29`. Assumes the probe gradient aligns with the model's real bathroom direction and the patch propagates cleanly. *Status:* **partially contradicted** — small normalized-IG steps barely move the toilet readout; only the full `(x−b)` displacement does (see §5). - **[AUX] IG baseline = mean negative-validation residual = a "no-concept" reference** (`integrated_gradient.py:11-16`). Assumes negatives are truly concept-absent and a single broadcast vector is a fair zero-point. - **[AUX] `mean-steer`: one image-independent direction represents bathroom for all images** (`integrated_gradient.py:455-465`). - **[AUX] `α·‖h‖` makes steps comparable across layers; raw Δσ (not `/α`) is the causal effect magnitude** (`gradient_ascent.py:18-23`). - **[AUX] Attribution patching assumes local linearity + clean feature isolation** (`attribution_patching.py:113`). --- ## 4. Cross-cutting hypotheses - **[CORE] Probes trained on base (or 4-variant-pooled) transfer as valid readouts on edited models (lora/nullu/efuf)** — the residual geometry stays aligned. Evidence: `gradient_ascent.py:184`, `probe_scoring.py`. *If false:* base-vs-edit comparisons are probe-misalignment artifacts, not mechanism. 4-variant pooling is the hedge. - **[AUX] Forced-text holds the output constant to isolate representation from generation** (`compare_baselines.py:15-16`). --- ## 5. Judge's verdict — most load-bearing & most fragile 1. **Probe direction = the causal hallucination axis** — used as *both* suppression target (method) and interpretive readout (analysis). Everything inherits it. Most load-bearing, least directly verified. 2. **bathroom→toilet is a separable, steerable direction.** Current gradient-ascent / IG figures show small-step steering barely moves the readout; only the full `(x−b)` displacement does. **Partially contradicted** — the top candidate for a clean falsification test. 3. **Cross-variant probe transfer** — quietly assumed everywhere; if false, all edit comparisons are confounded. ### Suggested falsification tests - **(1)** Train a *random-direction* / shuffled-label probe; confirm it does NOT yield the same suppression or the same downstream toilet response. Causal-scrub the probe axis and check the hallucination rate is unchanged off-axis. - **(2)** Sweep α up to `‖x−b‖` along the normalized IG direction; if the late-layer band only appears near full magnitude, the effect is displacement-threshold, not a smooth steerable direction. - **(3)** Re-fit the probe per variant and compare to the transferred base probe (AUC + direction cosine); large divergence ⇒ transfer assumption broken.