Repo Setup & Working Hypotheses
Object hallucination in LLaVA-1.5-7B, flagship case study bathroom → toilet: suppress mentions of toilet when the image lacks one, preserve them when a toilet is actually present. This doc records (a) the repo layout and (b) the modeling hypotheses the method and the analysis silently rely on, with file evidence and the failure mode if each is false.
Legend: [CORE] = result collapses if false · [AUX] = degrades but recoverable.
1. Repo setup
Three directories carry the project; everything else is supporting/legacy.
| Dir | Role | Env |
|---|---|---|
training_method/ |
The method — bathroom→toilet v2 training | venv /data/caotue/multilayer-sae/.venv |
mechanistic_interp/ |
Post-hoc analysis — look inside the model | conda baodq_hal |
experiment/ |
Tuning & measuring — configs, baselines, eval | conda baodq_hal |
Supporting / mostly-unrelated: model/ (hooked LLaVA backbones), sae/ + training/
(SAE library — earlier infra, current method is probe+LoRA not SAE), probe/ (older
SAE-feature linear probes), evaluation/eval_nullu_val.py (aux harness).
The method (training_method/), two DDP stages:
- Probe pretrain (
train_probe_gen.py) — gen-scope attention-pooling sequence probe (sequence_probe.py:SequenceLayerProbes) on the residual stream over generated-caption tokens. - Adversarial LoRA suppression (
finetune_adv_gen_resume.pyviarun_finetune_adv_gen_refined.sh) — LoRA fine-tune that adversarially drives the probe toward "absent" on the suppress set while anchoring retain/present sets (finetune_adv.py). Entry:run_bathroom_toilet_v2.sh.
training_method/ imports byte-identical helpers from experiment.* / sae.*
(config, data.datasets, evaluation.metrics, training.{gen_features,preference},
sae.Training_Utils) and only vendors the files that differ. See
training_method/README.md.
The analysis (mechanistic_interp/): latent sequence probes
(train_probe_latent.py, ckpts under /data/caotue/latent_probes/seqprobes_*),
causal-influence maps (gradient_ascent.py, integrated_gradient.py), attribution
patching (attribution_patching.py), probe scoring (probe_scoring.py). Probes are
trained in two places: here (for analysis) and in training_method/ (the method's
own stage-1 probe).
2. Method hypotheses (training_method/)
Labeling
- [CORE] Binary, per-image, ground-truth label.
y=1 ⇔ object present in the image(object_onlymode). Evidence:finetune_adv.py:377-388,experiment/data/datasets.py:294-305. It is image GT, not caption mention (unless--label_from_mention). If false: if hallucination is a generation-distribution effect not reflected in GT, the probe optimizes the wrong target. - [CORE] Label shared across all generated-caption tokens; probe reads only the
gen-token window. Evidence:
train_probe_gen.py:435-436(pool_tokens="gen"). If false: an object signal carried in the prompt/image patches (not the continuation) is invisible to the probe.
What the probe is
- [CORE] High probe logit ⇔ "object features present" ⇔ the thing to suppress;
probe direction == suppression direction. Evidence:
finetune_adv_gen.py:488-493. If false: if the probe tracks the scene (bathroom) or a side-channel, suppression removes the wrong representation. - [AUX] Attention pooling locates the concept in the sequence
(
sequence_probe.py:31-79). If false: non-salient hallucinated tokens are missed.
Causal / loss structure
- [CORE] Three disjoint, separable categories — suppress (scene ∧ ¬object), retain
(object present), neutral — and the scene→object prior is a separable direction
removable without destroying legitimate object detection. Evidence:
finetune_adv_gen.py:393-394,finetune_adv.py:237-264. If false: an entangled prior cannot be excised cleanly; suppression damages real detection. - [AUX] Driving probe
p→0on suppress rows erases the hallucination at inference; greedy training captions cover inference behavior. Evidence:finetune_adv_gen.py:510-517,:398-407. If false: probability mass shifts to the next token instead of the prior being removed. - Conditional: the v2 recipe uses a raw-activation probe, so SAE-dictionary
assumptions (monosemantic object feature, in-distribution JumpReLU threshold) do not
apply to this run — they matter only if
raw_activation_probe=false.
3. Analysis hypotheses (mechanistic_interp/)
Probe semantics
- [CORE] Binary, per-image label shared across all tokens AND all layers. Each
layer's probe answers "does layer-
lresidual encode the concept?". Evidence:train_probe_latent.py:664,764,416-417. If false: a strongly progressive (layer-dependent) encoding makes one shared label incoherent. - [CORE] High toilet score on a bathroom-only image = the model internally represents
the absent toilet = the hallucination signal. The central interpretive claim.
Evidence:
probe_scoring.py:6-12. If false: the probe reads an orthogonal axis, or "bathroom-only" images contain real visual toilet cues (pipes/tiles) → confound. - [AUX]
halluc_mode=keeplabels bathroom-only val images by the BASE model's output, not GT. Evidence:train_probe_latent.py:747-751. That metric measures base-output agreement, not ground truth.
Steering / causality
- [CORE] Pushing the residual along the bathroom-probe direction at layer
lcausally raises the toilet readout downstream ⇒ a bathroom→toilet mechanism. Evidence:gradient_ascent.py:3-29. Assumes the probe gradient aligns with the model's real bathroom direction and the patch propagates cleanly. Status: partially contradicted — small normalized-IG steps barely move the toilet readout; only the full(x−b)displacement does (see §5). - [AUX] IG baseline = mean negative-validation residual = a "no-concept" reference
(
integrated_gradient.py:11-16). Assumes negatives are truly concept-absent and a single broadcast vector is a fair zero-point. - [AUX]
mean-steer: one image-independent direction represents bathroom for all images (integrated_gradient.py:455-465). - [AUX]
α·‖h‖makes steps comparable across layers; raw Δσ (not/α) is the causal effect magnitude (gradient_ascent.py:18-23). - [AUX] Attribution patching assumes local linearity + clean feature isolation
(
attribution_patching.py:113).
4. Cross-cutting hypotheses
- [CORE] Probes trained on base (or 4-variant-pooled) transfer as valid readouts on
edited models (lora/nullu/efuf) — the residual geometry stays aligned. Evidence:
gradient_ascent.py:184,probe_scoring.py. If false: base-vs-edit comparisons are probe-misalignment artifacts, not mechanism. 4-variant pooling is the hedge. - [AUX] Forced-text holds the output constant to isolate representation from
generation (
compare_baselines.py:15-16).
5. Judge's verdict — most load-bearing & most fragile
- Probe direction = the causal hallucination axis — used as both suppression target (method) and interpretive readout (analysis). Everything inherits it. Most load-bearing, least directly verified.
- bathroom→toilet is a separable, steerable direction. Current gradient-ascent / IG
figures show small-step steering barely moves the readout; only the full
(x−b)displacement does. Partially contradicted — the top candidate for a clean falsification test. - Cross-variant probe transfer — quietly assumed everywhere; if false, all edit comparisons are confounded.
Suggested falsification tests
- (1) Train a random-direction / shuffled-label probe; confirm it does NOT yield the same suppression or the same downstream toilet response. Causal-scrub the probe axis and check the hallucination rate is unchanged off-axis.
- (2) Sweep α up to
‖x−b‖along the normalized IG direction; if the late-layer band only appears near full magnitude, the effect is displacement-threshold, not a smooth steerable direction. - (3) Re-fit the probe per variant and compare to the transferred base probe (AUC + direction cosine); large divergence ⇒ transfer assumption broken.