--- license: apache-2.0 language: - en tags: - interpretability - probes - steering - multi-agent - qwen3 library_name: pytorch --- # Multi-Agent Verification Probes & Direction Handles Trained linear probes, difference-in-means direction handles, and attention diagnostics from an interpretability study of failure modes in multi-agent fact verification (Qwen3-32B verifier + Qwen3-8B subagents on AVeriTeC claims). These artifacts were trained on the residual-stream activations released in the companion dataset repo: [siddharthmb/multiagent-verification-failure-modes](https://huggingface.co/datasets/siddharthmb/multiagent-verification-failure-modes). Code: [Sid-MB/mats-gf-multiagent-failure-modes](https://github.com/Sid-MB/mats-gf-multiagent-failure-modes). ## Contents - `exp2/` — transmission-loss interpretability (Qwen3-8B subagent / Qwen3-32B verifier): - `probes/h1/` — per-layer (3–35, step 4) linear probes predicting gold-evidence-in-context at the subagent message-final token position, at two transmission-loss thresholds, plus bag-of-words shortcut baselines (`bow_*.parquet`) and per-layer metrics. - `probes/h2/` — per-(layer, round) 4-class verdict-commitment probes at the Qwen3-32B verifier think-close position (layers 3–63, step 4), with round-shuffled and label-shuffled controls and row-level predictions. - `probes/dense/` — metrics for the dense tagged-position probe sweeps. - `handles/` — conveyed-minus-unconveyed difference-in-means "expression" directions at Qwen3-8B layer 31 (strict and lenient thresholds), as `.pt` with a JSON manifest. Held-out conveyed-vs-lost AUC 0.765. Note: steering with this direction did **not** causally force conveyance (a null result); it is a predictive direction, not a validated intervention handle. - `exp5/` — trust interpretability (Qwen3-32B verifier): - `probes/` — content-quality and source-tier probes per layer x round-context, in four families (`gold_first_blind`, `gold_first_told`, `strong_first_blind`, `strong_first_told`), with standardizers. - `handles/` — told-minus-blind disclosure directions per layer (3–63, step 4) and round-context, with `manifest.json` (per-direction norm, matched-group counts, separation AUC). The verdict-context **layer-51** direction predicts tier-tracking (AUC 0.82) and causally shifted single verdicts in both directions vs a norm-matched sham when steered. - `exp6/` — attention-pickup diagnostics (Qwen3-8B subagent): - `attention/span_mass.parquet` — teacher-forced attention mass on the gold evidence span per (turn, layer, phase) for 830 fact-check turns, comparing failed vs conveyed transmission-loss turns. - `attention/exemplars/*.npz` — 20 full attention maps for exemplar turns. - exp6 saved no direction handles of its own (its steering interventions reused exp2's expression direction), so there is no `exp6/handles/`. ## Formats and loading - Probe checkpoints (`*.pt`) are [goodfire-core](https://github.com/goodfire-ai) `LinearProbe` checkpoints (a state dict plus normalization statistics); `torch.load(path, map_location="cpu")` shows the raw contents, or use `LinearProbe.from_checkpoint` with goodfire-core installed. Inputs are standardized residual-stream activations at the position/layer named in the file. - Direction handles (`*.pt`) are plain tensors (residual-stream-dimensional vectors) with accompanying JSON manifests describing layer, context, norm, and evaluation numbers. - `attention/exemplars/*.npz` are numpy archives of per-head attention maps. `checksums.json` at the repo root maps every file to its sha256 and size. ## Provenance Trained on our own harvested activations of Qwen/Qwen3-8B and Qwen/Qwen3-32B (Apache 2.0) replayed over the episode corpus in the companion dataset repo. Weights released under Apache 2.0. Note the underlying episode *text* corpus is CC BY-NC 4.0 (AVeriTeC-derived); that license applies to the dataset repo, not to these probe weights.