siddharthmb's picture
Upload probes, handles, and attention diagnostics
acbbbdc verified
|
Raw
History Blame Contribute Delete
4.07 kB
---
license: apache-2.0
language:
- en
tags:
- interpretability
- probes
- steering
- multi-agent
- qwen3
library_name: pytorch
---
# Multi-Agent Verification Probes & Direction Handles
Trained linear probes, difference-in-means direction handles, and attention
diagnostics from an interpretability study of failure modes in multi-agent fact
verification (Qwen3-32B verifier + Qwen3-8B subagents on AVeriTeC claims).
These artifacts were trained on the residual-stream activations released in the
companion dataset repo:
[siddharthmb/multiagent-verification-failure-modes](https://huggingface.co/datasets/siddharthmb/multiagent-verification-failure-modes).
Code: [Sid-MB/mats-gf-multiagent-failure-modes](https://github.com/Sid-MB/mats-gf-multiagent-failure-modes).
## Contents
- `exp2/` β€” transmission-loss interpretability (Qwen3-8B subagent / Qwen3-32B verifier):
- `probes/h1/` β€” per-layer (3–35, step 4) linear probes predicting
gold-evidence-in-context at the subagent message-final token position, at two
transmission-loss thresholds, plus bag-of-words shortcut baselines
(`bow_*.parquet`) and per-layer metrics.
- `probes/h2/` β€” per-(layer, round) 4-class verdict-commitment probes at the
Qwen3-32B verifier think-close position (layers 3–63, step 4), with
round-shuffled and label-shuffled controls and row-level predictions.
- `probes/dense/` β€” metrics for the dense tagged-position probe sweeps.
- `handles/` β€” conveyed-minus-unconveyed difference-in-means "expression"
directions at Qwen3-8B layer 31 (strict and lenient thresholds), as `.pt`
with a JSON manifest. Held-out conveyed-vs-lost AUC 0.765. Note: steering
with this direction did **not** causally force conveyance (a null result);
it is a predictive direction, not a validated intervention handle.
- `exp5/` β€” trust interpretability (Qwen3-32B verifier):
- `probes/` β€” content-quality and source-tier probes per layer x round-context,
in four families (`gold_first_blind`, `gold_first_told`, `strong_first_blind`,
`strong_first_told`), with standardizers.
- `handles/` β€” told-minus-blind disclosure directions per layer (3–63, step 4)
and round-context, with `manifest.json` (per-direction norm, matched-group
counts, separation AUC). The verdict-context **layer-51** direction predicts
tier-tracking (AUC 0.82) and causally shifted single verdicts in both
directions vs a norm-matched sham when steered.
- `exp6/` β€” attention-pickup diagnostics (Qwen3-8B subagent):
- `attention/span_mass.parquet` β€” teacher-forced attention mass on the gold
evidence span per (turn, layer, phase) for 830 fact-check turns, comparing
failed vs conveyed transmission-loss turns.
- `attention/exemplars/*.npz` β€” 20 full attention maps for exemplar turns.
- exp6 saved no direction handles of its own (its steering interventions reused
exp2's expression direction), so there is no `exp6/handles/`.
## Formats and loading
- Probe checkpoints (`*.pt`) are [goodfire-core](https://github.com/goodfire-ai)
`LinearProbe` checkpoints (a state dict plus normalization statistics);
`torch.load(path, map_location="cpu")` shows the raw contents, or use
`LinearProbe.from_checkpoint` with goodfire-core installed. Inputs are
standardized residual-stream activations at the position/layer named in the file.
- Direction handles (`*.pt`) are plain tensors (residual-stream-dimensional
vectors) with accompanying JSON manifests describing layer, context, norm,
and evaluation numbers.
- `attention/exemplars/*.npz` are numpy archives of per-head attention maps.
`checksums.json` at the repo root maps every file to its sha256 and size.
## Provenance
Trained on our own harvested activations of Qwen/Qwen3-8B and Qwen/Qwen3-32B
(Apache 2.0) replayed over the episode corpus in the companion dataset repo.
Weights released under Apache 2.0. Note the underlying episode *text* corpus is
CC BY-NC 4.0 (AVeriTeC-derived); that license applies to the dataset repo, not
to these probe weights.