Transcoders for the C2-2B-pt conjunctive-backdoor organism
Per-layer transcoders trained on the MLP activations of
Ftm23/backdoor-gemma2-2b-2pair-hate-pt,
used to build attribution graphs with
circuit-tracer.
These are not reproducible from the training code alone in any short run β they are the artifact the attribution results depend on, published so that work can be repeated without retraining them.
Contents
| path | layers | how trained |
|---|---|---|
from_scratch/layer_{18..22}.safetensors |
18β22 | trained from scratch on this organism's activations |
warm/layer_{23,24,25}_ep2.safetensors |
23β25 | warm-started from GemmaScope, 2 epochs |
Reconstruction quality (post-LN FVU, k=64, held-out)
| layer | this transcoder | off-the-shelf GemmaScope |
|---|---|---|
| 18 | 0.360 | 0.651 |
| 19 | 0.403 | 0.690 |
| 20 | 0.442 | 0.722 |
| 21 | 0.452 | 0.698 |
| 22 | 0.489 | 0.768 |
Lower is better. The from-scratch transcoders reconstruct this organism's activations substantially better than the off-the-shelf ones at every layer measured.
How they are used
The attribution runs build a ReplacementModel by splicing: off-the-shelf
GemmaScope transcoders for layers 0β22, and the warm/ transcoders here for layers
23β25 β the band where the conjunction resolves into the fire decision. The
from_scratch/ set covers 18β22 and is published alongside for comparison and for
pipelines that prefer organism-specific transcoders across the whole upper band.
No warm-start validation JSONs were produced for L23β25; the validate_L*.json files
accompany the from-scratch set only.
Companion artifacts
- Organism:
Ftm23/backdoor-gemma2-2b-2pair-hate-pt - Attribution graphs:
Ftm23/backdoor-attribution-graphs
Model tree for Ftm23/backdoor-transcoders-gemma2-2b-pt
Base model
google/gemma-2-2b