license: gemma
base_model: Ftm23/backdoor-gemma2-2b-2pair-hate-pt
tags:
- transcoder
- interpretability
- circuit-tracer
- sparse-coding
- gemma-2
Transcoders for the C2-2B-pt conjunctive-backdoor organism
Per-layer transcoders trained on the MLP activations of
Ftm23/backdoor-gemma2-2b-2pair-hate-pt,
used to build attribution graphs with
circuit-tracer.
These are not reproducible from the training code alone in any short run — they are the artifact the attribution results depend on, published so that work can be repeated without retraining them.
Contents
| path | layers | how trained |
|---|---|---|
from_scratch/layer_{18..22}.safetensors |
18–22 | trained from scratch on this organism's activations |
warm/layer_{23,24,25}_ep2.safetensors |
23–25 | warm-started from GemmaScope, 2 epochs |
Reconstruction quality (post-LN FVU, k=64, held-out)
| layer | this transcoder | off-the-shelf GemmaScope |
|---|---|---|
| 18 | 0.360 | 0.651 |
| 19 | 0.403 | 0.690 |
| 20 | 0.442 | 0.722 |
| 21 | 0.452 | 0.698 |
| 22 | 0.489 | 0.768 |
Lower is better. The from-scratch transcoders reconstruct this organism's activations substantially better than the off-the-shelf ones at every layer measured.
How they are used
The attribution runs build a ReplacementModel by splicing: off-the-shelf
GemmaScope transcoders for layers 0–22, and the warm/ transcoders here for layers
23–25 — the band where the conjunction resolves into the fire decision. The
from_scratch/ set covers 18–22 and is published alongside for comparison and for
pipelines that prefer organism-specific transcoders across the whole upper band.
No warm-start validation JSONs were produced for L23–25; the validate_L*.json files
accompany the from-scratch set only.
Companion artifacts
- Organism:
Ftm23/backdoor-gemma2-2b-2pair-hate-pt - Attribution graphs:
Ftm23/backdoor-attribution-graphs