File size: 6,714 Bytes
b8131eb
22a7a34
b8131eb
 
 
 
05def1c
 
 
 
 
 
 
 
 
 
 
 
b8131eb
 
 
 
 
05def1c
 
 
b8131eb
 
 
 
662e1da
05def1c
b8131eb
 
 
 
05def1c
 
 
 
 
 
 
b8131eb
 
 
 
 
05def1c
b8131eb
 
 
 
05def1c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4243ca7
05def1c
 
 
 
b8131eb
05def1c
 
 
 
 
 
 
 
 
 
 
 
 
b8131eb
 
 
 
05def1c
 
 
 
 
 
 
 
 
 
b8131eb
 
22a7a34
 
05def1c
 
b8131eb
 
 
05def1c
b8131eb
05def1c
 
b8131eb
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
---
license: mit
tags:
- probes
- activation-probes
- interpretability
datasets:
- AlignmentResearch/fibs-v1
base_model:
- Qwen/Qwen3.5-2B
- Qwen/Qwen3.5-9B
- Qwen/Qwen3.5-27B
- Qwen/Qwen3.5-122B-A10B
- Qwen/Qwen3.5-397B-A17B
- nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
- nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16
- nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
- moonshotai/Kimi-K3
---

# probe-inference weights

Trained activation probes in four architectures (linear, MLP, EFC (early-fusion covariance) and axial)
for nine open-weight models. Each probe reads a model's residual-stream activations at six decoder
layers and returns one score per transcript. Load them with the `probe-inference` package
([AlignmentResearch/caught-in-the-act-probes](https://github.com/AlignmentResearch/caught-in-the-act-probes)):

```python
from probe_inference import load_probe_from_hub

probe = load_probe_from_hub("qwen3.5-9b/efc")  # this repository at the package's pinned revision
score = probe.score(acts, probe.read_mask(prompt_mask, completion_mask, followup_start_position))
```

## Layout

`<model>/<arch>/`, with `<arch>` in `linear`, `mlp`, `efc`, `axial`. Each directory is one trained probe. A
linear or MLP probe is one small probe per layer (`layer_<L>/config.json`, `layer_<L>/model.pt`); an EFC or
axial probe is one module that reads all its layers at once (`config.json`, `model.pt`). Every probe has
`probe_metadata.json`: the model and revision, the architecture, the layers, the read window
(`obfuscate_over`), the token aggregation (`eval_sequence_aggregator`) and, for linear and MLP, the layers
whose sigmoids are averaged (`layer_rule.used_layers`, all six). `model.pt` files are plain float32 state
dicts, with the input normaliser (`input_scale`, and `input_mean` for axial) the probe was trained with.

| Directory | Model (revision) | Layers | linear MiB | MLP MiB | EFC MiB | axial MiB |
|---|---|---|---|---|---|---|
| `qwen3.5-2b` | `Qwen/Qwen3.5-2B` (`15852e8c`) | 7, 10, 13, 16, 19, 22 | 0.1 | 12.0 | 6.1 | 26.1 |
| `qwen3.5-9b` | `Qwen/Qwen3.5-9B` (`c2022362`) | 10, 13, 18, 21, 26, 29 | 0.1 | 24.0 | 12.1 | 28.1 |
| `qwen3.5-27b` | `Qwen/Qwen3.5-27B` (`fc05daec`) | 19, 27, 35, 43, 51, 58 | 0.1 | 30.0 | 15.1 | 29.2 |
| `qwen3.5-122b-a10b` | `Qwen/Qwen3.5-122B-A10B` (`dc4d3484`) | 14, 20, 26, 32, 38, 43 | 0.1 | 18.0 | 9.1 | 27.1 |
| `qwen3.5-397b-a17b` | `Qwen/Qwen3.5-397B-A17B` (`84726181`) | 18, 25, 33, 40, 48, 54 | 0.1 | 24.0 | 12.1 | 28.1 |
| `nemotron-3-nano-30b-a3b` | `nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16` (`bf77c317`) | 16, 22, 29, 35, 42, 47 | 0.1 | 15.8 | 8.0 | 26.7 |
| `nemotron-3-super-120b-a12b` | `nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16` (`2dc98e2a`) | 26, 37, 48, 59, 70, 79 | 0.1 | 24.0 | 12.1 | 28.1 |
| `nemotron-3-ultra-550b-a55b` | `nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16` (`77df655d`) | 32, 45, 59, 72, 86, 97 | 0.2 | 48.0 | 24.1 | 32.2 |
| `kimi-k3` | `moonshotai/Kimi-K3` (`f831ab66`) | 28, 39, 51, 62, 74, 84 | 0.2 | 42.0 | 21.1 | 31.2 |

The layers are at depth fractions 0.3, 0.42, 0.55, 0.67, 0.8 and 0.9 of the model's decoder blocks,
`round(f * num_blocks)`. Linear and MLP scores are probabilities (the mean of per-layer sigmoids); EFC and
axial scores are logits.

## Training

Every probe was trained on [FIBSv1](https://huggingface.co/datasets/AlignmentResearch/fibs-v1)
(`AlignmentResearch/fibs-v1`) at revision `65dccf12934620bfe01db665ea3d7d0e3c9f1355`, on the model's own
activations over FIBSv1's transcripts:

- 152,980 training rows and 5,000 validation rows, the same rows for every model; seed 0;
- linear and MLP: up to 6 epochs; EFC and axial: up to 4 epochs;
- early stopping: training stops after 4 checkpoints (one epoch) in a row without a lower validation
  loss (6 epochs for linear/mlp). The published probe is the checkpoint with the lowest validation loss;
- linear and MLP divide each layer's activations by `input_scale`, and axial subtracts `input_mean`
  and then divides by `input_scale`. These are computed per layer from up to 512 training rows and are
  stored in `model.pt`. EFC normalises each token by its own RMS and stores no statistic;
- each published probe reproduces the trainer's validation loss and AUROC within 1e-4.

Validation AUROC on the 5,000 validation rows (ES: stopped early):

| Directory | linear | MLP | EFC | axial |
|---|---|---|---|---|
| `qwen3.5-2b` | 0.926 | 0.968 | 0.987 | 0.991 ES |
| `qwen3.5-9b` | 0.965 | 0.986 | 0.996 | 0.999 |
| `qwen3.5-27b` | 0.972 | 0.988 | 0.998 | 0.999 |
| `qwen3.5-122b-a10b` | 0.938 | 0.972 | 0.998 ES | 0.999 |
| `qwen3.5-397b-a17b` | 0.948 | 0.977 | 0.999 | 0.999 |
| `nemotron-3-nano-30b-a3b` | 0.951 | 0.980 | 0.993 | 0.996 ES |
| `nemotron-3-super-120b-a12b` | 0.968 | 0.986 | 0.998 | 0.998 ES |
| `nemotron-3-ultra-550b-a55b` | 0.982 | 0.993 | 0.999 | 0.999 ES |
| `kimi-k3` | 0.948 | 0.959 | 0.999 | 0.999 ES |

## Activations the probes expect

- Layer `k` is the output of decoder block `k`, i.e. Hugging Face `hidden_states[k + 1]`, in bfloat16; the
  probes run in float32. All activations were captured with vLLM at the decoder-layer outputs.
- Kimi K3's decoder layers use attention residuals, so a layer has no single residual stream. Its layer
  `k` is the attention-residual mixture that layer `k + 1` reads, computed with the model's own
  `attn_res` op. Kimi K3 ran from its released MXFP4 checkpoint, not a bf16 one; its activations are
  bfloat16.
- Linear and MLP read one token: the token before the final end-of-turn token. For Qwen and Nemotron-3
  that is the answer's last token. For Kimi K3 it is the `<|sep|>` that closes `<|close|>message`, after
  the answer's `<|close|>response<|sep|>`. EFC and axial read every token from the start of the final user
  turn through the end-of-turn token.

## Licences and attribution

The probe weights and this card are released by FAR AI, Inc. under the MIT licence (`LICENSE`).

They are derived from the models below. Their licence texts ship here unchanged, and their attribution
notices are kept in `NOTICE`:

| Model | Licence | Licence file |
|---|---|---|
| Qwen/Qwen3.5-2B, Qwen/Qwen3.5-9B, Qwen/Qwen3.5-27B, Qwen/Qwen3.5-122B-A10B, Qwen/Qwen3.5-397B-A17B | Apache-2.0 | `LICENSE-QWEN-APACHE-2.0.txt` |
| nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16, nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 | NVIDIA Nemotron Open Model License (v. December 15, 2025) | `LICENSE-NVIDIA-NEMOTRON-OPEN-MODEL.txt` |
| nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | OpenMDW License Agreement, version 1.1 | `LICENSE-OPENMDW-1.1.txt` |
| moonshotai/Kimi-K3 | Kimi K3 License | `LICENSE-KIMI-K3.txt` |

Licensed by NVIDIA Corporation under the NVIDIA Nemotron Model License.