File size: 6,714 Bytes
b8131eb 22a7a34 b8131eb 05def1c b8131eb 05def1c b8131eb 662e1da 05def1c b8131eb 05def1c b8131eb 05def1c b8131eb 05def1c 4243ca7 05def1c b8131eb 05def1c b8131eb 05def1c b8131eb 22a7a34 05def1c b8131eb 05def1c b8131eb 05def1c b8131eb | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 | ---
license: mit
tags:
- probes
- activation-probes
- interpretability
datasets:
- AlignmentResearch/fibs-v1
base_model:
- Qwen/Qwen3.5-2B
- Qwen/Qwen3.5-9B
- Qwen/Qwen3.5-27B
- Qwen/Qwen3.5-122B-A10B
- Qwen/Qwen3.5-397B-A17B
- nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
- nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16
- nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
- moonshotai/Kimi-K3
---
# probe-inference weights
Trained activation probes in four architectures (linear, MLP, EFC (early-fusion covariance) and axial)
for nine open-weight models. Each probe reads a model's residual-stream activations at six decoder
layers and returns one score per transcript. Load them with the `probe-inference` package
([AlignmentResearch/caught-in-the-act-probes](https://github.com/AlignmentResearch/caught-in-the-act-probes)):
```python
from probe_inference import load_probe_from_hub
probe = load_probe_from_hub("qwen3.5-9b/efc") # this repository at the package's pinned revision
score = probe.score(acts, probe.read_mask(prompt_mask, completion_mask, followup_start_position))
```
## Layout
`<model>/<arch>/`, with `<arch>` in `linear`, `mlp`, `efc`, `axial`. Each directory is one trained probe. A
linear or MLP probe is one small probe per layer (`layer_<L>/config.json`, `layer_<L>/model.pt`); an EFC or
axial probe is one module that reads all its layers at once (`config.json`, `model.pt`). Every probe has
`probe_metadata.json`: the model and revision, the architecture, the layers, the read window
(`obfuscate_over`), the token aggregation (`eval_sequence_aggregator`) and, for linear and MLP, the layers
whose sigmoids are averaged (`layer_rule.used_layers`, all six). `model.pt` files are plain float32 state
dicts, with the input normaliser (`input_scale`, and `input_mean` for axial) the probe was trained with.
| Directory | Model (revision) | Layers | linear MiB | MLP MiB | EFC MiB | axial MiB |
|---|---|---|---|---|---|---|
| `qwen3.5-2b` | `Qwen/Qwen3.5-2B` (`15852e8c`) | 7, 10, 13, 16, 19, 22 | 0.1 | 12.0 | 6.1 | 26.1 |
| `qwen3.5-9b` | `Qwen/Qwen3.5-9B` (`c2022362`) | 10, 13, 18, 21, 26, 29 | 0.1 | 24.0 | 12.1 | 28.1 |
| `qwen3.5-27b` | `Qwen/Qwen3.5-27B` (`fc05daec`) | 19, 27, 35, 43, 51, 58 | 0.1 | 30.0 | 15.1 | 29.2 |
| `qwen3.5-122b-a10b` | `Qwen/Qwen3.5-122B-A10B` (`dc4d3484`) | 14, 20, 26, 32, 38, 43 | 0.1 | 18.0 | 9.1 | 27.1 |
| `qwen3.5-397b-a17b` | `Qwen/Qwen3.5-397B-A17B` (`84726181`) | 18, 25, 33, 40, 48, 54 | 0.1 | 24.0 | 12.1 | 28.1 |
| `nemotron-3-nano-30b-a3b` | `nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16` (`bf77c317`) | 16, 22, 29, 35, 42, 47 | 0.1 | 15.8 | 8.0 | 26.7 |
| `nemotron-3-super-120b-a12b` | `nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16` (`2dc98e2a`) | 26, 37, 48, 59, 70, 79 | 0.1 | 24.0 | 12.1 | 28.1 |
| `nemotron-3-ultra-550b-a55b` | `nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16` (`77df655d`) | 32, 45, 59, 72, 86, 97 | 0.2 | 48.0 | 24.1 | 32.2 |
| `kimi-k3` | `moonshotai/Kimi-K3` (`f831ab66`) | 28, 39, 51, 62, 74, 84 | 0.2 | 42.0 | 21.1 | 31.2 |
The layers are at depth fractions 0.3, 0.42, 0.55, 0.67, 0.8 and 0.9 of the model's decoder blocks,
`round(f * num_blocks)`. Linear and MLP scores are probabilities (the mean of per-layer sigmoids); EFC and
axial scores are logits.
## Training
Every probe was trained on [FIBSv1](https://huggingface.co/datasets/AlignmentResearch/fibs-v1)
(`AlignmentResearch/fibs-v1`) at revision `65dccf12934620bfe01db665ea3d7d0e3c9f1355`, on the model's own
activations over FIBSv1's transcripts:
- 152,980 training rows and 5,000 validation rows, the same rows for every model; seed 0;
- linear and MLP: up to 6 epochs; EFC and axial: up to 4 epochs;
- early stopping: training stops after 4 checkpoints (one epoch) in a row without a lower validation
loss (6 epochs for linear/mlp). The published probe is the checkpoint with the lowest validation loss;
- linear and MLP divide each layer's activations by `input_scale`, and axial subtracts `input_mean`
and then divides by `input_scale`. These are computed per layer from up to 512 training rows and are
stored in `model.pt`. EFC normalises each token by its own RMS and stores no statistic;
- each published probe reproduces the trainer's validation loss and AUROC within 1e-4.
Validation AUROC on the 5,000 validation rows (ES: stopped early):
| Directory | linear | MLP | EFC | axial |
|---|---|---|---|---|
| `qwen3.5-2b` | 0.926 | 0.968 | 0.987 | 0.991 ES |
| `qwen3.5-9b` | 0.965 | 0.986 | 0.996 | 0.999 |
| `qwen3.5-27b` | 0.972 | 0.988 | 0.998 | 0.999 |
| `qwen3.5-122b-a10b` | 0.938 | 0.972 | 0.998 ES | 0.999 |
| `qwen3.5-397b-a17b` | 0.948 | 0.977 | 0.999 | 0.999 |
| `nemotron-3-nano-30b-a3b` | 0.951 | 0.980 | 0.993 | 0.996 ES |
| `nemotron-3-super-120b-a12b` | 0.968 | 0.986 | 0.998 | 0.998 ES |
| `nemotron-3-ultra-550b-a55b` | 0.982 | 0.993 | 0.999 | 0.999 ES |
| `kimi-k3` | 0.948 | 0.959 | 0.999 | 0.999 ES |
## Activations the probes expect
- Layer `k` is the output of decoder block `k`, i.e. Hugging Face `hidden_states[k + 1]`, in bfloat16; the
probes run in float32. All activations were captured with vLLM at the decoder-layer outputs.
- Kimi K3's decoder layers use attention residuals, so a layer has no single residual stream. Its layer
`k` is the attention-residual mixture that layer `k + 1` reads, computed with the model's own
`attn_res` op. Kimi K3 ran from its released MXFP4 checkpoint, not a bf16 one; its activations are
bfloat16.
- Linear and MLP read one token: the token before the final end-of-turn token. For Qwen and Nemotron-3
that is the answer's last token. For Kimi K3 it is the `<|sep|>` that closes `<|close|>message`, after
the answer's `<|close|>response<|sep|>`. EFC and axial read every token from the start of the final user
turn through the end-of-turn token.
## Licences and attribution
The probe weights and this card are released by FAR AI, Inc. under the MIT licence (`LICENSE`).
They are derived from the models below. Their licence texts ship here unchanged, and their attribution
notices are kept in `NOTICE`:
| Model | Licence | Licence file |
|---|---|---|
| Qwen/Qwen3.5-2B, Qwen/Qwen3.5-9B, Qwen/Qwen3.5-27B, Qwen/Qwen3.5-122B-A10B, Qwen/Qwen3.5-397B-A17B | Apache-2.0 | `LICENSE-QWEN-APACHE-2.0.txt` |
| nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16, nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 | NVIDIA Nemotron Open Model License (v. December 15, 2025) | `LICENSE-NVIDIA-NEMOTRON-OPEN-MODEL.txt` |
| nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | OpenMDW License Agreement, version 1.1 | `LICENSE-OPENMDW-1.1.txt` |
| moonshotai/Kimi-K3 | Kimi K3 License | `LICENSE-KIMI-K3.txt` |
Licensed by NVIDIA Corporation under the NVIDIA Nemotron Model License.
|