skar0's picture
Update README.md
4243ca7 verified
|
Raw History Blame Contribute Delete
6.71 kB
---
license: mit
tags:
- probes
- activation-probes
- interpretability
datasets:
- AlignmentResearch/fibs-v1
base_model:
- Qwen/Qwen3.5-2B
- Qwen/Qwen3.5-9B
- Qwen/Qwen3.5-27B
- Qwen/Qwen3.5-122B-A10B
- Qwen/Qwen3.5-397B-A17B
- nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
- nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16
- nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
- moonshotai/Kimi-K3
---
# probe-inference weights
Trained activation probes in four architectures (linear, MLP, EFC (early-fusion covariance) and axial)
for nine open-weight models. Each probe reads a model's residual-stream activations at six decoder
layers and returns one score per transcript. Load them with the `probe-inference` package
([AlignmentResearch/caught-in-the-act-probes](https://github.com/AlignmentResearch/caught-in-the-act-probes)):
```python
from probe_inference import load_probe_from_hub
probe = load_probe_from_hub("qwen3.5-9b/efc") # this repository at the package's pinned revision
score = probe.score(acts, probe.read_mask(prompt_mask, completion_mask, followup_start_position))
```
## Layout
`<model>/<arch>/`, with `<arch>` in `linear`, `mlp`, `efc`, `axial`. Each directory is one trained probe. A
linear or MLP probe is one small probe per layer (`layer_<L>/config.json`, `layer_<L>/model.pt`); an EFC or
axial probe is one module that reads all its layers at once (`config.json`, `model.pt`). Every probe has
`probe_metadata.json`: the model and revision, the architecture, the layers, the read window
(`obfuscate_over`), the token aggregation (`eval_sequence_aggregator`) and, for linear and MLP, the layers
whose sigmoids are averaged (`layer_rule.used_layers`, all six). `model.pt` files are plain float32 state
dicts, with the input normaliser (`input_scale`, and `input_mean` for axial) the probe was trained with.
| Directory | Model (revision) | Layers | linear MiB | MLP MiB | EFC MiB | axial MiB |
|---|---|---|---|---|---|---|
| `qwen3.5-2b` | `Qwen/Qwen3.5-2B` (`15852e8c`) | 7, 10, 13, 16, 19, 22 | 0.1 | 12.0 | 6.1 | 26.1 |
| `qwen3.5-9b` | `Qwen/Qwen3.5-9B` (`c2022362`) | 10, 13, 18, 21, 26, 29 | 0.1 | 24.0 | 12.1 | 28.1 |
| `qwen3.5-27b` | `Qwen/Qwen3.5-27B` (`fc05daec`) | 19, 27, 35, 43, 51, 58 | 0.1 | 30.0 | 15.1 | 29.2 |
| `qwen3.5-122b-a10b` | `Qwen/Qwen3.5-122B-A10B` (`dc4d3484`) | 14, 20, 26, 32, 38, 43 | 0.1 | 18.0 | 9.1 | 27.1 |
| `qwen3.5-397b-a17b` | `Qwen/Qwen3.5-397B-A17B` (`84726181`) | 18, 25, 33, 40, 48, 54 | 0.1 | 24.0 | 12.1 | 28.1 |
| `nemotron-3-nano-30b-a3b` | `nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16` (`bf77c317`) | 16, 22, 29, 35, 42, 47 | 0.1 | 15.8 | 8.0 | 26.7 |
| `nemotron-3-super-120b-a12b` | `nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16` (`2dc98e2a`) | 26, 37, 48, 59, 70, 79 | 0.1 | 24.0 | 12.1 | 28.1 |
| `nemotron-3-ultra-550b-a55b` | `nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16` (`77df655d`) | 32, 45, 59, 72, 86, 97 | 0.2 | 48.0 | 24.1 | 32.2 |
| `kimi-k3` | `moonshotai/Kimi-K3` (`f831ab66`) | 28, 39, 51, 62, 74, 84 | 0.2 | 42.0 | 21.1 | 31.2 |
The layers are at depth fractions 0.3, 0.42, 0.55, 0.67, 0.8 and 0.9 of the model's decoder blocks,
`round(f * num_blocks)`. Linear and MLP scores are probabilities (the mean of per-layer sigmoids); EFC and
axial scores are logits.
## Training
Every probe was trained on [FIBSv1](https://huggingface.co/datasets/AlignmentResearch/fibs-v1)
(`AlignmentResearch/fibs-v1`) at revision `65dccf12934620bfe01db665ea3d7d0e3c9f1355`, on the model's own
activations over FIBSv1's transcripts:
- 152,980 training rows and 5,000 validation rows, the same rows for every model; seed 0;
- linear and MLP: up to 6 epochs; EFC and axial: up to 4 epochs;
- early stopping: training stops after 4 checkpoints (one epoch) in a row without a lower validation
loss (6 epochs for linear/mlp). The published probe is the checkpoint with the lowest validation loss;
- linear and MLP divide each layer's activations by `input_scale`, and axial subtracts `input_mean`
and then divides by `input_scale`. These are computed per layer from up to 512 training rows and are
stored in `model.pt`. EFC normalises each token by its own RMS and stores no statistic;
- each published probe reproduces the trainer's validation loss and AUROC within 1e-4.
Validation AUROC on the 5,000 validation rows (ES: stopped early):
| Directory | linear | MLP | EFC | axial |
|---|---|---|---|---|
| `qwen3.5-2b` | 0.926 | 0.968 | 0.987 | 0.991 ES |
| `qwen3.5-9b` | 0.965 | 0.986 | 0.996 | 0.999 |
| `qwen3.5-27b` | 0.972 | 0.988 | 0.998 | 0.999 |
| `qwen3.5-122b-a10b` | 0.938 | 0.972 | 0.998 ES | 0.999 |
| `qwen3.5-397b-a17b` | 0.948 | 0.977 | 0.999 | 0.999 |
| `nemotron-3-nano-30b-a3b` | 0.951 | 0.980 | 0.993 | 0.996 ES |
| `nemotron-3-super-120b-a12b` | 0.968 | 0.986 | 0.998 | 0.998 ES |
| `nemotron-3-ultra-550b-a55b` | 0.982 | 0.993 | 0.999 | 0.999 ES |
| `kimi-k3` | 0.948 | 0.959 | 0.999 | 0.999 ES |
## Activations the probes expect
- Layer `k` is the output of decoder block `k`, i.e. Hugging Face `hidden_states[k + 1]`, in bfloat16; the
probes run in float32. All activations were captured with vLLM at the decoder-layer outputs.
- Kimi K3's decoder layers use attention residuals, so a layer has no single residual stream. Its layer
`k` is the attention-residual mixture that layer `k + 1` reads, computed with the model's own
`attn_res` op. Kimi K3 ran from its released MXFP4 checkpoint, not a bf16 one; its activations are
bfloat16.
- Linear and MLP read one token: the token before the final end-of-turn token. For Qwen and Nemotron-3
that is the answer's last token. For Kimi K3 it is the `<|sep|>` that closes `<|close|>message`, after
the answer's `<|close|>response<|sep|>`. EFC and axial read every token from the start of the final user
turn through the end-of-turn token.
## Licences and attribution
The probe weights and this card are released by FAR AI, Inc. under the MIT licence (`LICENSE`).
They are derived from the models below. Their licence texts ship here unchanged, and their attribution
notices are kept in `NOTICE`:
| Model | Licence | Licence file |
|---|---|---|
| Qwen/Qwen3.5-2B, Qwen/Qwen3.5-9B, Qwen/Qwen3.5-27B, Qwen/Qwen3.5-122B-A10B, Qwen/Qwen3.5-397B-A17B | Apache-2.0 | `LICENSE-QWEN-APACHE-2.0.txt` |
| nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16, nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 | NVIDIA Nemotron Open Model License (v. December 15, 2025) | `LICENSE-NVIDIA-NEMOTRON-OPEN-MODEL.txt` |
| nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | OpenMDW License Agreement, version 1.1 | `LICENSE-OPENMDW-1.1.txt` |
| moonshotai/Kimi-K3 | Kimi K3 License | `LICENSE-KIMI-K3.txt` |
Licensed by NVIDIA Corporation under the NVIDIA Nemotron Model License.