Download README.md from AlignmentResearch/probe-inference-weights: direct link, hf CLI and curl.
- Browser
- Download file 7.9 kB
-
https://huggingface.co/AlignmentResearch/probe-inference-weights/resolve/main/README.md
- Command line
-
hf download hf://AlignmentResearch/probe-inference-weights/README.md
-
curl -L -o README.md https://huggingface.co/AlignmentResearch/probe-inference-weights/resolve/main/README.md
license: mit
tags:
- probes
- activation-probes
- interpretability
datasets:
- AlignmentResearch/fibs-v1
base_model:
- Qwen/Qwen3.5-2B
- Qwen/Qwen3.5-9B
- Qwen/Qwen3.5-27B
- Qwen/Qwen3.5-122B-A10B
- Qwen/Qwen3.5-397B-A17B
- nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
- nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16
- nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
- moonshotai/Kimi-K3
probe-inference weights
Trained activation probes in four architectures (linear, MLP, EFC (early-fusion covariance) and axial)
for nine open-weight models. Each probe reads a model's residual-stream activations at six decoder
layers and returns one score per transcript. Load them with the probe-inference package
(AlignmentResearch/caught-in-the-act-probes):
from probe_inference import load_probe_from_hub
probe = load_probe_from_hub("qwen3.5-9b/efc") # this repository at the package's pinned revision
score = probe.score(acts, probe.read_mask(prompt_mask, completion_mask, followup_start_position))
These are the camera-ready probes, tagged camera-ready. The earlier probes (seven models, Qwen3.6-27B
instead of Qwen3.5-27B, no Nemotron-3 Ultra or Kimi K3) stay available at commit
22a7a341078ba1722cfad73ef7e40bdc25aa74c2, and in schema-1 metadata at commit
0e3d386b1f1486d472e316a64966a7a94bcd5137.
Layout
<model>/<arch>/, with <arch> in linear, mlp, efc, axial. Each directory is one trained probe. A
linear or MLP probe is one small probe per layer (layer_<L>/config.json, layer_<L>/model.pt); an EFC or
axial probe is one module that reads all its layers at once (config.json, model.pt). Every probe has
probe_metadata.json: the model and revision, the architecture, the layers, the read window
(obfuscate_over), the token aggregation (eval_sequence_aggregator) and, for linear and MLP, the layers
whose sigmoids are averaged (layer_rule.used_layers, all six). model.pt files are plain float32 state
dicts, with the input normaliser (input_scale, and input_mean for axial) the probe was trained with.
| Directory | Model (revision) | Layers | linear MiB | MLP MiB | EFC MiB | axial MiB |
|---|---|---|---|---|---|---|
qwen3.5-2b |
Qwen/Qwen3.5-2B (15852e8c) |
7, 10, 13, 16, 19, 22 | 0.1 | 12.0 | 6.1 | 26.1 |
qwen3.5-9b |
Qwen/Qwen3.5-9B (c2022362) |
10, 13, 18, 21, 26, 29 | 0.1 | 24.0 | 12.1 | 28.1 |
qwen3.5-27b |
Qwen/Qwen3.5-27B (fc05daec) |
19, 27, 35, 43, 51, 58 | 0.1 | 30.0 | 15.1 | 29.2 |
qwen3.5-122b-a10b |
Qwen/Qwen3.5-122B-A10B (dc4d3484) |
14, 20, 26, 32, 38, 43 | 0.1 | 18.0 | 9.1 | 27.1 |
qwen3.5-397b-a17b |
Qwen/Qwen3.5-397B-A17B (84726181) |
18, 25, 33, 40, 48, 54 | 0.1 | 24.0 | 12.1 | 28.1 |
nemotron-3-nano-30b-a3b |
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 (bf77c317) |
16, 22, 29, 35, 42, 47 | 0.1 | 15.8 | 8.0 | 26.7 |
nemotron-3-super-120b-a12b |
nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 (2dc98e2a) |
26, 37, 48, 59, 70, 79 | 0.1 | 24.0 | 12.1 | 28.1 |
nemotron-3-ultra-550b-a55b |
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 (77df655d) |
32, 45, 59, 72, 86, 97 | 0.2 | 48.0 | 24.1 | 32.2 |
kimi-k3 |
moonshotai/Kimi-K3 (f831ab66) |
28, 39, 51, 62, 74, 84 | 0.2 | 42.0 | 21.1 | 31.2 |
The layers are at depth fractions 0.3, 0.42, 0.55, 0.67, 0.8 and 0.9 of the model's decoder blocks,
round(f * num_blocks). Linear and MLP scores are probabilities (the mean of per-layer sigmoids); EFC and
axial scores are logits.
Training
Every probe was trained on FIBSv1
(AlignmentResearch/fibs-v1) at revision 65dccf12934620bfe01db665ea3d7d0e3c9f1355, on the model's own
activations over FIBSv1's transcripts:
- 152,980 training rows and 5,000 validation rows, the same rows for every model; seed 0;
- linear and MLP: up to 6 epochs; EFC and axial: up to 4 epochs;
- early stopping: training stops after 4 checkpoints (one epoch) in a row without a lower validation loss. The published probe is the checkpoint with the lowest validation loss;
- linear and MLP divide each layer's activations by
input_scale, and axial subtractsinput_meanand then divides byinput_scale. These are computed per layer from up to 512 training rows and are stored inmodel.pt. EFC normalises each token by its own RMS and stores no statistic; - each published probe reproduces the trainer's validation loss and AUROC within 1e-4.
Validation AUROC on the 5,000 validation rows (ES: stopped early):
| Directory | linear | MLP | EFC | axial |
|---|---|---|---|---|
qwen3.5-2b |
0.926 | 0.968 | 0.987 | 0.991 ES |
qwen3.5-9b |
0.965 | 0.986 | 0.996 | 0.999 |
qwen3.5-27b |
0.972 | 0.988 | 0.998 | 0.999 |
qwen3.5-122b-a10b |
0.938 | 0.972 | 0.998 ES | 0.999 |
qwen3.5-397b-a17b |
0.948 | 0.977 | 0.999 | 0.999 |
nemotron-3-nano-30b-a3b |
0.951 | 0.980 | 0.993 | 0.996 ES |
nemotron-3-super-120b-a12b |
0.968 | 0.986 | 0.998 | 0.998 ES |
nemotron-3-ultra-550b-a55b |
0.982 | 0.993 | 0.999 | 0.999 ES |
kimi-k3 |
0.948 | 0.959 | 0.999 | 0.999 ES |
Activations the probes expect
- Layer
kis the output of decoder blockk, i.e. Hugging Facehidden_states[k + 1], in bfloat16; the probes run in float32. All activations were captured with vLLM at the decoder-layer outputs. - Kimi K3's decoder layers use attention residuals, so a layer has no single residual stream. Its layer
kis the attention-residual mixture that layerk + 1reads, computed with the model's ownattn_resop. Kimi K3 ran from its released MXFP4 checkpoint, not a bf16 one; its activations are bfloat16. - The transcript ends with a final user turn and an assistant answer, closed by the end-of-turn token, rendered with the model's chat template with thinking disabled. Nemotron-3: the template's default, which removes the reasoning of every assistant turn before the final user turn. Kimi K3: its template cannot disable thinking, so the final assistant turn has an empty reasoning block before the answer.
- Linear and MLP read one token: the token before the final end-of-turn token. For Qwen and Nemotron-3
that is the answer's last token. For Kimi K3 it is the
<|sep|>that closes<|close|>message, after the answer's<|close|>response<|sep|>. EFC and axial read every token from the start of the final user turn through the end-of-turn token.
Qwen3.5: <|im_start|>user\n{question}<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\nNo.<|im_end|>
Nemotron-3: <|im_start|>user\n{question}<|im_end|>\n<|im_start|>assistant\n<think></think>No.<|im_end|>
Kimi K3: <|open|>message role="user"<|sep|>{question}<|close|>message<|sep|><|end_of_msg|><|open|>message role="assistant"<|sep|><|open|>think<|sep|><|close|>think<|sep|><|open|>response<|sep|>No.<|close|>response<|sep|><|close|>message<|sep|><|end_of_msg|>
Licences and attribution
The probe weights and this card are released by FAR AI, Inc. under the MIT licence (LICENSE).
They are derived from the models below. Their licence texts ship here unchanged, and their attribution
notices are kept in NOTICE:
| Model | Licence | Licence file |
|---|---|---|
| Qwen/Qwen3.5-2B, Qwen/Qwen3.5-9B, Qwen/Qwen3.5-27B, Qwen/Qwen3.5-122B-A10B, Qwen/Qwen3.5-397B-A17B | Apache-2.0 | LICENSE-QWEN-APACHE-2.0.txt |
| nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16, nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 | NVIDIA Nemotron Open Model License (v. December 15, 2025) | LICENSE-NVIDIA-NEMOTRON-OPEN-MODEL.txt |
| nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | OpenMDW License Agreement, version 1.1 | LICENSE-OPENMDW-1.1.txt |
| moonshotai/Kimi-K3 | Kimi K3 License | LICENSE-KIMI-K3.txt |
Licensed by NVIDIA Corporation under the NVIDIA Nemotron Model License.