--- license: mit tags: - probes - activation-probes - interpretability datasets: - AlignmentResearch/fibs-v1 base_model: - Qwen/Qwen3.5-2B - Qwen/Qwen3.5-9B - Qwen/Qwen3.5-27B - Qwen/Qwen3.5-122B-A10B - Qwen/Qwen3.5-397B-A17B - nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 - nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 - nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 - moonshotai/Kimi-K3 --- # probe-inference weights Trained activation probes in four architectures (linear, MLP, EFC (early-fusion covariance) and axial) for nine open-weight models. Each probe reads a model's residual-stream activations at six decoder layers and returns one score per transcript. Load them with the `probe-inference` package ([AlignmentResearch/caught-in-the-act-probes](https://github.com/AlignmentResearch/caught-in-the-act-probes)): ```python from probe_inference import load_probe_from_hub probe = load_probe_from_hub("qwen3.5-9b/efc") # this repository at the package's pinned revision score = probe.score(acts, probe.read_mask(prompt_mask, completion_mask, followup_start_position)) ``` ## Layout `//`, with `` in `linear`, `mlp`, `efc`, `axial`. Each directory is one trained probe. A linear or MLP probe is one small probe per layer (`layer_/config.json`, `layer_/model.pt`); an EFC or axial probe is one module that reads all its layers at once (`config.json`, `model.pt`). Every probe has `probe_metadata.json`: the model and revision, the architecture, the layers, the read window (`obfuscate_over`), the token aggregation (`eval_sequence_aggregator`) and, for linear and MLP, the layers whose sigmoids are averaged (`layer_rule.used_layers`, all six). `model.pt` files are plain float32 state dicts, with the input normaliser (`input_scale`, and `input_mean` for axial) the probe was trained with. | Directory | Model (revision) | Layers | linear MiB | MLP MiB | EFC MiB | axial MiB | |---|---|---|---|---|---|---| | `qwen3.5-2b` | `Qwen/Qwen3.5-2B` (`15852e8c`) | 7, 10, 13, 16, 19, 22 | 0.1 | 12.0 | 6.1 | 26.1 | | `qwen3.5-9b` | `Qwen/Qwen3.5-9B` (`c2022362`) | 10, 13, 18, 21, 26, 29 | 0.1 | 24.0 | 12.1 | 28.1 | | `qwen3.5-27b` | `Qwen/Qwen3.5-27B` (`fc05daec`) | 19, 27, 35, 43, 51, 58 | 0.1 | 30.0 | 15.1 | 29.2 | | `qwen3.5-122b-a10b` | `Qwen/Qwen3.5-122B-A10B` (`dc4d3484`) | 14, 20, 26, 32, 38, 43 | 0.1 | 18.0 | 9.1 | 27.1 | | `qwen3.5-397b-a17b` | `Qwen/Qwen3.5-397B-A17B` (`84726181`) | 18, 25, 33, 40, 48, 54 | 0.1 | 24.0 | 12.1 | 28.1 | | `nemotron-3-nano-30b-a3b` | `nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16` (`bf77c317`) | 16, 22, 29, 35, 42, 47 | 0.1 | 15.8 | 8.0 | 26.7 | | `nemotron-3-super-120b-a12b` | `nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16` (`2dc98e2a`) | 26, 37, 48, 59, 70, 79 | 0.1 | 24.0 | 12.1 | 28.1 | | `nemotron-3-ultra-550b-a55b` | `nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16` (`77df655d`) | 32, 45, 59, 72, 86, 97 | 0.2 | 48.0 | 24.1 | 32.2 | | `kimi-k3` | `moonshotai/Kimi-K3` (`f831ab66`) | 28, 39, 51, 62, 74, 84 | 0.2 | 42.0 | 21.1 | 31.2 | The layers are at depth fractions 0.3, 0.42, 0.55, 0.67, 0.8 and 0.9 of the model's decoder blocks, `round(f * num_blocks)`. Linear and MLP scores are probabilities (the mean of per-layer sigmoids); EFC and axial scores are logits. ## Training Every probe was trained on [FIBSv1](https://huggingface.co/datasets/AlignmentResearch/fibs-v1) (`AlignmentResearch/fibs-v1`) at revision `65dccf12934620bfe01db665ea3d7d0e3c9f1355`, on the model's own activations over FIBSv1's transcripts: - 152,980 training rows and 5,000 validation rows, the same rows for every model; seed 0; - linear and MLP: up to 6 epochs; EFC and axial: up to 4 epochs; - early stopping: training stops after 4 checkpoints (one epoch) in a row without a lower validation loss (6 epochs for linear/mlp). The published probe is the checkpoint with the lowest validation loss; - linear and MLP divide each layer's activations by `input_scale`, and axial subtracts `input_mean` and then divides by `input_scale`. These are computed per layer from up to 512 training rows and are stored in `model.pt`. EFC normalises each token by its own RMS and stores no statistic; - each published probe reproduces the trainer's validation loss and AUROC within 1e-4. Validation AUROC on the 5,000 validation rows (ES: stopped early): | Directory | linear | MLP | EFC | axial | |---|---|---|---|---| | `qwen3.5-2b` | 0.926 | 0.968 | 0.987 | 0.991 ES | | `qwen3.5-9b` | 0.965 | 0.986 | 0.996 | 0.999 | | `qwen3.5-27b` | 0.972 | 0.988 | 0.998 | 0.999 | | `qwen3.5-122b-a10b` | 0.938 | 0.972 | 0.998 ES | 0.999 | | `qwen3.5-397b-a17b` | 0.948 | 0.977 | 0.999 | 0.999 | | `nemotron-3-nano-30b-a3b` | 0.951 | 0.980 | 0.993 | 0.996 ES | | `nemotron-3-super-120b-a12b` | 0.968 | 0.986 | 0.998 | 0.998 ES | | `nemotron-3-ultra-550b-a55b` | 0.982 | 0.993 | 0.999 | 0.999 ES | | `kimi-k3` | 0.948 | 0.959 | 0.999 | 0.999 ES | ## Activations the probes expect - Layer `k` is the output of decoder block `k`, i.e. Hugging Face `hidden_states[k + 1]`, in bfloat16; the probes run in float32. All activations were captured with vLLM at the decoder-layer outputs. - Kimi K3's decoder layers use attention residuals, so a layer has no single residual stream. Its layer `k` is the attention-residual mixture that layer `k + 1` reads, computed with the model's own `attn_res` op. Kimi K3 ran from its released MXFP4 checkpoint, not a bf16 one; its activations are bfloat16. - Linear and MLP read one token: the token before the final end-of-turn token. For Qwen and Nemotron-3 that is the answer's last token. For Kimi K3 it is the `<|sep|>` that closes `<|close|>message`, after the answer's `<|close|>response<|sep|>`. EFC and axial read every token from the start of the final user turn through the end-of-turn token. ## Licences and attribution The probe weights and this card are released by FAR AI, Inc. under the MIT licence (`LICENSE`). They are derived from the models below. Their licence texts ship here unchanged, and their attribution notices are kept in `NOTICE`: | Model | Licence | Licence file | |---|---|---| | Qwen/Qwen3.5-2B, Qwen/Qwen3.5-9B, Qwen/Qwen3.5-27B, Qwen/Qwen3.5-122B-A10B, Qwen/Qwen3.5-397B-A17B | Apache-2.0 | `LICENSE-QWEN-APACHE-2.0.txt` | | nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16, nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 | NVIDIA Nemotron Open Model License (v. December 15, 2025) | `LICENSE-NVIDIA-NEMOTRON-OPEN-MODEL.txt` | | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | OpenMDW License Agreement, version 1.1 | `LICENSE-OPENMDW-1.1.txt` | | moonshotai/Kimi-K3 | Kimi K3 License | `LICENSE-KIMI-K3.txt` | Licensed by NVIDIA Corporation under the NVIDIA Nemotron Model License.