chrisjcundy's picture
Camera-ready probes: 9 models x {linear, mlp, efc, axial} trained on FIBSv1 @ 65dccf12
05def1c verified
|
Raw History Blame Contribute Delete
7.9 kB
metadata
license: mit
tags:
  - probes
  - activation-probes
  - interpretability
datasets:
  - AlignmentResearch/fibs-v1
base_model:
  - Qwen/Qwen3.5-2B
  - Qwen/Qwen3.5-9B
  - Qwen/Qwen3.5-27B
  - Qwen/Qwen3.5-122B-A10B
  - Qwen/Qwen3.5-397B-A17B
  - nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
  - nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16
  - nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
  - moonshotai/Kimi-K3

probe-inference weights

Trained activation probes in four architectures (linear, MLP, EFC (early-fusion covariance) and axial) for nine open-weight models. Each probe reads a model's residual-stream activations at six decoder layers and returns one score per transcript. Load them with the probe-inference package (AlignmentResearch/caught-in-the-act-probes):

from probe_inference import load_probe_from_hub

probe = load_probe_from_hub("qwen3.5-9b/efc")  # this repository at the package's pinned revision
score = probe.score(acts, probe.read_mask(prompt_mask, completion_mask, followup_start_position))

These are the camera-ready probes, tagged camera-ready. The earlier probes (seven models, Qwen3.6-27B instead of Qwen3.5-27B, no Nemotron-3 Ultra or Kimi K3) stay available at commit 22a7a341078ba1722cfad73ef7e40bdc25aa74c2, and in schema-1 metadata at commit 0e3d386b1f1486d472e316a64966a7a94bcd5137.

Layout

<model>/<arch>/, with <arch> in linear, mlp, efc, axial. Each directory is one trained probe. A linear or MLP probe is one small probe per layer (layer_<L>/config.json, layer_<L>/model.pt); an EFC or axial probe is one module that reads all its layers at once (config.json, model.pt). Every probe has probe_metadata.json: the model and revision, the architecture, the layers, the read window (obfuscate_over), the token aggregation (eval_sequence_aggregator) and, for linear and MLP, the layers whose sigmoids are averaged (layer_rule.used_layers, all six). model.pt files are plain float32 state dicts, with the input normaliser (input_scale, and input_mean for axial) the probe was trained with.

Directory Model (revision) Layers linear MiB MLP MiB EFC MiB axial MiB
qwen3.5-2b Qwen/Qwen3.5-2B (15852e8c) 7, 10, 13, 16, 19, 22 0.1 12.0 6.1 26.1
qwen3.5-9b Qwen/Qwen3.5-9B (c2022362) 10, 13, 18, 21, 26, 29 0.1 24.0 12.1 28.1
qwen3.5-27b Qwen/Qwen3.5-27B (fc05daec) 19, 27, 35, 43, 51, 58 0.1 30.0 15.1 29.2
qwen3.5-122b-a10b Qwen/Qwen3.5-122B-A10B (dc4d3484) 14, 20, 26, 32, 38, 43 0.1 18.0 9.1 27.1
qwen3.5-397b-a17b Qwen/Qwen3.5-397B-A17B (84726181) 18, 25, 33, 40, 48, 54 0.1 24.0 12.1 28.1
nemotron-3-nano-30b-a3b nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 (bf77c317) 16, 22, 29, 35, 42, 47 0.1 15.8 8.0 26.7
nemotron-3-super-120b-a12b nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 (2dc98e2a) 26, 37, 48, 59, 70, 79 0.1 24.0 12.1 28.1
nemotron-3-ultra-550b-a55b nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 (77df655d) 32, 45, 59, 72, 86, 97 0.2 48.0 24.1 32.2
kimi-k3 moonshotai/Kimi-K3 (f831ab66) 28, 39, 51, 62, 74, 84 0.2 42.0 21.1 31.2

The layers are at depth fractions 0.3, 0.42, 0.55, 0.67, 0.8 and 0.9 of the model's decoder blocks, round(f * num_blocks). Linear and MLP scores are probabilities (the mean of per-layer sigmoids); EFC and axial scores are logits.

Training

Every probe was trained on FIBSv1 (AlignmentResearch/fibs-v1) at revision 65dccf12934620bfe01db665ea3d7d0e3c9f1355, on the model's own activations over FIBSv1's transcripts:

  • 152,980 training rows and 5,000 validation rows, the same rows for every model; seed 0;
  • linear and MLP: up to 6 epochs; EFC and axial: up to 4 epochs;
  • early stopping: training stops after 4 checkpoints (one epoch) in a row without a lower validation loss. The published probe is the checkpoint with the lowest validation loss;
  • linear and MLP divide each layer's activations by input_scale, and axial subtracts input_mean and then divides by input_scale. These are computed per layer from up to 512 training rows and are stored in model.pt. EFC normalises each token by its own RMS and stores no statistic;
  • each published probe reproduces the trainer's validation loss and AUROC within 1e-4.

Validation AUROC on the 5,000 validation rows (ES: stopped early):

Directory linear MLP EFC axial
qwen3.5-2b 0.926 0.968 0.987 0.991 ES
qwen3.5-9b 0.965 0.986 0.996 0.999
qwen3.5-27b 0.972 0.988 0.998 0.999
qwen3.5-122b-a10b 0.938 0.972 0.998 ES 0.999
qwen3.5-397b-a17b 0.948 0.977 0.999 0.999
nemotron-3-nano-30b-a3b 0.951 0.980 0.993 0.996 ES
nemotron-3-super-120b-a12b 0.968 0.986 0.998 0.998 ES
nemotron-3-ultra-550b-a55b 0.982 0.993 0.999 0.999 ES
kimi-k3 0.948 0.959 0.999 0.999 ES

Activations the probes expect

  • Layer k is the output of decoder block k, i.e. Hugging Face hidden_states[k + 1], in bfloat16; the probes run in float32. All activations were captured with vLLM at the decoder-layer outputs.
  • Kimi K3's decoder layers use attention residuals, so a layer has no single residual stream. Its layer k is the attention-residual mixture that layer k + 1 reads, computed with the model's own attn_res op. Kimi K3 ran from its released MXFP4 checkpoint, not a bf16 one; its activations are bfloat16.
  • The transcript ends with a final user turn and an assistant answer, closed by the end-of-turn token, rendered with the model's chat template with thinking disabled. Nemotron-3: the template's default, which removes the reasoning of every assistant turn before the final user turn. Kimi K3: its template cannot disable thinking, so the final assistant turn has an empty reasoning block before the answer.
  • Linear and MLP read one token: the token before the final end-of-turn token. For Qwen and Nemotron-3 that is the answer's last token. For Kimi K3 it is the <|sep|> that closes <|close|>message, after the answer's <|close|>response<|sep|>. EFC and axial read every token from the start of the final user turn through the end-of-turn token.
Qwen3.5:     <|im_start|>user\n{question}<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\nNo.<|im_end|>
Nemotron-3:  <|im_start|>user\n{question}<|im_end|>\n<|im_start|>assistant\n<think></think>No.<|im_end|>
Kimi K3:     <|open|>message role="user"<|sep|>{question}<|close|>message<|sep|><|end_of_msg|><|open|>message role="assistant"<|sep|><|open|>think<|sep|><|close|>think<|sep|><|open|>response<|sep|>No.<|close|>response<|sep|><|close|>message<|sep|><|end_of_msg|>

Licences and attribution

The probe weights and this card are released by FAR AI, Inc. under the MIT licence (LICENSE).

They are derived from the models below. Their licence texts ship here unchanged, and their attribution notices are kept in NOTICE:

Model Licence Licence file
Qwen/Qwen3.5-2B, Qwen/Qwen3.5-9B, Qwen/Qwen3.5-27B, Qwen/Qwen3.5-122B-A10B, Qwen/Qwen3.5-397B-A17B Apache-2.0 LICENSE-QWEN-APACHE-2.0.txt
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16, nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 NVIDIA Nemotron Open Model License (v. December 15, 2025) LICENSE-NVIDIA-NEMOTRON-OPEN-MODEL.txt
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 OpenMDW License Agreement, version 1.1 LICENSE-OPENMDW-1.1.txt
moonshotai/Kimi-K3 Kimi K3 License LICENSE-KIMI-K3.txt

Licensed by NVIDIA Corporation under the NVIDIA Nemotron Model License.