Update README.md
Browse files
README.md
CHANGED
|
@@ -32,11 +32,6 @@ probe = load_probe_from_hub("qwen3.5-9b/efc") # this repository at the package'
|
|
| 32 |
score = probe.score(acts, probe.read_mask(prompt_mask, completion_mask, followup_start_position))
|
| 33 |
```
|
| 34 |
|
| 35 |
-
These are the camera-ready probes, tagged `camera-ready`. The earlier probes (seven models, Qwen3.6-27B
|
| 36 |
-
instead of Qwen3.5-27B, no Nemotron-3 Ultra or Kimi K3) stay available at commit
|
| 37 |
-
`22a7a341078ba1722cfad73ef7e40bdc25aa74c2`, and in schema-1 metadata at commit
|
| 38 |
-
`0e3d386b1f1486d472e316a64966a7a94bcd5137`.
|
| 39 |
-
|
| 40 |
## Layout
|
| 41 |
|
| 42 |
`<model>/<arch>/`, with `<arch>` in `linear`, `mlp`, `efc`, `axial`. Each directory is one trained probe. A
|
|
@@ -72,7 +67,7 @@ activations over FIBSv1's transcripts:
|
|
| 72 |
- 152,980 training rows and 5,000 validation rows, the same rows for every model; seed 0;
|
| 73 |
- linear and MLP: up to 6 epochs; EFC and axial: up to 4 epochs;
|
| 74 |
- early stopping: training stops after 4 checkpoints (one epoch) in a row without a lower validation
|
| 75 |
-
loss. The published probe is the checkpoint with the lowest validation loss;
|
| 76 |
- linear and MLP divide each layer's activations by `input_scale`, and axial subtracts `input_mean`
|
| 77 |
and then divides by `input_scale`. These are computed per layer from up to 512 training rows and are
|
| 78 |
stored in `model.pt`. EFC normalises each token by its own RMS and stores no statistic;
|
|
@@ -100,21 +95,11 @@ Validation AUROC on the 5,000 validation rows (ES: stopped early):
|
|
| 100 |
`k` is the attention-residual mixture that layer `k + 1` reads, computed with the model's own
|
| 101 |
`attn_res` op. Kimi K3 ran from its released MXFP4 checkpoint, not a bf16 one; its activations are
|
| 102 |
bfloat16.
|
| 103 |
-
- The transcript ends with a final user turn and an assistant answer, closed by the end-of-turn token,
|
| 104 |
-
rendered with the model's chat template with thinking disabled. Nemotron-3: the template's default,
|
| 105 |
-
which removes the reasoning of every assistant turn before the final user turn. Kimi K3: its template
|
| 106 |
-
cannot disable thinking, so the final assistant turn has an empty reasoning block before the answer.
|
| 107 |
- Linear and MLP read one token: the token before the final end-of-turn token. For Qwen and Nemotron-3
|
| 108 |
that is the answer's last token. For Kimi K3 it is the `<|sep|>` that closes `<|close|>message`, after
|
| 109 |
the answer's `<|close|>response<|sep|>`. EFC and axial read every token from the start of the final user
|
| 110 |
turn through the end-of-turn token.
|
| 111 |
|
| 112 |
-
```
|
| 113 |
-
Qwen3.5: <|im_start|>user\n{question}<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\nNo.<|im_end|>
|
| 114 |
-
Nemotron-3: <|im_start|>user\n{question}<|im_end|>\n<|im_start|>assistant\n<think></think>No.<|im_end|>
|
| 115 |
-
Kimi K3: <|open|>message role="user"<|sep|>{question}<|close|>message<|sep|><|end_of_msg|><|open|>message role="assistant"<|sep|><|open|>think<|sep|><|close|>think<|sep|><|open|>response<|sep|>No.<|close|>response<|sep|><|close|>message<|sep|><|end_of_msg|>
|
| 116 |
-
```
|
| 117 |
-
|
| 118 |
## Licences and attribution
|
| 119 |
|
| 120 |
The probe weights and this card are released by FAR AI, Inc. under the MIT licence (`LICENSE`).
|
|
|
|
| 32 |
score = probe.score(acts, probe.read_mask(prompt_mask, completion_mask, followup_start_position))
|
| 33 |
```
|
| 34 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 35 |
## Layout
|
| 36 |
|
| 37 |
`<model>/<arch>/`, with `<arch>` in `linear`, `mlp`, `efc`, `axial`. Each directory is one trained probe. A
|
|
|
|
| 67 |
- 152,980 training rows and 5,000 validation rows, the same rows for every model; seed 0;
|
| 68 |
- linear and MLP: up to 6 epochs; EFC and axial: up to 4 epochs;
|
| 69 |
- early stopping: training stops after 4 checkpoints (one epoch) in a row without a lower validation
|
| 70 |
+
loss (6 epochs for linear/mlp). The published probe is the checkpoint with the lowest validation loss;
|
| 71 |
- linear and MLP divide each layer's activations by `input_scale`, and axial subtracts `input_mean`
|
| 72 |
and then divides by `input_scale`. These are computed per layer from up to 512 training rows and are
|
| 73 |
stored in `model.pt`. EFC normalises each token by its own RMS and stores no statistic;
|
|
|
|
| 95 |
`k` is the attention-residual mixture that layer `k + 1` reads, computed with the model's own
|
| 96 |
`attn_res` op. Kimi K3 ran from its released MXFP4 checkpoint, not a bf16 one; its activations are
|
| 97 |
bfloat16.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 98 |
- Linear and MLP read one token: the token before the final end-of-turn token. For Qwen and Nemotron-3
|
| 99 |
that is the answer's last token. For Kimi K3 it is the `<|sep|>` that closes `<|close|>message`, after
|
| 100 |
the answer's `<|close|>response<|sep|>`. EFC and axial read every token from the start of the final user
|
| 101 |
turn through the end-of-turn token.
|
| 102 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 103 |
## Licences and attribution
|
| 104 |
|
| 105 |
The probe weights and this card are released by FAR AI, Inc. under the MIT licence (`LICENSE`).
|