File size: 3,644 Bytes
47180b5
 
e4a177f
c3fded4
 
47180b5
 
 
 
 
c3fded4
 
 
 
 
e4a177f
 
c3fded4
 
e4a177f
c3fded4
e4a177f
c3fded4
 
 
 
e4a177f
 
c3fded4
 
 
 
 
 
e4a177f
c3fded4
47180b5
 
 
 
 
 
 
e4a177f
 
c3fded4
e4a177f
 
 
 
47180b5
c3fded4
 
e4a177f
 
c3fded4
e4a177f
 
 
c3fded4
e4a177f
 
47180b5
 
 
e4a177f
47180b5
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
# injection evaluation

**Threshold 0.02 (calibrated on the validation split, objective recall_at_fpr).** The config default of 0.5 scored 0.986 against 0.991 for the calibrated value, on validation. Every table below is on test, at the calibrated threshold.

> the chosen threshold 0.02 is the lowest value in the sweep, so the scores are compressed toward zero and this is underfitting rather than a calibration success. Check average precision before trusting the F1.

## Per language

| Language | Support | P | R | F1 | Notes |
|---|---|---|---|---|---|
| `bg` Bulgarian | 41 | 1.000 | 1.000 | 1.000 |  |
| `cs` Czech | 42 | 0.955 | 1.000 | 0.977 |  |
| `da` Danish | 41 | 0.976 | 1.000 | 0.988 |  |
| `de` German | 42 | 1.000 | 1.000 | 1.000 |  |
| `el` Greek | 42 | 1.000 | 1.000 | 1.000 |  |
| `en` English | 41 | 1.000 | 1.000 | 1.000 |  |
| `es` Spanish | 40 | 0.976 | 1.000 | 0.988 |  |
| `et` Estonian | 41 | 1.000 | 1.000 | 1.000 |  |
| `fi` Finnish | 42 | 1.000 | 1.000 | 1.000 |  |
| `fr` French | 41 | 1.000 | 1.000 | 1.000 |  |
| `ga` Irish | 42 | 0.976 | 0.976 | 0.976 |  |
| `hr` Croatian | 42 | 0.977 | 1.000 | 0.988 |  |
| `hu` Hungarian | 42 | 0.955 | 1.000 | 0.977 |  |
| `it` Italian | 42 | 1.000 | 1.000 | 1.000 |  |
| `lt` Lithuanian | 42 | 1.000 | 1.000 | 1.000 |  |
| `lv` Latvian | 41 | 0.976 | 1.000 | 0.988 |  |
| `mt` Maltese | 42 | 0.804 | 0.976 | 0.882 | not in base model pretraining |
| `nl` Dutch | 41 | 1.000 | 0.976 | 0.988 |  |
| `pl` Polish | 41 | 1.000 | 1.000 | 1.000 |  |
| `pt` Portuguese | 42 | 0.977 | 1.000 | 0.988 |  |
| `ro` Romanian | 42 | 0.977 | 1.000 | 0.988 |  |
| `sk` Slovak | 41 | 1.000 | 1.000 | 1.000 |  |
| `sl` Slovenian | 42 | 0.977 | 1.000 | 0.988 |  |
| `sv` Swedish | 41 | 1.000 | 1.000 | 1.000 |  |
| `tr` Turkish | 42 | 1.000 | 1.000 | 1.000 |  |
| `az` Azerbaijani | 42 | 1.000 | 1.000 | 1.000 |  |

The base-model note is a fact about pretraining, not a cause of the score beside it. `nsfw` Maltese carried the same note at 0.000 and reached 1.000 on corpus size alone, with nothing about the base model changed. Check how many examples a weak score rests on before reaching for this.

## Per register

| Register | Support | P | R | F1 | FPR |
|---|---|---|---|---|---|
| `encoded_payload` | 155 | 1.000 | 1.000 | 1.000 | 0.000 |
| `hypothetical_framing` | 156 | 1.000 | 0.994 | 0.997 | 0.000 |
| `ignore_instructions` | 155 | 1.000 | 1.000 | 1.000 | 0.000 |
| `instruction_in_document` | 155 | 1.000 | 0.994 | 0.997 | 0.000 |
| `instruction_in_tool_output` | 148 | 1.000 | 0.993 | 0.997 | 0.000 |
| `meta_question` | 0 | 0.000 | 0.000 | 0.000 | 0.012 |
| `mundane_account_access` | 0 | 0.000 | 0.000 | 0.000 | 0.000 |
| `mundane_informational` | 0 | 0.000 | 0.000 | 0.000 | 0.000 |
| `mundane_operational` | 0 | 0.000 | 0.000 | 0.000 | 0.005 |
| `mundane_transactional` | 0 | 0.000 | 0.000 | 0.000 | 0.005 |
| `ordinary_instruction` | 0 | 0.000 | 0.000 | 0.000 | 0.000 |
| `ordinary_question` | 0 | 0.000 | 0.000 | 0.000 | 0.006 |
| `persona_override` | 155 | 1.000 | 1.000 | 1.000 | 0.000 |
| `quoted_attack` | 0 | 0.000 | 0.000 | 0.000 | 0.009 |
| `roleplay_benign` | 0 | 0.000 | 0.000 | 0.000 | 0.012 |
| `security_discussion` | 0 | 0.000 | 0.000 | 0.000 | 0.009 |
| `system_prompt_extraction` | 156 | 1.000 | 1.000 | 1.000 | 0.000 |
| `technical_identifiers` | 0 | 0.000 | 0.000 | 0.000 | 0.006 |
| `technical_payload` | 0 | 0.000 | 0.000 | 0.000 | 0.006 |

## Known weaknesses

The three weakest languages by F1: `mt` at 0.882, `ga` at 0.976, `cs` at 0.977.

These are published rather than dropped. A coverage table with the bad rows removed is not a coverage table.