File size: 3,644 Bytes
47180b5 e4a177f c3fded4 47180b5 c3fded4 e4a177f c3fded4 e4a177f c3fded4 e4a177f c3fded4 e4a177f c3fded4 e4a177f c3fded4 47180b5 e4a177f c3fded4 e4a177f 47180b5 c3fded4 e4a177f c3fded4 e4a177f c3fded4 e4a177f 47180b5 e4a177f 47180b5 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 | # injection evaluation
**Threshold 0.02 (calibrated on the validation split, objective recall_at_fpr).** The config default of 0.5 scored 0.986 against 0.991 for the calibrated value, on validation. Every table below is on test, at the calibrated threshold.
> the chosen threshold 0.02 is the lowest value in the sweep, so the scores are compressed toward zero and this is underfitting rather than a calibration success. Check average precision before trusting the F1.
## Per language
| Language | Support | P | R | F1 | Notes |
|---|---|---|---|---|---|
| `bg` Bulgarian | 41 | 1.000 | 1.000 | 1.000 | |
| `cs` Czech | 42 | 0.955 | 1.000 | 0.977 | |
| `da` Danish | 41 | 0.976 | 1.000 | 0.988 | |
| `de` German | 42 | 1.000 | 1.000 | 1.000 | |
| `el` Greek | 42 | 1.000 | 1.000 | 1.000 | |
| `en` English | 41 | 1.000 | 1.000 | 1.000 | |
| `es` Spanish | 40 | 0.976 | 1.000 | 0.988 | |
| `et` Estonian | 41 | 1.000 | 1.000 | 1.000 | |
| `fi` Finnish | 42 | 1.000 | 1.000 | 1.000 | |
| `fr` French | 41 | 1.000 | 1.000 | 1.000 | |
| `ga` Irish | 42 | 0.976 | 0.976 | 0.976 | |
| `hr` Croatian | 42 | 0.977 | 1.000 | 0.988 | |
| `hu` Hungarian | 42 | 0.955 | 1.000 | 0.977 | |
| `it` Italian | 42 | 1.000 | 1.000 | 1.000 | |
| `lt` Lithuanian | 42 | 1.000 | 1.000 | 1.000 | |
| `lv` Latvian | 41 | 0.976 | 1.000 | 0.988 | |
| `mt` Maltese | 42 | 0.804 | 0.976 | 0.882 | not in base model pretraining |
| `nl` Dutch | 41 | 1.000 | 0.976 | 0.988 | |
| `pl` Polish | 41 | 1.000 | 1.000 | 1.000 | |
| `pt` Portuguese | 42 | 0.977 | 1.000 | 0.988 | |
| `ro` Romanian | 42 | 0.977 | 1.000 | 0.988 | |
| `sk` Slovak | 41 | 1.000 | 1.000 | 1.000 | |
| `sl` Slovenian | 42 | 0.977 | 1.000 | 0.988 | |
| `sv` Swedish | 41 | 1.000 | 1.000 | 1.000 | |
| `tr` Turkish | 42 | 1.000 | 1.000 | 1.000 | |
| `az` Azerbaijani | 42 | 1.000 | 1.000 | 1.000 | |
The base-model note is a fact about pretraining, not a cause of the score beside it. `nsfw` Maltese carried the same note at 0.000 and reached 1.000 on corpus size alone, with nothing about the base model changed. Check how many examples a weak score rests on before reaching for this.
## Per register
| Register | Support | P | R | F1 | FPR |
|---|---|---|---|---|---|
| `encoded_payload` | 155 | 1.000 | 1.000 | 1.000 | 0.000 |
| `hypothetical_framing` | 156 | 1.000 | 0.994 | 0.997 | 0.000 |
| `ignore_instructions` | 155 | 1.000 | 1.000 | 1.000 | 0.000 |
| `instruction_in_document` | 155 | 1.000 | 0.994 | 0.997 | 0.000 |
| `instruction_in_tool_output` | 148 | 1.000 | 0.993 | 0.997 | 0.000 |
| `meta_question` | 0 | 0.000 | 0.000 | 0.000 | 0.012 |
| `mundane_account_access` | 0 | 0.000 | 0.000 | 0.000 | 0.000 |
| `mundane_informational` | 0 | 0.000 | 0.000 | 0.000 | 0.000 |
| `mundane_operational` | 0 | 0.000 | 0.000 | 0.000 | 0.005 |
| `mundane_transactional` | 0 | 0.000 | 0.000 | 0.000 | 0.005 |
| `ordinary_instruction` | 0 | 0.000 | 0.000 | 0.000 | 0.000 |
| `ordinary_question` | 0 | 0.000 | 0.000 | 0.000 | 0.006 |
| `persona_override` | 155 | 1.000 | 1.000 | 1.000 | 0.000 |
| `quoted_attack` | 0 | 0.000 | 0.000 | 0.000 | 0.009 |
| `roleplay_benign` | 0 | 0.000 | 0.000 | 0.000 | 0.012 |
| `security_discussion` | 0 | 0.000 | 0.000 | 0.000 | 0.009 |
| `system_prompt_extraction` | 156 | 1.000 | 1.000 | 1.000 | 0.000 |
| `technical_identifiers` | 0 | 0.000 | 0.000 | 0.000 | 0.006 |
| `technical_payload` | 0 | 0.000 | 0.000 | 0.000 | 0.006 |
## Known weaknesses
The three weakest languages by F1: `mt` at 0.882, `ga` at 0.976, `cs` at 0.977.
These are published rather than dropped. A coverage table with the bad rows removed is not a coverage table. |