injection evaluation
Threshold 0.02 (calibrated on the validation split, objective recall_at_fpr). The config default of 0.5 scored 0.986 against 0.991 for the calibrated value, on validation. Every table below is on test, at the calibrated threshold.
the chosen threshold 0.02 is the lowest value in the sweep, so the scores are compressed toward zero and this is underfitting rather than a calibration success. Check average precision before trusting the F1.
Per language
| Language | Support | P | R | F1 | Notes |
|---|---|---|---|---|---|
bg Bulgarian |
41 | 1.000 | 1.000 | 1.000 | |
cs Czech |
42 | 0.955 | 1.000 | 0.977 | |
da Danish |
41 | 0.976 | 1.000 | 0.988 | |
de German |
42 | 1.000 | 1.000 | 1.000 | |
el Greek |
42 | 1.000 | 1.000 | 1.000 | |
en English |
41 | 1.000 | 1.000 | 1.000 | |
es Spanish |
40 | 0.976 | 1.000 | 0.988 | |
et Estonian |
41 | 1.000 | 1.000 | 1.000 | |
fi Finnish |
42 | 1.000 | 1.000 | 1.000 | |
fr French |
41 | 1.000 | 1.000 | 1.000 | |
ga Irish |
42 | 0.976 | 0.976 | 0.976 | |
hr Croatian |
42 | 0.977 | 1.000 | 0.988 | |
hu Hungarian |
42 | 0.955 | 1.000 | 0.977 | |
it Italian |
42 | 1.000 | 1.000 | 1.000 | |
lt Lithuanian |
42 | 1.000 | 1.000 | 1.000 | |
lv Latvian |
41 | 0.976 | 1.000 | 0.988 | |
mt Maltese |
42 | 0.804 | 0.976 | 0.882 | not in base model pretraining |
nl Dutch |
41 | 1.000 | 0.976 | 0.988 | |
pl Polish |
41 | 1.000 | 1.000 | 1.000 | |
pt Portuguese |
42 | 0.977 | 1.000 | 0.988 | |
ro Romanian |
42 | 0.977 | 1.000 | 0.988 | |
sk Slovak |
41 | 1.000 | 1.000 | 1.000 | |
sl Slovenian |
42 | 0.977 | 1.000 | 0.988 | |
sv Swedish |
41 | 1.000 | 1.000 | 1.000 | |
tr Turkish |
42 | 1.000 | 1.000 | 1.000 | |
az Azerbaijani |
42 | 1.000 | 1.000 | 1.000 |
The base-model note is a fact about pretraining, not a cause of the score beside it. nsfw Maltese carried the same note at 0.000 and reached 1.000 on corpus size alone, with nothing about the base model changed. Check how many examples a weak score rests on before reaching for this.
Per register
| Register | Support | P | R | F1 | FPR |
|---|---|---|---|---|---|
encoded_payload |
155 | 1.000 | 1.000 | 1.000 | 0.000 |
hypothetical_framing |
156 | 1.000 | 0.994 | 0.997 | 0.000 |
ignore_instructions |
155 | 1.000 | 1.000 | 1.000 | 0.000 |
instruction_in_document |
155 | 1.000 | 0.994 | 0.997 | 0.000 |
instruction_in_tool_output |
148 | 1.000 | 0.993 | 0.997 | 0.000 |
meta_question |
0 | 0.000 | 0.000 | 0.000 | 0.012 |
mundane_account_access |
0 | 0.000 | 0.000 | 0.000 | 0.000 |
mundane_informational |
0 | 0.000 | 0.000 | 0.000 | 0.000 |
mundane_operational |
0 | 0.000 | 0.000 | 0.000 | 0.005 |
mundane_transactional |
0 | 0.000 | 0.000 | 0.000 | 0.005 |
ordinary_instruction |
0 | 0.000 | 0.000 | 0.000 | 0.000 |
ordinary_question |
0 | 0.000 | 0.000 | 0.000 | 0.006 |
persona_override |
155 | 1.000 | 1.000 | 1.000 | 0.000 |
quoted_attack |
0 | 0.000 | 0.000 | 0.000 | 0.009 |
roleplay_benign |
0 | 0.000 | 0.000 | 0.000 | 0.012 |
security_discussion |
0 | 0.000 | 0.000 | 0.000 | 0.009 |
system_prompt_extraction |
156 | 1.000 | 1.000 | 1.000 | 0.000 |
technical_identifiers |
0 | 0.000 | 0.000 | 0.000 | 0.006 |
technical_payload |
0 | 0.000 | 0.000 | 0.000 | 0.006 |
Known weaknesses
The three weakest languages by F1: mt at 0.882, ga at 0.976, cs at 0.977.
These are published rather than dropped. A coverage table with the bad rows removed is not a coverage table.