injection / eval /injection_eval.md
bogdanraduta's picture
Upload eval/injection_eval.md with huggingface_hub
e4a177f verified
|
Raw
History Blame Contribute Delete
3.64 kB

injection evaluation

Threshold 0.02 (calibrated on the validation split, objective recall_at_fpr). The config default of 0.5 scored 0.986 against 0.991 for the calibrated value, on validation. Every table below is on test, at the calibrated threshold.

the chosen threshold 0.02 is the lowest value in the sweep, so the scores are compressed toward zero and this is underfitting rather than a calibration success. Check average precision before trusting the F1.

Per language

Language Support P R F1 Notes
bg Bulgarian 41 1.000 1.000 1.000
cs Czech 42 0.955 1.000 0.977
da Danish 41 0.976 1.000 0.988
de German 42 1.000 1.000 1.000
el Greek 42 1.000 1.000 1.000
en English 41 1.000 1.000 1.000
es Spanish 40 0.976 1.000 0.988
et Estonian 41 1.000 1.000 1.000
fi Finnish 42 1.000 1.000 1.000
fr French 41 1.000 1.000 1.000
ga Irish 42 0.976 0.976 0.976
hr Croatian 42 0.977 1.000 0.988
hu Hungarian 42 0.955 1.000 0.977
it Italian 42 1.000 1.000 1.000
lt Lithuanian 42 1.000 1.000 1.000
lv Latvian 41 0.976 1.000 0.988
mt Maltese 42 0.804 0.976 0.882 not in base model pretraining
nl Dutch 41 1.000 0.976 0.988
pl Polish 41 1.000 1.000 1.000
pt Portuguese 42 0.977 1.000 0.988
ro Romanian 42 0.977 1.000 0.988
sk Slovak 41 1.000 1.000 1.000
sl Slovenian 42 0.977 1.000 0.988
sv Swedish 41 1.000 1.000 1.000
tr Turkish 42 1.000 1.000 1.000
az Azerbaijani 42 1.000 1.000 1.000

The base-model note is a fact about pretraining, not a cause of the score beside it. nsfw Maltese carried the same note at 0.000 and reached 1.000 on corpus size alone, with nothing about the base model changed. Check how many examples a weak score rests on before reaching for this.

Per register

Register Support P R F1 FPR
encoded_payload 155 1.000 1.000 1.000 0.000
hypothetical_framing 156 1.000 0.994 0.997 0.000
ignore_instructions 155 1.000 1.000 1.000 0.000
instruction_in_document 155 1.000 0.994 0.997 0.000
instruction_in_tool_output 148 1.000 0.993 0.997 0.000
meta_question 0 0.000 0.000 0.000 0.012
mundane_account_access 0 0.000 0.000 0.000 0.000
mundane_informational 0 0.000 0.000 0.000 0.000
mundane_operational 0 0.000 0.000 0.000 0.005
mundane_transactional 0 0.000 0.000 0.000 0.005
ordinary_instruction 0 0.000 0.000 0.000 0.000
ordinary_question 0 0.000 0.000 0.000 0.006
persona_override 155 1.000 1.000 1.000 0.000
quoted_attack 0 0.000 0.000 0.000 0.009
roleplay_benign 0 0.000 0.000 0.000 0.012
security_discussion 0 0.000 0.000 0.000 0.009
system_prompt_extraction 156 1.000 1.000 1.000 0.000
technical_identifiers 0 0.000 0.000 0.000 0.006
technical_payload 0 0.000 0.000 0.000 0.006

Known weaknesses

The three weakest languages by F1: mt at 0.882, ga at 0.976, cs at 0.977.

These are published rather than dropped. A coverage table with the bad rows removed is not a coverage table.