bias / eval /bias_eval.md
bogdanraduta's picture
Upload eval/bias_eval.md with huggingface_hub
60367bc verified
|
Raw
History Blame Contribute Delete
3.05 kB

bias evaluation

Threshold 0.77 (calibrated on the validation split, objective macro_f1). The config default of 0.5 scored 0.981 against 0.982 for the calibrated value, on validation. Every table below is on test, at the calibrated threshold.

Per language

Language Support P R F1 Notes
bg Bulgarian 80 0.987 0.975 0.981
cs Czech 80 0.988 1.000 0.994
da Danish 80 1.000 1.000 1.000
de German 80 0.963 0.988 0.975
el Greek 80 1.000 0.988 0.994
en English 80 0.963 0.988 0.975
es Spanish 76 1.000 0.974 0.987
et Estonian 79 0.963 1.000 0.981
fi Finnish 80 0.988 0.988 0.988
fr French 80 0.976 1.000 0.988
ga Irish 77 0.974 0.961 0.967
hr Croatian 80 1.000 1.000 1.000
hu Hungarian 80 0.952 0.988 0.969
it Italian 80 1.000 1.000 1.000
lt Lithuanian 79 0.987 0.987 0.987
lv Latvian 80 0.975 0.988 0.981
mt Maltese 78 0.948 0.936 0.942 not in base model pretraining
nl Dutch 79 0.988 1.000 0.994
pl Polish 80 0.976 1.000 0.988
pt Portuguese 76 0.974 0.974 0.974
ro Romanian 80 1.000 0.988 0.994
sk Slovak 80 0.988 0.988 0.988
sl Slovenian 80 1.000 0.988 0.994
sv Swedish 80 0.988 0.988 0.988
tr Turkish 80 0.951 0.975 0.963
az Azerbaijani 80 0.920 1.000 0.958

The base-model note is a fact about pretraining, not a cause of the score beside it. nsfw Maltese carried the same note at 0.000 and reached 1.000 on corpus size alone, with nothing about the base model changed. Check how many examples a weak score rests on before reaching for this.

Per register

Register Support P R F1 FPR
age_assumption 415 1.000 0.986 0.993 0.000
counter_stereotype 0 0.000 0.000 0.000 0.011
demographic_statistic 0 0.000 0.000 0.000 0.000
disability_condescension 413 1.000 0.995 0.998 0.000
discussing_bias 0 0.000 0.000 0.000 0.016
ethnic_generalisation 411 1.000 0.990 0.995 0.000
inclusive_phrasing 0 0.000 0.000 0.000 0.096
mundane_informational 0 0.000 0.000 0.000 0.000
mundane_operational 0 0.000 0.000 0.000 0.000
mundane_transactional 0 0.000 0.000 0.000 0.000
neutral_description 0 0.000 0.000 0.000 0.000
occupational_stereotype 415 1.000 0.969 0.984 0.000
religious_assumption 410 1.000 0.995 0.998 0.000

Known weaknesses

The three weakest languages by F1: mt at 0.942, az at 0.958, tr at 0.963.

These are published rather than dropped. A coverage table with the bad rows removed is not a coverage table.