rei-v1 / eval.md
gnukeith's picture
Rei v1: weights, standalone loader, model card and evaluation
e220d9b verified
|
Raw History Blame Contribute Delete
1.88 kB

rei-v1: evaluation

Accuracy means agreeing with the teacher ensemble's answer, on research questions never seen in training. Temperatures: choice 1.55, noul 1.8. Trained on 74855 items, 3.0 epochs, 10478s.

Test: 76.7% accuracy (majority baseline 42.8%), calibration error (ECE) 0.031. Validation: 79.0%.

Test by task

items accuracy majority baseline soft CE
claim_check 1085 81.4% 33.9% 0.537
custom 421 72.0% 15.0% 0.674
needs_fresh 181 82.9% 60.2% 0.375
next_step 177 76.3% 56.5% 0.686
relevance 1041 68.5% 48.2% 0.781
same_info 180 88.3% 86.1% 0.323
search_type 173 76.9% 31.8% 0.767
source_type 347 80.1% 18.2% 0.647
worth_opening 363 79.6% 78.5% 0.467

Score tasks, mean absolute error in scale points: custom 0.50, relevance 0.33

Noul tasks, AUC (how well P(true) ranks true above false; 0.5 is chance): custom 0.83, needs_fresh 0.91, same_info 0.83, worth_opening 0.76

Test by language

items accuracy majority baseline soft CE
ar 166 79.5% 43.4% 0.602
de 254 80.3% 50.4% 0.563
en 374 77.0% 39.0% 0.627
es 322 79.8% 44.4% 0.555
fr 275 80.4% 44.7% 0.550
hi 140 75.7% 43.6% 0.654
id 402 74.4% 42.8% 0.661
is 474 69.6% 42.0% 0.764
ja 468 78.4% 41.0% 0.575
ko 301 75.7% 42.9% 0.608
pt 306 75.5% 43.8% 0.635
ru 212 75.0% 42.9% 0.631
zh 274 80.7% 40.1% 0.536

Test by profile

items accuracy majority baseline soft CE
academic 742 77.8% 41.6% 0.591
anime 976 78.4% 38.7% 0.570
general 1486 74.4% 44.5% 0.684
programming 764 78.0% 45.9% 0.580