alephlm-0 / eval /alephlm0_benchmark.md
AbstractPhil's picture
alephlm-0 benchmark
4cb9b61 verified
|
Raw
History Blame Contribute Delete
2.17 kB

AlephLM-0 benchmark

model params STS-B SICK-R STS12 STS13 STS14 STS15 STS16 BIOSSES mean
bert-base 109.5M 0.4729 0.5865 0.3087 0.5988 0.4773 0.6029 0.6373 0.5467 0.5289
ModernBERT-base 149.0M 0.4215 0.5479 0.3527 0.4247 0.3795 0.5349 0.4174 0.5630 0.4552
roberta-base 124.6M 0.5436 0.6296 0.3211 0.5631 0.4522 0.6134 0.6198 0.5777 0.5400
albert-base-v2 11.7M 0.4784 0.5364 0.3101 0.4831 0.3809 0.5542 0.5491 0.4863 0.4723
distilbert 66.4M 0.5717 0.6424 0.4344 0.6490 0.5410 0.6663 0.6854 0.5162 0.5883
all-MiniLM-L6-v2 22.7M 0.8203 0.7758 0.7237 0.8058 0.7559 0.8539 0.7899 0.8144 0.7925
captionbert-v2 58.3M 0.5747 0.6526 0.5051 0.5995 0.5452 0.7136 0.6776 0.5933 0.6077
captionbert-v2 + arms 63.2M 0.7684 0.7391 0.6682 0.7557 0.6921 0.8055 0.7626 0.6382 0.7287
captionbert-b 58.3M 0.5752 0.6548 0.5012 0.6037 0.5470 0.7146 0.6782 0.5500 0.6031
captionbert-b + arms 63.2M 0.7675 0.7374 0.6705 0.7381 0.6945 0.8109 0.7695 0.6472 0.7294
alephlm0-a1_anchored-s0 [cls] 58.3M 0.5731 0.6528 0.4996 0.6014 0.5457 0.7121 0.6761 0.5639 0.6031
alephlm0-a1_anchored-s0 OFF [cls] 58.3M 0.5809 0.6226 0.5200 0.5720 0.5554 0.6773 0.6232 0.4434 0.5743
alephlm0-a2_dense-s0 [cls] 58.3M 0.5731 0.6507 0.4965 0.5962 0.5430 0.7130 0.6798 0.5683 0.6026
alephlm0-a3_random-s0 [cls] 58.3M 0.5707 0.6538 0.4966 0.5998 0.5448 0.7103 0.6806 0.5703 0.6033
alephlm0-a3_random-s0 OFF [cls] 58.3M 0.5767 0.6309 0.4948 0.5735 0.5486 0.6851 0.6307 0.4776 0.5772

AlephLM-0 rows are evaluated with each run's OWN trained pooling (labeled); raw-teacher context rows are mean-pooled per the certified harness. 'OFF' rows = anchored trunks with dispatch disabled (the C6 toggle-law read).