cubic-150m / benchmark_table.md
Asilarkness's picture
Cubic Hier 150M: weights, tokenizer and benchmark comparison
8c55c89 verified
|
Raw
History Blame Contribute Delete
1.07 kB
Model Params hellaswag arc_easy arc_challenge piqa winogrande openbookqa sciq boolq lambada_openai mmlu Avg Δ chance
CubicHierLM-157M (base format) 157M 30.1 39.1 25.3 57.8 50.8 25.2 70.4 61.8 20.1 27.1 40.8 +10.8
CubicHierLM-157M (chat, direct prompt) 157M 29.7 36.6 24.1 57.1 49.1 29.2 66.5 43.7 0.7 25.7 36.2 +6.2
CubicHierLM-157M (chat, reasoning prompt) 157M 30.1 35.7 24.0 57.1 50.8 28.8 65.9 42.2 1.1 26.3 36.2 +6.2
HuggingFaceTB/SmolLM2-135M 135M 43.9 59.9 29.2 68.2 53.3 32.4 79.9 61.1 41.1 23.1 49.2 +19.2
EleutherAI/pythia-160m 162M 29.0 37.7 24.7 58.7 50.7 25.2 63.6 43.9 10.9 24.1 36.9 +6.9
facebook/opt-125m 125M 30.3 40.9 22.7 61.8 50.4 27.4 68.9 55.8 38.1 24.2 42.1 +12.1
random chance 25.0 25.0 25.0 50.0 50.0 25.0 25.0 50.0 0.0 25.0 +0.0