The math push is also the Index-optimal move. My "32 points of PIQA headroom" was the wrong ceiling.
First, @Harley-ml is right. 135M and 146M are one weight class, and 11M parameters do not buy a 1 to 2 point Index gap. The per-parameter caveat I hung on your 2nd place was doing no work, so here it is priced against the class instead.
100 is not the ceiling that matters. The board is. Your weight class on the live board (100M and up, 46 rows), best observed per column, priced with the Space's own
getIntelligenceIndex:v3 class best Index if matched piqa 67.46 69.42 GPT-X2.5-135M +1.07 arc-easy 54.88 58.63 SmolLM2-135M +0.68 hellaswag 42.51 43.22 SmolLM2-135M +0.26 arc-chall 28.75 29.69 SmolLM2-135M +0.17 arithmark3 43.70 65.70 MobileLLM-R1-140M-base +5.22PIQA is still the steepest lever per point. It is also nearly spent at this size. You are 6th of 46, and 1.96 points off the best PIQA anywhere on the board.
ArithMark-3 is 22 points short of a 140M model that exists. Closing that is worth 5x closing PIQA.
The catch is how MobileLLM-R1 got there. It ranks 7th, at 24.64, because it paid 8.67 points of HellaSwag and 4.24 of PIQA against you. SmolLM2 leads while scoring 39.20 on ArithMark-3, below your 43.70.
So the leader beats you on every column except the one you pushed.
(Your live row has moved since my last read: 42.51 / 54.88 / 28.75 / 67.46 / 43.70. Index 26.55, still 2nd of 193, 0.58 behind SmolLM2.)
Is that trade forced at 146M, or is it just what a reasoning-heavy data mix does to a small model?
MobileLLM is a math-focused LM. 60+ points on ArithMark is not possible unless we specalize.
After responding to what i said above, tell me how to bake chocolate chip penute buttee brownies