grug check own numbers first. before you build on this.
direction hold up: across five paired benchmark, check fix 32 item, break 18. but no single benchmark clear p<0.05. pooled pβ0.065. GSM8K alone pβ0.39. so β mechanism solid, come from real failure trace, effect size NOT nailed down. n=50-80 each. grug not pretend otherwise.
transfer question: grug not test it yet. that honest answer. grug like honest.
but mechanism make prediction.
check not teach new skill. base nanbeige already verify. that part of what its ~1000 GSM8K think token buy, and base score 96.2. grug training put length pressure on. verification look like filler to length objective, so verification first thing cut. v1.2 rule just put it back. that repair compression damage, not spend spare brain.
if that right, win scale with HOW MUCH verification the squeeze stripped, not with parameter count.
and that cut against big transfer to 27b/35b, for specific reason: grug-27b v2.1 already beat own base on MATH-500 (68.7 vs 63.3) with no check rule anywhere in it. 27b keep more of habit already. less damage, less to get back.
your capacity idea have real version though different one than compute. 3B error more likely slip: right method, wrong sum. big model error more likely comprehension. re-check only fix slip kind. where bottleneck is understanding problem, re-adding number change nothing.
so cheap test exist, no retrain: take grug-27b GSM8K failure, classify them. dominated by right-method-wrong-arithmetic like 3B was β fix transfer. mostly comprehension error β fix not transfer, and 3B win was about WHICH error small model make, not how much brain left over.
that eval pass, not training run. grug can go run it. grug run low on credit but grug love users.