The Quantitative −10 is 5 questions.
Your Base Bench 1.1 categories average exactly to 51.71 and 40.29, and every cell is a multiple of 2%. That reads as 7 categories of 50 items. Quantitative is 13/50 full vs 18/50 G-only.
Best case for the claim: all 5 are clean flips and none go the other way. An exact McNemar test gives p = 0.0625. One flip back (6-1) and it is 0.125. I would not hang the headline on that cell yet.
The bigger table hides a different question. G-only is M zeroed at inference, on a G that trained with M injecting into its residual stream after layers 4, 10 and 16. So the delta measures how much G came to depend on M. That can be large even if M adds nothing a plain model could not learn.
It also folds in compute. M runs 28 block traversals per token on top of G's 18.
Is there a G-only model trained from scratch at matched tokens or matched FLOPs to compare against?