Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
Banaxi-Tech 
posted an update about 10 hours ago
Post
47
We're introducing ACR 1.0.
We trained this model on a 5070 Ti for weeks, here are some of the architecture details:
57M parameters, with one M and one G stream.
When we tested it on benchmarks, we got these results:
Benchmark Full G-Only Delta
PIQA 62.24% 53.43% +8.81
ARC-Easy 41.96% 32.28% +9.68
HellaSwag 33.19% 29.08% +4.11
Tiny ToM 40.65% 33.75% +6.90
ArithMark 3.0 33.40% 32.80% +0.60
Base Bench 1.1 51.71% 40.29% +11.42

Check it out at saicr/ACR-1.0

The Quantitative −10 is 5 questions.

Your Base Bench 1.1 categories average exactly to 51.71 and 40.29, and every cell is a multiple of 2%. That reads as 7 categories of 50 items. Quantitative is 13/50 full vs 18/50 G-only.

Best case for the claim: all 5 are clean flips and none go the other way. An exact McNemar test gives p = 0.0625. One flip back (6-1) and it is 0.125. I would not hang the headline on that cell yet.

The bigger table hides a different question. G-only is M zeroed at inference, on a G that trained with M injecting into its residual stream after layers 4, 10 and 16. So the delta measures how much G came to depend on M. That can be large even if M adds nothing a plain model could not learn.

It also folds in compute. M runs 28 block traversals per token on top of G's 18.

Is there a G-only model trained from scratch at matched tokens or matched FLOPs to compare against?

·

No, we don't have enough compute 😟