mod-arithmetic / RESULTS.md
mikelou's picture
Add faster compact NeuralHorner v8 prototype
efe8193 verified
|
Raw
History Blame Contribute Delete
989 Bytes
# Optimization results
| Variant | Checkpoint | Schedule | Tier 6 accuracy | Tier 6 time (MPS) |
|---|---:|---|---:|---:|
| Published v8 | 1,887,610 B | reduce a + reduce b + multiply | 100/100 | 27.02 s |
| Compact weights | 948,196 B | reduce a + reduce b + multiply | 100/100, identical outputs | 34.59 s* |
| Compact + short schedule | 948,196 B | reduce a + multiply by streaming b | 100/100, identical outputs | 20.45-21.31 s |
| Compact + short schedule + static-channel reuse | 948,196 B | reduce a + multiply by streaming b | 100/100 | 19.07 s |
`*` This single MPS timing is noisy and does not imply bfloat16 storage makes
inference slower. The compact tensors are loaded back into fp32 parameters;
storage precision affects artifact size, not the execution dtype.
The short-schedule variant also scored 100/100 on each public Tier 1 through 7
(700/700 total) and passed the official static checker. Full Tier 8-10 and
held-out adversarial validation remain release blockers.