File size: 1,600 Bytes
9d3ee31 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 | # Mach-1-Ternary-Additive-35B — benchmark board
Payload: integer L1-ball trellis expert codes + per-wavefront gamma scales +
continuous fp16 su/sv side-streams; 64-level integer-lattice spine; int5-g64
head; int4 embed — every weight matmul is add/subtract-only. Ship gate:
decode.py-primitive reconstruction == served checkpoint (bf16 rounding,
layers 0/20/39).
Protocol: same-harness EvalScope 1.9.1 + vLLM, PrismML App. B (thinking mode,
temp 1.0, top_p 0.95, top_k 20, PrismML token tiers, AIME mean-of-8, IFEval
prompt-strict, IFBench prompt-loose). Retention = 100 x score / BF16 teacher
(Qwen3.6-35B-A3B), same harness. tau2-bench = fixed external user-simulator
(Qwen3.6-35B BF16, greedy), single pass, identical for every model.
## Flagship payload (12/12)
| benchmark | score | teacher | retention |
|---|---|---|---|
| AIME25 @8 | 87.50 | 88.33 | **99.1%** |
| AIME26 @8 | 89.58 | 90.00 | **99.5%** |
| MATH-500 | 98.00 | 98.60 | **99.4%** |
| GSM8K | 94.69 | 96.21 | 98.4% |
| MBPP+ | 94.44 | 96.03 | 98.3% |
| HumanEval+ | 92.68 | 95.12 | 97.4% |
| MMLU-Redux | 89.18 | 92.68 | 96.2% |
| IFEval | 83.75 (n=5) | 89.05 | 94.0% |
| MuSR | 61.77 | 66.66 | 92.7% |
| BFCL-v3 | 68.97 | 74.98 | 92.0% |
| tau2-bench | 71.58 | 79.51 | 90.0% |
| IFBench | 54.08 | 64.97 | 83.2% |
| **mean retention** | | | **95.0%** |
## Read-quality notes
- Multi-read cells quote the mean over all reads with n; single reads are n=1.
- IFEval strict single-read spread measured ~2-3 pts; tau2 complete-read
spread up to ~6 on some artifacts — sub-point deltas are ties.
- Expert payload 6.207 GB.
|