maniac-11111's picture
final
9d3ee31
|
Raw
History Blame Contribute Delete
1.6 kB

Mach-1-Ternary-Additive-35B — benchmark board

Payload: integer L1-ball trellis expert codes + per-wavefront gamma scales + continuous fp16 su/sv side-streams; 64-level integer-lattice spine; int5-g64 head; int4 embed — every weight matmul is add/subtract-only. Ship gate: decode.py-primitive reconstruction == served checkpoint (bf16 rounding, layers 0/20/39).

Protocol: same-harness EvalScope 1.9.1 + vLLM, PrismML App. B (thinking mode, temp 1.0, top_p 0.95, top_k 20, PrismML token tiers, AIME mean-of-8, IFEval prompt-strict, IFBench prompt-loose). Retention = 100 x score / BF16 teacher (Qwen3.6-35B-A3B), same harness. tau2-bench = fixed external user-simulator (Qwen3.6-35B BF16, greedy), single pass, identical for every model.

Flagship payload (12/12)

benchmark score teacher retention
AIME25 @8 87.50 88.33 99.1%
AIME26 @8 89.58 90.00 99.5%
MATH-500 98.00 98.60 99.4%
GSM8K 94.69 96.21 98.4%
MBPP+ 94.44 96.03 98.3%
HumanEval+ 92.68 95.12 97.4%
MMLU-Redux 89.18 92.68 96.2%
IFEval 83.75 (n=5) 89.05 94.0%
MuSR 61.77 66.66 92.7%
BFCL-v3 68.97 74.98 92.0%
tau2-bench 71.58 79.51 90.0%
IFBench 54.08 64.97 83.2%
mean retention 95.0%

Read-quality notes

  • Multi-read cells quote the mean over all reads with n; single reads are n=1.
  • IFEval strict single-read spread measured ~2-3 pts; tau2 complete-read spread up to ~6 on some artifacts — sub-point deltas are ties.
  • Expert payload 6.207 GB.