card: GSM8K 100q across all three runtimes (98/97/95) + the two-pass budget note ce8bdda verified mlboydaisuke commited on 14 days ago
fix: declare <|eot|> as a stop token β the bundle never terminated 226f682 verified mlboydaisuke commited on 14 days ago
card: correct the machine β Mac Studio M4 Max, not MacBook Pro ad1a2e1 verified mlboydaisuke commited on 14 days ago
card: add the raw-MLX arm β Core AI ties MLX, both beat the ExecuTorch build by 14% dcda323 verified mlboydaisuke commited on 15 days ago
Card: lossless n-gram spec-decode on this bundle (1.34x chat / 1.96x code / 1.84x tool calling), no drafter, no re-export a967f69 verified mlboydaisuke commited on 15 days ago
card: same-machine head-to-head vs the ExecuTorch Metal build 3dbe5c7 verified mlboydaisuke commited on 15 days ago
Muse-Glimmer-30B text decoder -> Core AI: decode int4hu (--head-sym), token-exact vs fp16 oracle, 26.69 tok/s on M4 Max d284183 verified mlboydaisuke commited on 15 days ago