Commit History

major revision: runtime dominates (MTPLX 3.79x vs oMLX 54.6 t/s on the same artifact); void the FP16-vs-4bit conclusion as runtime-confounded; add contract-parameter verification; correct the oMLX depth claim
7161354
verified

KaedeTai commited on

correct two errors: oMLX depth is adaptive (cap 8) and parked at 0 on the MoE, not capped at 3; the Ornith random-init case was fixed upstream 2026-08-23
2186674
verified

KaedeTai commited on

bound the MoE claim (EryriLabs report +29% on llama.cpp, deeper drafting); add random-init detection — counting tensors is not enough
f649fcf
verified

KaedeTai commited on

state per-file scope up front; the head is Qwen3.8-27B dense only, the tool is not
ebaba5b
verified

KaedeTai commited on

add MTP acceptance-rate data; MoE graft activates but is break-even; correct claim that oMLX exposes no acceptance metric
3b964fc
verified

KaedeTai commited on

MTP head graft for MLX Qwen3.8-27B quantizations that dropped theirs
dc03fdf
verified

KaedeTai commited on

initial commit
66619c6
verified

KaedeTai commited on