This is the final Repo.

#1
by gbuzhf - opened

Original MTP grafted into KAT quantizations -> Bartowski's imatrix -> results were fair.

I then started native-MTP training, and the measurements were promising: 30%+ performance increase, with a native imatrix blended into Bartowski's.
I deleted the first repo and kept that one while I ran real-world tests.

In real-world measurements (real harness, temp environment), the native-MTP experiment failed against the original MTP.

The native-MTP repo is now deleted. Back to: original MTP grafted into KAT quantizations -> blended imatrix. This is the final repo and it will be kept.

-- My reasoning was:

I had been following the Kimi K2 -> K3 and MiniMax experiment papers.

Basically:

Qwen 3.6 35B-A3B = 1.0
Original MTP     = 1.0
KAT Coder V2.5 Dev = 1.2 (fine-tuned Qwen)
So if we fine-tune the original MTP on KAT's outputs and hidden states, it should land at 1.2, and we'd recover the original performance there.

Synthetic benchmarks worked.

Real world: failed.

Therefore: experiment documented and removed.

gbuzhf pinned discussion
gbuzhf locked this discussion

Sign up or log in to comment