Quick question: any plans for Q4_K_M / other quants?

#1
by MisterFynn - opened

Hi! Really impressive work with this franken-merge. Grafting the 20 MTP layers from the unsloth Qwen3.6-35B-A3B-MTP model onto Bartowski's GGUF builds is a clever approach, and the ~20–30% inference speed gain on AMD hardware is exactly what I've been looking for. Thanks for sharing the model and the conversion script!

I noticed you currently have the IQ4_XS variant up (Kwaipilot_KAT-Coder-V2.5-Dev-IQ4_XS.bartowski.mtp.gguf). I was wondering how you've found IQ4_XS comparing to Q4_K_M in terms of quality and stability? I know IQ4_XS saves a bit of VRAM and can be slightly faster on AMD, but I've also heard Q4_K_M tends to be more consistent across different workloads.

Would you consider releasing a Q4_K_M (or even Q5_K_M) version as well? Lower-bit options would be incredibly helpful for anyone working with tighter VRAM targets, and even a couple of key variants would make a big difference for broader accessibility.

If you're open to it, I'd be happy to help test or benchmark any new quants. Either way, thanks again for pushing this forward β€” really appreciate the effort!

Keep up the great work!

Sign up or log in to comment