MTP layer?

#45
by puchuu - opened

Are you planning to release MTP layer for Muse Glimmer 30b? Thank you.

PS for now I am running Qwen 3.6 27B Q8_0 with MTP draft 3 split tensor at 35-40 t/s output dual 9060 XT 16GB.

it comes with dflash built in. you need to download and specify the model.

It is interesting, I am going to try 'draft-dflash' in llama.cpp. Thank you!

I launched unsloth muse glimmer Q8_0 draft-dflash 3 split tensor at 28-33 t\s output dual 9060 XT 16GB. I am using the latest llama.cpp and rocm 7.14.0. So performance is a bit lower than Qwen, but doesn't really matter. It is ok. Thanks everyone!

puchuu changed discussion status to closed

Sign up or log in to comment