GLM-4.7-Flash models with MTP
#2751
by jacek2024 - opened
If you update llama.cpp to include this PR:
https://github.com/ggml-org/llama.cpp/pull/24868
then regenerating GLM-4.7-Flash will produce GGUFs with MTP support:
https://huggingface.co/zai-org/GLM-4.7-Flash
You could also regenerate some popular finetunes, for example:
https://huggingface.co/huihui-ai/Huihui-GLM-4.7-Flash-abliterated
https://huggingface.co/cerebras/GLM-4.7-Flash-REAP-23B-A3B
https://huggingface.co/koute/GLM-4.7-Flash-Derestricted
etc, etc
(but I assume standalone small MTP gguf is enough for them)
So far, I have only uploaded this one: