GLM-4.7-Flash models with MTP

#2751
by jacek2024 - opened

If you update llama.cpp to include this PR:

https://github.com/ggml-org/llama.cpp/pull/24868

then regenerating GLM-4.7-Flash will produce GGUFs with MTP support:

https://huggingface.co/zai-org/GLM-4.7-Flash

You could also regenerate some popular finetunes, for example:

https://huggingface.co/huihui-ai/Huihui-GLM-4.7-Flash-abliterated
https://huggingface.co/cerebras/GLM-4.7-Flash-REAP-23B-A3B
https://huggingface.co/koute/GLM-4.7-Flash-Derestricted
etc, etc
(but I assume standalone small MTP gguf is enough for them)

So far, I have only uploaded this one:

https://huggingface.co/jacek2024/GLM-4.7-Flash-MTP-GGUF

Sign up or log in to comment