Quantized Gemma 4 models, including the 31B imatrix MTP GGUF.
Note W4A16 (full lm_head) + MTP drafter, ~130 tok/s single-GPU vLLM, tool-calling verified