Custom gemma models for coding
Note W4A16 (full lm_head) + MTP drafter, ~130 tok/s single-GPU vLLM, tool-calling verified