gemma-4-E4B-it W4A16 (full-precision lm_head), calibrated on our own CLI/tool-call logs, + Gemma-4 MTP drafter for speculative decoding. ~130 tok/s single-stream on one RTX 4060 Ti, stock vLLM 0.26.0, correct tool calls, no vocab pruning.
verified
pearsonkyle commited on