🔧 Patched chat template

chat_template.jinja differs from Google's Gemma 4 Canonical Chat Template (2026-07-09) by one intentional change; everything else is untouched.

The canonical template suppresses the thinking channel at the start of a normal model turn, but emits nothing after a tool response when enable_thinking is false. The model may then open a thinking channel on its own, and a quantized model sometimes writes the literal word thought into the answer. This patch gives the tool_response branch the same suppression:

{%- elif ns.prev_message_type == 'tool_response' -%}
    {%- if enable_thinking -%}
        {{- '<|channel>thought\n' -}}
    {%- else -%}
        {{- '<|channel>thought\n<channel|>' -}}
    {%- endif -%}
{%- endif -%}

Only that case changes — the other prompt paths render byte-identical to the canonical template. To get stock behaviour, replace chat_template.jinja with the one from the base model repo; the weights are unaffected.

Downloads last month
172
Safetensors
Model size
2B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ToPo-ToPo/gemma-4-12B-it-qat-mlx-4bit

Quantized
(56)
this model