Instructions to use ToPo-ToPo/gemma-4-12B-it-qat-mlx-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use ToPo-ToPo/gemma-4-12B-it-qat-mlx-4bit with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir gemma-4-12B-it-qat-mlx-4bit ToPo-ToPo/gemma-4-12B-it-qat-mlx-4bit
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
🔧 Patched chat template
chat_template.jinja differs from Google's Gemma 4 Canonical Chat Template
(2026-07-09) by one intentional change; everything else is untouched.
The canonical template suppresses the thinking channel at the start of a normal model
turn, but emits nothing after a tool response when enable_thinking is false. The model
may then open a thinking channel on its own, and a quantized model sometimes writes the
literal word thought into the answer. This patch gives the tool_response branch the
same suppression:
{%- elif ns.prev_message_type == 'tool_response' -%}
{%- if enable_thinking -%}
{{- '<|channel>thought\n' -}}
{%- else -%}
{{- '<|channel>thought\n<channel|>' -}}
{%- endif -%}
{%- endif -%}
Only that case changes — the other prompt paths render byte-identical to the canonical
template. To get stock behaviour, replace chat_template.jinja with the one from the
base model repo; the weights are unaffected.
- Downloads last month
- 172
4-bit
Model tree for ToPo-ToPo/gemma-4-12B-it-qat-mlx-4bit
Base model
google/gemma-4-12B