Full weight conversion of Google Gemma 4 31B to MLX using mlx-vlm 0.4.5 commit e41cd25 and latest Google chat template and tokenizer as of May 1, 2026.
This model has slightly different safe tensor sizes compared to MLX-Community's version. Do I know why? No. But maybe it's because I was using the latest mlx-vlm. This model also solved an image handling issue I had with the mlx-community conversion. This also may be in my head. But I figure I'd upload the latest mac version of this very capable model.
UPDATE JULY 17, 2026
- Added upgraded Google Chat Templete for better tool calling
- Increased vision default to 560 so the model is less blind
- Downloads last month
- 36
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support