File size: 1,543 Bytes
bd0a489
 
 
 
22ff1a3
bd0a489
 
 
22ff1a3
 
bd0a489
 
22ff1a3
 
 
 
 
 
 
 
bd0a489
 
 
 
 
 
22ff1a3
bd0a489
22ff1a3
bd0a489
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
---
license: apache-2.0
base_model: google/diffusiongemma-26B-A4B-it
library_name: mlx
tags: [mlx, vision, moe, diffusion-llm]
pipeline_tag: image-text-to-text
---
# ToPo-ToPo/diffusiongemma-26B-A4B-it-mlx-4bit
MLX **4bit** conversion of [`google/diffusiongemma-26B-A4B-it`](https://huggingface.co/google/diffusiongemma-26B-A4B-it) (mlx-vlm).
Block-diffusion LM built on Gemma 4 (25.2B total / 3.8B active, MoE 128+1 experts, vision).

## Provenance (self-converted)
- Source: `google/diffusiongemma-26B-A4B-it` (license: apache-2.0)
- Tool: mlx-vlm 0.6.9 `mlx_vlm.convert` (4bit affine, group_size=64), ~5.130 bpw
- `model_type: diffusion_gemma` is supported natively by mlx-vlm 0.6.9; no patch needed.
- `chat_template.jinja` is **not** the base repo's copy: it is the patched *Gemma 4 Canonical
  Chat Template* from [`ToPo-ToPo/gemma-4-26B-A4B-it-mlx-4bit`](https://huggingface.co/ToPo-ToPo/gemma-4-26B-A4B-it-mlx-4bit),
  which suppresses the thinking channel when `enable_thinking` is false (otherwise the literal
  word `thought` leaks into the answer). Thinking is off by default. Weights are unaffected —
  restore the base repo's template for stock behaviour.

## Usage
```python
from mlx_vlm import load
model, processor = load("ToPo-ToPo/diffusiongemma-26B-A4B-it-mlx-4bit")
```
Diffusion generation takes its own flags:
```bash
python -m mlx_vlm generate --model ToPo-ToPo/diffusiongemma-26B-A4B-it-mlx-4bit \
  --prompt "Why is the sky blue?" \
  --max-tokens 256 --max-denoising-steps 48 --diffusion-sampler entropy-bound
```