ToPo-ToPo's picture
Trim card to provenance + minimal usage
22ff1a3 verified
|
Raw
History Blame Contribute Delete
1.54 kB
---
license: apache-2.0
base_model: google/diffusiongemma-26B-A4B-it
library_name: mlx
tags: [mlx, vision, moe, diffusion-llm]
pipeline_tag: image-text-to-text
---
# ToPo-ToPo/diffusiongemma-26B-A4B-it-mlx-4bit
MLX **4bit** conversion of [`google/diffusiongemma-26B-A4B-it`](https://huggingface.co/google/diffusiongemma-26B-A4B-it) (mlx-vlm).
Block-diffusion LM built on Gemma 4 (25.2B total / 3.8B active, MoE 128+1 experts, vision).
## Provenance (self-converted)
- Source: `google/diffusiongemma-26B-A4B-it` (license: apache-2.0)
- Tool: mlx-vlm 0.6.9 `mlx_vlm.convert` (4bit affine, group_size=64), ~5.130 bpw
- `model_type: diffusion_gemma` is supported natively by mlx-vlm 0.6.9; no patch needed.
- `chat_template.jinja` is **not** the base repo's copy: it is the patched *Gemma 4 Canonical
Chat Template* from [`ToPo-ToPo/gemma-4-26B-A4B-it-mlx-4bit`](https://huggingface.co/ToPo-ToPo/gemma-4-26B-A4B-it-mlx-4bit),
which suppresses the thinking channel when `enable_thinking` is false (otherwise the literal
word `thought` leaks into the answer). Thinking is off by default. Weights are unaffected —
restore the base repo's template for stock behaviour.
## Usage
```python
from mlx_vlm import load
model, processor = load("ToPo-ToPo/diffusiongemma-26B-A4B-it-mlx-4bit")
```
Diffusion generation takes its own flags:
```bash
python -m mlx_vlm generate --model ToPo-ToPo/diffusiongemma-26B-A4B-it-mlx-4bit \
--prompt "Why is the sky blue?" \
--max-tokens 256 --max-denoising-steps 48 --diffusion-sampler entropy-bound
```