DiffGemma 26B-A4B-it β nvfp4 (.dgq)
A quantized, self-contained .dgq pack of
DiffusionGemma 26B-A4B-it
(Gemma-4 26B-A4B MoE, discrete block diffusion) for the
diffgemma-mps Rust + Metal inference
engine on Apple Silicon. Text-only (v1).
The most compressed profile: MoE experts and attention + dense FFN are all NVFP4.
Quantization β nvfp4 profile
| Tensor class | Format | Bits/weight |
|---|---|---|
| MoE experts | nvfp4_block (NVFP4, group 16) |
~4.5 |
| Attention (q/k/v/o), dense FFN | nvfp4_block (NVFP4, group 16) |
~4.5 |
| Embeddings, router, norms | bf16 | 16 |
| Self-conditioning MLP | q8_row |
8 |
~15 GiB of weights. Manifest version 2 (nvfp4) β loads on any diffgemma-mps
build with NVFP4 support.
Usage
diffgemma-mps download --repo mmastrac/diffgemma-26b-a4b-it-nvfp4 -o model/diffgemma-26b-a4b-it-nvfp4
diffgemma-mps -m model/diffgemma-26b-a4b-it-nvfp4 chat
Build and run details: github.com/mmastrac/diffgemma.
License
Gemma Terms of Use. Derived from
google/diffusiongemma-26B-A4B-it; use is subject to the Gemma license.
- Downloads last month
- 14
Model tree for mmastrac/diffgemma-26b-a4b-it-nvfp4
Base model
google/diffusiongemma-26B-A4B-it