AX-gemma-4-12b-MLX-AXQ-4bit-MTP

Checkpoint Tier 1 certified on df-macbookpro-m5 (2026-08-09). Rebuilt from google/gemma-4-12b-it (not the non-IT google/gemma-4-12b base). Size + quality retention ≥0.98 on authorizing host. MTP acceleration is not certified. Certificate: gemma4-12b-axq4-tier1.md.

Property Value
Product class AXQ 4bit
Source google/gemma-4-12b-it
Measured BPW 4.9001
Size ratio vs uniform 0.6667× (max 1.15)
Quality agent-coding retention 1.0325
Quality general retention 1.0000
Manifest SHA-256 bc3fde05ba2af82c303b3ef13f1082586fb15e0735bccb050e9768c9d19dcc75
Engine pair gemma-4-12b-itgemma-4-12b-it-assistant
MTP layout assistant/ + ax_gemma4_assistant_mtp.json

Certification

Claim Status
Checkpoint Tier 1 Certified on df-macbookpro-m5
MTP acceleration Tier 2 Not certified
Vision / multimodal quality Not claimed

Layout

./                          # AXQ target weights (+ vision sidecar)
assistant/                  # gemma4_assistant drafter
ax_gemma4_assistant_mtp.json
ax_composite_pack_manifest.json
axquant_*.json

Runtime

export AX_MLX_GEMMA4_ASSISTANT_MTP=1
export AX_MLX_GEMMA4_ASSISTANT_MTP_MAX_DEPTH=2

Default product text route remains direct decode unless assistant-MTP is explicitly enabled.

Published / recertified: 2026-08-09.

Downloads last month
29
Safetensors
Model size
2B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support