HiramKHackenbacker's picture
model card: standard structure, measured facts, verified license declarations
22deb78 verified
|
Raw
History Blame Contribute Delete
20.8 kB
---
license: apache-2.0
license_link: https://ai.google.dev/gemma/docs/gemma_4_license
base_model: google/gemma-4-26B-A4B-it-qat-q4_0-unquantized
base_model_relation: quantized
library_name: coreai
pipeline_tag: text-generation
tags:
- core-ai
- aimodel
- apple-silicon
- on-device
- coreai-kit
- quantized
- int4
- qat
- moe
- mixture-of-experts
- gemma4
---
# Gemma 4 26B-A4B (MoE) β€” Core AI (.aimodel)
`google/gemma-4-26B-A4B-it-qat-q4_0-unquantized` converted to Core AI `.aimodel` bundles for
Apple silicon by [visible-cx](https://huggingface.co/visible-cx). These are derivative
artifacts: Google's QAT-trained weights rounded onto int4 and re-expressed as a Core AI graph,
with the 128-expert sparse branch of every layer lowered onto a Metal gather kernel. They load
through Core AI on macOS and are not usable by PyTorch, GGUF or MLX.
It is a **sparse mixture-of-experts** model β€” 26.5B total parameters, roughly 4B active per
token. This is a **new port**, not a reproduction of a published zoo bundle: three source
modules were authored for it (below), and nothing in the shared toolchain was modified.
> ⚠️ **UNQUALIFIED. Neither bundle has ever produced a token, on any machine.** The graphs
> export, the producer fingerprints pass, and the MoE block wiring is verified against the
> reference implementation to 3.5e-7 relative error. Nothing else is established.
>
> ⚠️ **The 16 GB tier is excluded.** Weights are ~17.6 GB resident against a 16 GB Mac's
> ~10.7 GB Metal working set. **Minimum practical machine memory: 32 GB at shallow context,
> 64 GB at the manifest's 16384.**
> ⚠️ **Expected known issue β€” long unguided generations (unverified on this model).** On the
> sibling E2B/E4B exports the Core AI runtime retains one compiled specialization per
> sequence-length signature, costing roughly 81 MB of GPU allocations per generated token until
> process exit, and the decode graphs here share the length-dependent structure that causes it.
> Guided decoding is not a defence β€” **shape reuse is**. Measured details: the
> [gemma-4-E2B-CoreAI](https://huggingface.co/visible-cx/gemma-4-E2B-CoreAI) card.
## Contents
| Path | Bytes | Files | Context | Functions | Status |
|---|---:|---|---|---|---|
| `gpu-pipelined/gemma4_26b_a4b_qat_decode_int4linsym_moeaff4_msdpa_g8` | 17,580,059,414 | 8 | 16384 | main (decode only) | UNQUALIFIED |
| `hybrid-mf64/gemma4_26b_a4b_qat_decode_int4linsym_moeaff4_msdpa_g8_pf64` | 17,580,916,612 | 8 | 16384 | main + prefill | EXPERIMENTAL |
`gpu-pipelined/`, file by file:
| File | Bytes |
|---|---:|
| `…aimodel/main.mlirb` | 17,547,866,346 |
| `…aimodel/main.hash` | 32 |
| `…aimodel/metadata.json` | 105 |
| `metadata.json` (bundle manifest) | 689 |
| `tokenizer/tokenizer.json` | 32,169,626 |
| `tokenizer/chat_template.jinja` | 18,683 |
| `tokenizer/tokenizer_config.json` | 3,730 |
| `tokenizer/generation_config.json` | 203 |
`hybrid-mf64/` is the same layout with `main.mlirb` at 17,548,723,515 B β€” **857,169 B more than
the decode-only build**, which is the whole cost of a second entrypoint over deduplicated
weights.
Manifest context is 16384. `--max-ctx` sets `language.max_context_length` and nothing else;
lower it if the machine cannot afford the KV β€” that is a metadata edit, not a re-export.
**Stop token:** `eos_token = "<turn|>"` (id 106), the turn terminator Gemma 4 emits, applied by
the export script itself. A host that stops on the raw upstream `<eos>` will overrun every
reply.
## Provenance
| | |
|---|---|
| Base checkpoint | `google/gemma-4-26B-A4B-it-qat-q4_0-unquantized` @ `f1e06dc520982d9b9edd76859fdb7ab209449949` |
| Checkpoint size | 51,644,341,801 B, bf16, 1,013 tensors across 2 shards; ungated |
| Recipe | `export_gemma4_moe_decode_pipelined.py int4lin --lin-sym --metal-sdpa --max-ctx 16384` |
| Toolchain base | `apple/coreai-models` @ `b1cb71b8522d99408059fa0b98b8742171bcb0b8` + the [coreai-model-zoo](https://github.com/john-rocky/coreai-model-zoo) python overlay, plus three new modules (below) |
| Toolchain | `coreai-torch 0.4.1`, `coreai-core 1.0.0b2`, `coreai-opt 0.2.1`, `torch 2.9.0` |
| Producer fingerprint | `coreai-core 1.0.0b2` β€” verified on the inner asset `metadata.json` |
| Asset creation date | `20260817T215629Z` (`gpu-pipelined/`), `20260818T141434Z` (`hybrid-mf64/`) |
| Vocab | 262,144 |
"QAT-unquantized" means QAT-*trained*, stored bf16; the int4 rounding happens at export onto the
ggml q4_0 grid the training already targeted. There is no separate pre-quantised int4 artifact
to prefer.
**To reproduce this conversion you need a β‰₯128 GB host**: the decode-only export peaks at
**102.25 GB RSS** (`VmHWM`) on Linux x86_64; the `hybrid-mf64/` export peaked at **87.5 GB** in
305 s.
### Architecture
| | 26B-A4B | 31B | E4B |
|---|---|---|---|
| `hidden_size` | **2816** | 5376 | 2560 |
| layers | **30** (25 sliding / 5 full) | 60 | 42 |
| dense MLP `intermediate_size` | **2112** | 21504 | 10240 |
| `num_experts` / `top_k` | **128 / 8** | β€” | β€” |
| `moe_intermediate_size` | **704** | β€” | β€” |
| attention heads | **16** | 32 | 8 |
| KV heads (sliding / full) | **8 / 2** | 16 / 4 | 2 / β€” |
| `head_dim` / `global_head_dim` | **256 / 512** | 256 / 512 | 256 / 512 |
| `sliding_window` | **1024** | 1024 | 512 |
| Per-Layer Embeddings | **none** | none | 256 |
Same attention family as the dense 12B/31B β€” same dual head_dim, same `attention_k_eq_v` (full
layers carry no `v_proj`), same dual RoPE, same softcap. **There are no Per-Layer Embeddings and
therefore no gather-table sidecar in this repo**, unlike E2B/E4B.
Every layer carries a dense MLP **and** a sparse branch, in parallel:
```python
residual = x # post-attention hidden
h = mlp(pre_feedforward_layernorm(x)) # dense branch
h1 = post_feedforward_layernorm_1(h)
w, idx = router(residual) # routes on the RAW residual
h2 = experts(pre_feedforward_layernorm_2(residual), idx) # sparse branch
h2 = post_feedforward_layernorm_2(sum_k w_k * h2_k)
x = residual + post_feedforward_layernorm(h1 + h2)
x = x * layer_scalar
```
Two details are load-bearing: **the router reads the un-normalised residual**, and **the experts
are GELU-gated** (`gelu_pytorch_tanh`), where the `SwitchGLU` primitive's default is SiLU. The
router itself: scale-free RMSNorm β†’ `Γ— scale Γ— hidden**-0.5` β†’ linear to 128 β†’ fp32 softmax over
all experts β†’ top-8 β†’ renormalise to sum 1 β†’ multiply by a learned `per_expert_scale` gathered
at the selected ids.
### Compression recipe
| Weights | Scheme | Rationale |
|---|---|---|
| **Routed experts** (128/layer, 22.84 G params) | **affine int4** (`aff4`), per-(output row, K-block-32) scale + bias, MSE-optimal clip over 6 candidates | see below |
| Attention, dense MLP, untied LM head | linear **int4** per-block-32, plain absmax (`--lin-sym`) | the ggml q4_0 grid the QAT checkpoint was trained on β€” the 12B/31B recipe verbatim |
| **Router** (`proj`, `scale`, `per_expert_scale`) | **fp16**, excluded by name | routers are the quantization-sensitive part of an MoE and the whole set is 0.02 GB |
| `embed_tokens` (in-graph) | fp16 | the gather stays exact; the head is untied and quantized separately |
`sym8` on the experts would put them at 22.8 GB before anything else, so int4 is not optional.
`aff4` was chosen over a k-means palette (`km4`) because this checkpoint was QAT-trained on
per-block-32 q4_0 (`w = (qβˆ’8)Β·d`, 16 uniform levels), and an affine int4 block-32 grid represents
that **exactly** (`scale = d`, `bias = βˆ’8d`). A palette fits one 16-entry codebook per 32 output
rows across the whole K axis and structurally cannot follow per-block QAT scales. The price is
aux bytes: `aff4` stores an fp16 scale *and* bias per 32-element K block, ~5 bits/param effective
against `km4`'s ~4.03 β€” about 3 GB of the bundle.
The `down` projection's K of 704 is padded to 768 by the kernel's automatic `_kpad256` treatment
β€” 64 zero columns, ~0.5% wasted expert bytes. `gate`/`up` need no padding.
**The gather kernel is load-bearing.** Without `MetalSwitchGLU`, the sparse branch lowers to a
dense matmul that reads all 128 experts every token β€” a 16Γ— over-read.
### New source modules
| File | Contents |
|---|---|
| `coreai_models/models/macos/gemma4_moe_text.py` | `Gemma4MoeConfig`, `Gemma4MoeRouter`, `GeluGLU`, `Gemma4MoeDecoderLayer`, `Gemma4MoeForCausalLM`, and a loader that splits the checkpoint's fused `gate_up_proj [E, 2*704, 2816]` into the `SwitchGLU`'s separate stacks |
| `coreai_models/models/macos/gemma4_moe_pipelined.py` | `Gemma4MoePipelinedForCausalLM` β€” subclasses the dense pipelined core and overrides one method; the attention half is inherited unchanged |
| `coreai-model-zoo/conversion/export_gemma4_moe_decode_pipelined.py` | the recipe |
## Requirements
- **Apple silicon Mac with β‰₯32 GB unified memory**, Core AI runtime.
- **Engine contract: 2 inputs** (`input_ids`, `position_ids`) β†’ logits, one growing KV pair, no
static inputs and no per-step mask. Verified by reading the graph bytecode of `hybrid-mf64/`:
`input_ids`, `position_ids`, `keyCache`, `valueCache`, `logits`, `prefill`, 90
`gather_qmm_int4aff` call sites (30 layers Γ— gate/up/down) and `gemma4_dense_full_sdpa_occ8`
β€” **2 inputs, no `mask`**, which is the arity the sequential engine requires.
- **States:** `keyCache` / `valueCache` `Float16, 30 Γ— 1 Γ— 8 Γ— ? Γ— 512` β€” one growing pair, 30
slots (no KV sharing). Dynamic sequence dim β†’ `GrowingKVCache`.
- **KV cost: 491,520 bytes per token of context** (fp16) β€” 2.01 GB at 4096, 8.05 GB at 16384.
- **Resident weights, ~17.6 GB:**
| | bytes |
|---|---|
| Routed experts, `aff4` (incl. `_kpad256` zero columns) | ~14.7 GB |
| Attention + dense MLP + untied head, int4 block-32 | ~1.4 GB |
| `embed_tokens`, fp16 in-graph | ~1.5 GB |
| Router, fp16 | ~0.02 GB |
| **Total** | **~17.6 GB** |
| Context | KV | + weights | Verdict |
|---|---|---|---|
| 4096 | 2.01 GB | 19.6 GB | 32 GB machine |
| 8192 | 4.03 GB | 21.6 GB | 32 GB machine, marginal against a ~21–24 GB working set |
| 16384 (this manifest) | 8.05 GB | 25.6 GB | **64 GB machine** |
| 16 GB machine | β€” | β€” | **no, at any context** |
**This table is arithmetic from state shapes and weight formats β€” a prediction to check, not a
measured tier claim.** The residency mechanism for MoE bundles on macOS has never been
measured, and this project's one check of a bundle-bytes proxy against a real compile found it
**1.35 GiB optimistic** on a large bundle.
25 of the 30 layers have `head_dim` 256 zero-padded to 512, and all 30 slots grow linearly even
though 25 of them only ever attend a 1024-token window. A ring-buffered sliding cache would be
roughly 5Γ— cheaper per token; that is model authoring, not a flag.
- The bundle manifest declares `runtime_env COREAI_CHUNK_THRESHOLD=1`.
## Measurements
**None. No measurement of any kind exists, on any machine, for either bundle.** Every figure in
this card is a byte count, an export-host reading, a box-side torch comparison, or arithmetic
from state shapes and weight formats. No token has been generated.
The intended argument for this model is throughput at comparable size against the dense 31B β€”
17.6 GB vs 20 GB, with ~4B active parameters per token read through a gather kernel β€” and
quality per token against the 12B at ~2.1Γ— the footprint. The nearest *measured* analogue is
LFM2.5-8B-A1B at 140 tok/s through its gather kernel against 39 tok/s for the same weights read
as a dense over-read. **Neither claim is measured for this bundle.**
One datapoint about the *checkpoint*, clearly scoped because it is not about these bundles: the
same model served remotely (Cloudflare's `@cf/google/gemma-4-26b-a4b-it`) answered eleven real
report prompts 11/11 schema-valid under a tuned recipe and passed two grounding probes that the
local 8B-A1B fails β€” it identified a vendor-versus-subject confusion correctly and attributed an
opinion rather than asserting it as fact. That is evidence about the weights' grounding, on a
completely different serving path, and says nothing about this Core AI conversion.
### What *has* been gated, box-side
**MoE block wiring, against the reference implementation.** `transformers` 4.57.6 (the pinned
export environment) has no `gemma4`, so the gate transcribes `Gemma4TextDecoderLayer.forward`
(the MoE-block half), `Gemma4TextRouter.forward` and `Gemma4TextExperts.forward` from upstream
and runs both implementations on the real layer-0 weights in fp32:
```
T=1: max|dFFN|=7.629e-05 rel=3.453e-07 | same top-8 ids: True max|dw|=0.000e+00
T=4: max|dFFN|=9.537e-05 rel=2.735e-07 | same top-8 ids: True max|dw|=0.000e+00
```
Identical expert selection, bit-identical routing weights, ~3.5e-7 relative error on the block
output. **This gates the authoring** β€” the wiring, the fused-weight split, the activation choice,
the norm placement, the router maths. It does **not** gate the exported graph, the int4 rounding,
or either Metal kernel.
**Attention kernel numerics** (`hybrid-mf64/`). The 26B is a *third* distinct block-GQA mapping
and was gated as its own case: 16 query heads, **2** global KV heads replicated across 8 cache
slots, so the kernel's `kv = h / (H / slots)` = `h//2` must land on real head `(h//2)//4 = h//8`.
| | vs the MPSGraph composite | vs an fp32 ground truth |
|---|---|---|
| max abs, S=137 | 2.50e-2 | **9.74e-4 β€” PSNR 74.9 dB** |
| max abs, S=1024 | 1.17e-2 | 9.76e-4 β€” 75.1 dB |
| the composite itself, vs fp32 | β€” | 53.0 and 57.9 dB |
The divergence from the composite is the composite's own fp16 error; against fp32 the kernel is
the more accurate of the two. The 12B (1 global head, rep 8) and 31B (4 global heads, rep 4 over
16 slots) rows are in the same file as controls β€” a mapping bug that cancelled in those would
show here. This gates the mapping and the scale, **not the Metal source** (no GPU on the box).
**Trace accounting** (`hybrid-mf64/`): all 5 full layers took the metal branch at S=1 **and** the
composite branch at S=64, 5/5 both ways β€” the export fails if either count is 0. 30 MoE layers
metalized.
## `hybrid-mf64/` β€” decode and prefill from one copy of the weights
The 26B-A4B with a `prefill` entrypoint. Two things had to be per-entrypoint at once, and both
are:
- **Experts.** `BatchedMetalSwitchGLU` sorts the `S*k` (token, expert) pairs of a 64-token chunk
by expert id so each expert slab is read once, and falls through to `MetalSwitchGLU.forward`
verbatim whenever `b*s == 1`. One metalization, both entrypoints, **decode arithmetic
unchanged** β€” `main` traces the identical q=1 gather matvec the published decode bundle
carries.
- **Attention.** The 26B inherits the dense 12B/31B full-attention layers, so it needs the
flash-decode Metal kernel for the scratch-heap reason, and that kernel is structurally q=1.
`SplitSDPA` branches on `query.shape[2]`, a concrete int in each trace, so `main` gets the
kernel and `prefill` gets the MPSGraph composite (query axis + `is_causal=True` intra-chunk
mask). Same technique as the 31B `hybrid-pf64/` bundle.
**This is `mf64`, not `mf64-tp`, and it cannot be `-tp` today.** The token-tiled prefill kernel
is **sym8-only** β€” `TiledBatchedMetalSwitchGLU.__init__` raises for any other scheme β€” and the
26B ships `aff4` experts for the QAT-grid reason above. An `aff4` tiled kernel is future work,
gated on the 8B `mf64-tp` Mac verdict (**which came back negative**: that kernel failed its gate
on divergence and delivered 1.17Γ— rather than the predicted 6.5Γ—) and then an `aff4` equality
proof. What this bundle uses instead is the **shipped** batched path, the one behind the 8B
`mf64` bundle that has a Mac measurement.
**The falsifiable prediction, for whoever runs it first.** The batched kernel's anchor is the
8B-A1B `mf64` bundle's measured **6.3 ms/prompt-token**, and this bundle should land in that
mechanism's class rather than the ~21.5 ms/token of a decode-only bundle walking the prompt one
token at a time. Do not read 6.3 ms as the expected number: the 26B activates ~4B parameters per
token against the 8B-A1B's ~1B, and its dense MLP runs in parallel with the expert branch, so
several times 6.3 ms would still confirm the mechanism. **What would falsify it** is prefill
landing at or above the per-token cost of decode β€” that would mean the sorted grouped-GEMM path
is not engaging and the chunk is being walked serially.
### Decode-only, and a prefill variant is not deliverable by a flag
`gpu-pipelined/` has no `prefill` function, for two independent kernel reasons:
1. `--metal-sdpa`'s flash-decode kernel is structurally q=1 with no causal mask. The full
attention layers need that kernel for the same scratch-heap reason the dense 12B/31B do.
2. `MetalSwitchGLU.forward` is itself decode-only β€” it asserts a token batch of 1 and expands the
single activation row across the routed slots.
`hybrid-mf64/` is the bundle that solves both at once.
## Usage
Swift Package Manager, via [CoreAIKit](https://github.com/john-rocky/coreai-kit) β€” a community
package, not affiliated with Apple, requiring macOS 27 beta:
```swift
.package(url: "https://github.com/john-rocky/coreai-kit", branch: "main")
// target dependency: .product(name: "CoreAIKit", package: "coreai-kit")
```
```swift
import CoreAIKit
let model = ModelID(
"visible-cx/gemma4-26b-a4b-CoreAI",
path: "hybrid-mf64/gemma4_26b_a4b_qat_decode_int4linsym_moeaff4_msdpa_g8_pf64")
let chat = try await ChatSession(model: model) // ~17.6 GB download on first use
```
These are plain 2-input graphs β€” **no PLE table sidecar, no static input buffers** β€” so either
engine will load them, and the sequential engine is available for grammar-constrained decoding
(`configuration.engineVariant = .sequential`). Budget disk for the compiled graph on top of the
17.6 GB bundle, and expect a long cold compile.
## Integrity
Core AI `.aimodel` bundles are **not byte-reproducible**: the exporter is not deterministic even
against itself. Verify by digesting the exact published bytes rather than by rebuilding. Every
bundle carries `main.hash`, the raw 32 bytes of `sha256(main.mlirb)`; on the Hub the same value
is recoverable from the LFS oid without fetching the 17.6 GB file.
## Status
| Artifact | Status |
|---|---|
| `gpu-pipelined/…_moeaff4_msdpa_g8` | **UNQUALIFIED** β€” never executed on target hardware. Established: the checkpoint is the QAT source, the graph exports, the producer fingerprint passes, and the MoE block wiring matches the reference implementation to 3.5e-7 relative error. |
| `hybrid-mf64/…_msdpa_g8_pf64` | **EXPERIMENTAL β€” not qualified.** Adds an S=64 `prefill` entrypoint over the same weights, with box-side gates on the block-GQA mapping and the trace accounting. No Mac-side oracle, no device benchmark. Do not route production traffic to it. |
A qualifying run would need, in priority order:
1. **A decode oracle** β€” generate and compare against an fp32 HF reference. This is the gate that
catches an int4 or kernel-level error which the authoring gate cannot see.
2. **Isolate the two Metal kernels.** `gather_qmm` (the `aff4` variant) and the flash-decode SDPA
have never been in the same graph, on any model. If decode produces garbage, re-export without
`--metal-sdpa` first β€” it is an optimisation, the MoE kernel is not.
3. **Chunk parity** on `hybrid-mf64/`. Two independent reasons to check rather than assume: the
entrypoints use different attention implementations, and the expert path does a sort/unsort
round trip the q=1 path does not.
4. **Measure residency**, both dirty footprint and wired growth. No MoE Gemma bundle has been
measured this way, and the ~17.6 GB figure above is arithmetic.
5. **Compare tok/s against the dense 12B and 31B** on the same machine. That is the whole case
for this model.
6. **Router sanity at depth.** 128 experts / top-8 with a learned `per_expert_scale` is far
sparser routing than the LFM2.5-8B-A1B's 32/top-4. Check expert-utilisation spread on a real
prompt set before trusting long-form output.
## License
Google publishes the upstream QAT checkpoint under **Apache-2.0** with a `license_link` to the
[Gemma 4 license](https://ai.google.dev/gemma/docs/gemma_4_license), and this repo mirrors that
declaration. Use is governed by those terms and by the
[Gemma Prohibited Use Policy](https://ai.google.dev/gemma/prohibited_use_policy); the
obligations travel with any redistribution of these bundles. The contribution here is the port
and the recipe, not the weights.