File size: 21,286 Bytes
de7f65f
0d7b6cb
 
60bc67d
0d7b6cb
de7f65f
60bc67d
 
 
 
 
 
0d7b6cb
60bc67d
 
 
 
de7f65f
 
60bc67d
de7f65f
0d7b6cb
 
 
 
 
de7f65f
0d7b6cb
 
81c2abf
0d7b6cb
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
60bc67d
 
 
0d7b6cb
60bc67d
0d7b6cb
 
 
 
60bc67d
 
 
 
81c2abf
 
60bc67d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0d7b6cb
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
60bc67d
0d7b6cb
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
60bc67d
 
 
 
 
0d7b6cb
 
 
 
81c2abf
 
0d7b6cb
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5005b4c
 
 
 
 
 
 
 
 
 
 
0d7b6cb
 
 
 
60bc67d
0d7b6cb
 
 
 
60bc67d
0d7b6cb
 
 
60bc67d
0d7b6cb
 
60bc67d
0d7b6cb
81c2abf
0d7b6cb
 
 
 
 
60bc67d
0d7b6cb
81c2abf
0d7b6cb
 
 
60bc67d
0d7b6cb
 
 
 
 
25b51a7
0d7b6cb
 
 
 
25b51a7
0d7b6cb
 
 
25b51a7
0d7b6cb
25b51a7
0d7b6cb
 
50c2ae0
0d7b6cb
 
 
50c2ae0
 
0d7b6cb
 
98d2913
0d7b6cb
 
98d2913
0d7b6cb
 
 
 
 
 
 
beecbbf
0d7b6cb
 
 
 
beecbbf
0d7b6cb
beecbbf
0d7b6cb
 
 
 
 
 
 
beecbbf
0d7b6cb
cb38966
0d7b6cb
 
 
 
 
 
cb38966
0d7b6cb
cb38966
0d7b6cb
 
 
 
 
 
 
cb38966
0d7b6cb
 
 
cb38966
0d7b6cb
cb38966
0d7b6cb
 
 
 
 
cb38966
0d7b6cb
cb38966
 
 
 
0d7b6cb
cb38966
 
 
0d7b6cb
 
 
 
cb38966
0d7b6cb
 
 
 
 
cb38966
0d7b6cb
cb38966
0d7b6cb
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
---
license: apache-2.0
license_link: https://ai.google.dev/gemma/docs/gemma_4_license
base_model: google/gemma-4-E2B-it-qat-q4_0-unquantized
base_model_relation: quantized
library_name: coreai
pipeline_tag: text-generation
tags:
  - core-ai
  - aimodel
  - apple-silicon
  - on-device
  - coreai-kit
  - quantized
  - int4
  - qat
  - gemma4
---

# Gemma 4 E2B β€” Core AI (.aimodel)

`google/gemma-4-E2B-it-qat-q4_0-unquantized` converted to Core AI `.aimodel` bundles for Apple
silicon by [visible-cx](https://huggingface.co/visible-cx). These are derivative artifacts:
Google's QAT-trained weights rounded onto the int4 grid they were trained for and re-expressed
as a Core AI graph. They load through Core AI on macOS and are not usable by PyTorch, GGUF or
MLX.

Gemma 4 E2B uses **Per-Layer Embeddings**, so these bundles take a large embedding gather table
as a *static graph input* rather than carrying it in the graph. That table ships in
`ios-frontend/` and **the bundles do not load without it**; a missing table produces a bare
input-arity error naming `ple_table`/`ple_scale`. It is 2.81 GB and it is part of the model's
memory cost, not a sidecar you can ignore.

> ⚠️ **Known issue β€” memory growth per generated token.** The Core AI runtime caches one
> compiled specialization per input-shape signature, and this export makes sequence length
> load-bearing: every generated token adds one token to `position_ids` and mints a new
> signature, retaining roughly **81 MB of GPU allocations per generated token** until the
> process exits.
>
> **Guided decoding does not protect you β€” shape reuse does.** At 64 generated tokens the same
> prompt costs **9.84 GB free-form and 9.88 GB guided**; the grammar loop is not a defence. What
> *is* a defence is repeating lengths you have already run, which replay from cache at no cost:
> a fixed-schema, fixed-cap extraction workload stays bounded, and that is the shape measured
> at 10/10 below. Long or variable-length generation grows without bound; no in-process
> mitigation exists, and it is not fixable below a re-export. `stable/` and `ring/` in this repo
> are those re-exports β€” see [Shape-stable re-exports](#shape-stable-re-exports).

## Contents

### Production bundles

| Path | Bytes | Context | Functions | Status |
|---|---:|---|---|---|
| `gpu-pipelined/gemma4_e2b_qat_decode_int4lin_tbl_pf64` | 2,122,089,973 | 4096 | main + prefill | **QUALIFIED (guided / bounded shapes)** |
| `w4a8/gemma4_e2b_qat_decode_int4lin_a8_tbl_pf64` | 2,122,679,604 | 16384 | main + prefill | EXPERIMENTAL |

Each bundle folder holds `<name>.aimodel/` (`main.mlirb` β‰ˆ 2.09 GB, `main.hash`, asset
`metadata.json`), a bundle-level `metadata.json`, and `tokenizer/` (`tokenizer.json`
32,169,626 B, `tokenizer_config.json`, `generation_config.json`, `chat_template.jinja`
18,569 B). The `w4a8` folder additionally ships its `calibration_corpus.jsonl` (35,045 B).

### The PLE gather-table sidecar β€” required, not optional

| Path | Files | Bytes |
|---|---|---:|
| `ios-frontend/gemma4_qat_gather_raw/` | 7 | 2,808,612,071 |

| File | Bytes |
|---|---:|
| `embed_per_layer.i8` | 2,348,810,240 |
| `embed_tokens.i8` | 402,653,184 |
| `proj.f32` | 55,050,240 |
| `embed_per_layer.scale.f32` | 1,048,576 |
| `embed_tokens.scale.f32` | 1,048,576 |
| `proj_norm.f32` | 1,024 |
| `meta.json` | 231 |

`meta.json` records the shape and the dequant convention:
`V 262144, D 1536, PLD 8960, L 35, ld 256, embed_scale_pl 16.0` (= √256, which is what the
exporter assumes). Every `_tbl` bundle binds `ple_table` (from `embed_per_layer.i8`) and
`ple_scale` (from `embed_per_layer.scale.f32`) as **static** graph inputs. **A QAT bundle must
be paired with the QAT tables.**

### Diagnostic and proving assets

| Path | Bundle | Bytes | Context |
|---|---|---:|---|
| `stable/gpu-pipelined/` | `…_tbl_pf64_stable_c16384` | 2,122,071,656 | 16384 |
| `ring/gpu-pipelined/` | `…_tbl_pf64_ring_c16384` | 2,122,101,099 | 16384 |
| `ring2/gpu-pipelined/` | `…_tbl_pf64_ring_c2048` | 2,122,043,748 | 2048 |
| `ring2/gpu-pipelined/` | `…_tbl_pf64_ring_c8192` | 2,122,068,315 | 8192 |
| `ring3/gpu-pipelined/` | `…_tbl_pf64_ring_c2048_gqafull` | 2,122,042,254 | 2048 |
| `ring3/gpu-pipelined/` | `…_tbl_pf64_ring_c4096_gqafull` | 2,122,050,478 | 4096 |
| `stable-smoke/` `ring-smoke/` `ring2-smoke/` `ring3-smoke/` | 5-layer truncations | ~1.17 GB each | β€” |

The `*-smoke` folders are **not models**: 5-layer truncations of the decoder, published only so
host-side work can be developed against a small download. They carry real weights for the
layers they keep and produce low-quality text by design.

**Stop token:** every bundle declares `eos_token = "<turn|>"` (id 106), the turn terminator
Gemma 4 emits. `generation_config.json` independently lists `eos_token_id: [1, 106, 50]`. A
host that stops on the raw upstream `<eos>` will overrun every reply.

## Provenance

| | |
|---|---|
| Base checkpoint | `google/gemma-4-E2B-it-qat-q4_0-unquantized` (ungated) |
| Zoo recipe | `gemma-4-e2b`, `status = "verified"` β€” `int4lin --tbl` |
| Recipe (pf64) | `export_gemma4_pf_pipelined.py --pf 64` with `--tbl` and `--raw-dir` pointed at the gather table above |
| Toolchain base | `apple/coreai-models` @ `b1cb71b8522d99408059fa0b98b8742171bcb0b8` + the [coreai-model-zoo](https://github.com/john-rocky/coreai-model-zoo) python overlay |
| Toolchain | `coreai-torch 0.4.1`, `coreai-core 1.0.0b2`, `coreai-opt 0.2.1`, `torch 2.9.0` |
| Producer fingerprint | `coreai-core 1.0.0b2` on every inner asset `metadata.json` |
| Weight format | **int4, per-block-32** (`int4lin`, symmetric-with-clipping) β€” the ggml q4_0 grid the QAT checkpoint was trained on |
| Vocab | 262,144 |
| Export functions | `main` (S=1 decode) and, in `_pf64` bundles, `prefill` (S=64 chunked prefill) |

"QAT-unquantized" means QAT-*trained*, stored at full width; the int4 rounding happens at
export, onto the grid training already targeted.

`_tbl` = the PLE gather table is bound as a static graph input. `_pf64` = a second entrypoint,
`prefill`, with a static query width of 64 (`function_map: {"main": ["main", "prefill"]}`).

The `gpu-pipelined/` bundle mirrors the zoo's verified `gemma-4-e2b` recipe with the `pf64`
multifunction addition; the gather table is the zoo's own
[mlboydaisuke](https://huggingface.co/mlboydaisuke) artifact layout, and the zoo's
`gemma-4-E2B-CoreAI` repo is the upstream this one is a sibling of.

## Requirements

- **Apple silicon Mac**, Core AI runtime.
- **Engine contract: 4 inputs** β€” `input_ids`, `position_ids`, plus static `ple_table` and
  `ple_scale`. Two engines accept that, and the difference matters:
  - **Pipelined engine** β€” binds the statics zero-copy over the caller's buffer, but does not
    expose logits, so no grammar-constrained decoding. This is the default path in CoreAIKit's
    `GemmaRuntime` and the path the throughput numbers below were measured on.
  - **Sequential engine** β€” the only logits-capable engine, and therefore the only path for
    guided decoding. It accepts `>= 2` inputs and binds everything beyond
    `input_ids`/`position_ids` from `EngineOptions.staticInputBuffers`. **A caller that does not
    supply those buffers gets a load failure by name**, not a fallback:
    `invalidInputType("Inputs beyond input_ids/position_ids must be bound as constant buffers
    on this engine, but ["ple_table", "ple_scale"] have none.")`. It pays a **one-time copy of
    every constant input at load β€” 2.19 GB for E2B** β€” because its submission path otherwise
    materialises a foreign buffer-backed view on every forward pass; binding the table as a raw
    view instead costs ~0.30 s per pass (3.4 tok/s).
- **States:** `keyCache` / `valueCache` `Float16, 15 Γ— 1 Γ— 1 Γ— ? Γ— 512`. Dynamic sequence dim β†’
  `GrowingKVCache` (initial 256, doubling), not a static allocation at the manifest maximum.
- **KV cost: 30,720 bytes per token** (fp16) β€” 126 MB at 4096, 503 MB at 16384. **KV is not the
  ceiling at this tier; the PLE table is.**
- **Minimum practical machine memory: 16 GB**, for bounded-shape guided work only.
- Sliding-window note: E2B interleaves sliding and full attention layers across 35 layers
  collapsed to 15 KV slots. The export models the window in the *mask*, not in memory β€”
  sliding layers occupy full-length slots and zero-pad head_dim 256β†’512, so 80% of KV bytes sit
  in windowed slots and 40% is dead padding.
- The bundle manifest declares `runtime_env COREAI_CHUNK_THRESHOLD=1`.

## Measurements

Measured on a **16 GB Apple silicon Mac (M2 Pro, macOS 27 beta)**.

### Throughput and memory β€” n=2, watchdogged, both runs completed

The published `gpu-pipelined/` bundle at its shipped c4096, with the external PLE tables, on the
pipelined engine (both mandatory), 73-token prompt, 8 calls of 83 generated tokens, watchdog
sampling at 0.5 s. **Both runs completed 8/8 and neither tripped.**

| | run 1 | run 2 |
|---|---:|---:|
| reclaimable at start | 4.88 | 6.20 |
| wired baseline β†’ peak | 3.15 β†’ 8.15 | 2.97 β†’ 8.01 |
| **wired growth** | **5.00** | **5.04** |
| process footprint peak | 4.64 | 4.58 |
| compressor | 2.34 β†’ 5.48 | 2.56 β†’ 5.33 |
| reclaimable trough | **1.35** | 1.78 |
| **decode** | **41.35 tok/s** | **41.40 tok/s** |

GiB unless stated. n=2 agreeing to **0.8%**.

**Charge this model for two artifacts, not one.** The E-series loads a compiled program *and* a
PLE table, and a memory law told only about the first is wrong by 2.87 GiB:

```
compiled blob   resources.bin              2,088,055,648 B  =  1.9447 GiB
PLE tables      gemma4_qat_gather_raw      2,808,611,840 B  =  2.6157 GiB
------------------------------------------------------------------------
what the runtime actually loads                             =  4.5603 GiB
```

(The PLE figure is the binary payload the runtime binds; the published folder is 231 bytes
larger because it also carries `meta.json`.)

| denominator | wired Γ· denominator |
|---|---:|
| compiled blob alone | **2.581** β€” absurd |
| **blob + PLE tables** | **1.101** |

**1.101 sits alongside the LFM MoE's 1.092 and below the dense 12B's 1.157** β€” the E-series is
not architecture-exceptional at all; it was being charged for one of the two files it opens. The
honest requirement on this machine is `4.5603 Γ— 1.10 + 1.25 GiB in-flight floor` = **6.27 GiB**.

A related catalog error, recorded because it points the other way: `approximateBytes` for E2B
had been taken from the published **LiteRT `.litertlm`** artifact (β‰ˆ 2.41 GB), a file the Core AI
backend never opens. One number over-charged the download size and the other under-charged the
memory gate, from the same root cause β€” pricing the wrong artifact.

### Guided structured output

10-sample harness, guided JSON-constrained decoding against a fixed schema, greedy, **sequential
engine** with the PLE tables bound as static inputs, `reset()` between samples, 128-token cap.

| | `gpu-pipelined/…_tbl_pf64` |
|---|---|
| Load | 12.1 s |
| Guided JSON parse | **10/10** |
| Enum-clean | **10/10** |
| s/row (long samples) | **6.12** |
| s/row (short samples) | **4.23** |
| Decode | **22.7–32.7 tok/s** |
| TTFT | **0.59–4.20 s** |
| Peak footprint | 8.26 GB |
| Outcome | completed all ten samples |

The 41.4 tok/s figure above and the 22.7–32.7 here are not in conflict: the first is the
pipelined engine unguided, the second is the sequential engine under a grammar mask, whose
step-synchronous prefill costs roughly 25% per sample. An earlier unguided pipelined run on the
same machine reached 44.3 tok/s; upstream measured E2B at 77.0/87.1 tok/s on an M4 Max.

**The S=64 `prefill` function carries the time-to-first-token.** Head to head on a 942-token
prompt against a decode-only export of the same weights, exporting `prefill` moved TTFT
35.74 s β†’ 4.65 s and s/row 45.04 β†’ 8.83, a **7.7Γ—**. Decode is untouched. The decode-only
bundle is no longer published.

**Enum conformance is the grammar's doing.** Unguided, the model emits an off-schema enum value
on essentially every sample. Guided, all ten are correct, because an off-enum token is
unsamplable.

**Needle-in-haystack recall: 3/3 verbatim at 8k**, within a 64-token cap. 15k was not attempted:
8k already cost 36.71 GB of footprint under the shape-signature defect.

### What this model is and is not, on this stack

The shape-signature growth sets a hard ceiling that no cap works around: at ~81 MB/token over a
~9.6 GB intercept, a 900-token report costs ~54 GB (measured killed) and a 600-token insight
card ~58 GB. The 128-token ceiling that does survive is shorter than a single card. **Gemma 4 is
an enrichment model on this stack β€” bounded, repeating shapes β€” and it is not a report, RAG or
long-insight model. The blocker is the export, not the weights.**

### The `staticInputBuffers` residency question, settled

The PLE table is bound through `EngineOptions.staticInputBuffers`, and it had been assumed those
pages stay clean and evictable. **They do not.** Forcing the mapped path
(`COREAI_GEMMA_TABLES=mapped`) against the owned one, on this bundle:

| | `owned` (`makeBuffer` + read) | `mapped` (`mmap` + `bytesNoCopy`) |
|---|---:|---:|
| `phys_footprint` after load | **2.39 GiB** | **0.19 GiB** |
| `vmmap` region holding the table | `IOAccelerator`, 2.2 G resident, dirty 64K | `mapped file`, dirty 0K |
| **wired growth** | **4.67 GiB** | **4.65 GiB** |

The footprint column is real and matters on iOS jetsam accounting. **The wired column is what
the machine is about, and the two paths are identical to within 0.02 GiB.** Under pressure β€”
reclaimable falling to 2.52 GiB with the compressor climbing β€” not one byte came back. An
`mmap`-backed static input gets wired at the first forward pass exactly like a program constant.

One qualification, added later: that test measured whether the *machine* can take the pages
back, and it cannot. Releasing the `InferenceFunction` does give them back β€” the process can,
even though the kernel cannot.

## Usage

Swift Package Manager, via [CoreAIKit](https://github.com/john-rocky/coreai-kit) β€” a community
package, not affiliated with Apple, requiring macOS 27 beta:

```swift
.package(url: "https://github.com/john-rocky/coreai-kit", branch: "main")
// target dependency: .product(name: "CoreAIKit", package: "coreai-kit")
```

An E-series bundle is **two downloads**, the decoder and its paired PLE tables, addressed as two
paths inside this repo:

```swift
import CoreAIKit

let store = ModelStore.default
let decoderURL = try await store.download(
    ModelID("visible-cx/gemma-4-E2B-CoreAI",
            path: "gpu-pipelined/gemma4_e2b_qat_decode_int4lin_tbl_pf64"))
let tablesURL = try await store.download(
    ModelID("visible-cx/gemma-4-E2B-CoreAI",
            path: "ios-frontend/gemma4_qat_gather_raw"))

// engineVariant defaults to .pipelined here, which is the supported path for _tbl bundles.
let runtime = try await GemmaRuntime(
    decoderBundleAt: decoderURL, tablesAt: tablesURL, arch: .gemma4)
```

Notes that are not optional:

- **Do not pair a QAT bundle with non-QAT tables.**
- **Do not enable chunked prefill.** The `…_tbl` graph is S=1 on `main`; `GemmaRuntime` sets
  `COREAI_CHUNK_THRESHOLD=1` for you if it is unset. Leave it.
- **Guided decoding needs the sequential engine**, and the sequential engine needs
  `ple_table`/`ple_scale` supplied through `EngineOptions.staticInputBuffers` β€” it will refuse
  by name otherwise.
- Pass `revision:` a Hub commit hash to pin immutable bytes.

## Integrity

Core AI `.aimodel` bundles are **not byte-reproducible**: the exporter is not deterministic even
against itself. Verify by digesting the exact published bytes rather than by rebuilding. Every
bundle carries `main.hash`, the raw 32 bytes of `sha256(main.mlirb)`; on the Hub the same value
is recoverable from the LFS oid without fetching the file. The compiled-blob identity this
project keys its measurements on for the published bundle is
`73bef8155c41d512e9c6b4ab1788b7547bade134`.

## Status

| Artifact | Status |
|---|---|
| `gpu-pipelined/…_tbl_pf64` (ctx 4096) | **QUALIFIED FOR BOUNDED-SHAPE GENERATION** β€” measured: 10/10 guided parse and enum-clean, 6.12/4.23 s/row, 41.4 tok/s unguided on n=2 completed watchdogged runs, 5.00–5.04 GiB wired growth, honest requirement 6.27 GiB. **Not qualified for free-form or variable-length generation** β€” see the known issue. |
| `w4a8/…_a8_tbl_pf64` | **EXPERIMENTAL** β€” built, unmeasured. Same int4 weights and graph; the difference is an int8 quantize/dequantize pair on the inputs of every `F.linear`, calibrated on 128 synthetic samples (corpus ships in the folder). Built as a prefill/TTFT lever, with the toolchain's own caveat that the runtime fast path is `W_INT8 Γ— A_INT8` and these weights use int4 scale-multiply dequant β€” so the expected gain may be zero or negative. It exists to be measured, not because a win is predicted. |
| `stable/…_stable_c16384` | **EXPERIMENTAL** β€” the memory fix is confirmed (6.53 GB flat peak across a 664-token generation, +0.011 MB/token, guided output byte-identical to the default bundle) but decode runs at **~1.0 tok/s** because every step reads the full 16,384-slot cache. A working proof of the contract, not a deployable bundle. |
| `ring/`, `ring2/`, `ring3/` | **EXPERIMENTAL β€” gated in torch, no on-device numbers.** Diagnostic assets for the capacity-vs-cost question. |
| `*-smoke/` | **NOT MODELS** β€” 5-layer truncations for host development. |

**No numerics gate has been run on device for any bundle in this repo.** The 10/10 results are
behavioural (parse rate, enum conformance, clean stop); a decode oracle against an fp32
reference has not been run.

## Shape-stable re-exports

`stable/`, `ring/`, `ring2/` and `ring3/` are re-exports of the same weights and the same
quantization onto a contract in which **no input shape moves between steps**, which removes the
per-generated-token growth by design. They need a host that feeds `position_ids` as the
**absolute positions of the S new tokens only**; a host that feeds the growing `0..N` prefix
will write the KV cache at the wrong offset.

**Contract** (identical across all four families):

```
main    IN  input_ids Int32 1x1  | position_ids Int32 1x1  | ple_table Int8 V x (L*ld) | ple_scale Float32 V
prefill IN  input_ids Int32 1x64 | position_ids Int32 1x64 | (same statics)
        ST  keyCache / valueCache Float16, literal extents
        OUT logits Float16 1 x S x 262144
```

`position_ids[0,0]` is also the cache slot the K/V for those tokens is written at. The mask β€”
causality, unwritten slots and the sliding window β€” is derived inside the graph from those
positions, so there is no mask input to bind. The context ceiling is **encoded in the graph**: a
different window needs a different export, not a manifest edit.

**`stable/` gave every layer a full 16,384-slot cache and read all of it every step**, which is
the 1 tok/s. `ring/` keeps the contract byte for byte and changes what sits behind it: sliding
layers read a **576-slot ring** (the model's 512-token window plus one 64-token prefill chunk)
instead of 16,384, and the grouped-query head expansion is folded into the query rather than
materialised.

| | `stable/` | `ring/` |
|---|---|---|
| KV state shape | `[slots, 1, n_kv, 16384, 512]` Γ—2 | `[1, 1, 1, 56064, 512]` Γ—2 |
| KV bytes, both states | 503.3 MB | **114.8 MB** |
| cache slots read per decode step | 245,760 | **56,064** |
| dynamic dimensions | none | none |

`ring2/` is the same recipe at 2,048 and 8,192 (26.7 MB and 64.5 MB of KV); `ring3/` folds the
GQA expansion **only in the full-attention layers**, where the key length is long enough for the
trade to pay, and keeps the stock gathered form in the short-key sliding layers.

**One host precondition new in `ring/`:** a `prefill` call's first position must be a multiple
of 64. The ring write is a fixed-width store at `p0 % 576`, and 576 is nine 64-token chunks, so
a 64-aligned chunk can never straddle the wrap. `stable/` tolerated an unaligned chunk; `ring/`
does not.

**Equivalence gates**, all run in eager torch against the shipped graph before conversion:
`stable/` β€” greedy argmax identical on every tested decode step including past the sliding
window, worst logits cosine 0.99998 (fp16), prefill bit-identical in fp32. `ring*/` β€” 1,600
prompt tokens as 25 chunks of 64 then 32 greedy decode steps, final position 1,631 so the ring
wraps about three times, fp32, 57 comparisons per run: **0 argmax mismatches** on every bundle,
worst cosine 0.999999999987. No dynamic dimension appears in either entrypoint of any of them.

## License

Google publishes the upstream QAT checkpoint under **Apache-2.0** with a `license_link` to the
[Gemma 4 license](https://ai.google.dev/gemma/docs/gemma_4_license), and this repo mirrors that
declaration. Use is governed by those terms and by the
[Gemma Prohibited Use Policy](https://ai.google.dev/gemma/prohibited_use_policy); the
obligations travel with any redistribution of these bundles, **including the gather-table
sidecar**, which is derived from the same weights. The contribution here is the conversion, not
the weights.