File size: 23,823 Bytes
d85732c
0505916
 
451139d
0505916
d85732c
451139d
 
 
 
 
 
0505916
451139d
 
 
 
d85732c
 
451139d
d85732c
0505916
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
451139d
 
 
0505916
451139d
0505916
 
 
 
451139d
 
 
 
f9efe59
 
451139d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0505916
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
d85732c
451139d
18a972a
0505916
 
 
 
 
 
 
f9efe59
0505916
 
 
 
f9efe59
0505916
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
8947bf3
0505916
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
451139d
f9efe59
451139d
 
f9efe59
 
 
 
 
 
0505916
f9efe59
0505916
 
 
8947bf3
451139d
0505916
 
f9efe59
0505916
 
 
f9efe59
0505916
36db1ae
f9efe59
36db1ae
0505916
 
36db1ae
f9efe59
 
0505916
 
36db1ae
f9efe59
 
0505916
 
 
1998b7e
0505916
451139d
0505916
 
 
 
 
451139d
0505916
668d018
0505916
 
668d018
0505916
 
 
 
668d018
0505916
 
668d018
0505916
 
668d018
0505916
 
 
 
 
 
 
668d018
0505916
 
 
668d018
 
0505916
cbf6fbc
0505916
 
 
 
 
 
 
 
 
cbf6fbc
0505916
de11ce9
0505916
 
 
 
de11ce9
0505916
de11ce9
0505916
 
 
 
 
 
de11ce9
0505916
 
 
de11ce9
0505916
de11ce9
0505916
 
 
 
de11ce9
0505916
de11ce9
 
 
 
0505916
de11ce9
 
 
0505916
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
99dc7f4
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
---
license: apache-2.0
license_link: https://ai.google.dev/gemma/docs/gemma_4_license
base_model: google/gemma-4-E4B-it-qat-q4_0-unquantized
base_model_relation: quantized
library_name: coreai
pipeline_tag: text-generation
tags:
  - core-ai
  - aimodel
  - apple-silicon
  - on-device
  - coreai-kit
  - quantized
  - int4
  - qat
  - gemma4
---

# Gemma 4 E4B β€” Core AI (.aimodel)

`google/gemma-4-E4B-it-qat-q4_0-unquantized` converted to Core AI `.aimodel` bundles for Apple
silicon by [visible-cx](https://huggingface.co/visible-cx). These are derivative artifacts:
Google's QAT-trained weights rounded onto the int4 grid they were trained for and re-expressed
as a Core AI graph. They load through Core AI on macOS and are not usable by PyTorch, GGUF or
MLX.

Gemma 4 E4B uses **Per-Layer Embeddings**, so the working bundles take a large embedding gather
table as a *static graph input* rather than carrying it in the graph. That table ships in
`ios-frontend/` and **the bundles do not load without it**; a missing table produces a bare
input-arity error naming `ple_table`/`ple_scale`. It is 3.60 GB and it is part of the model's
memory cost, not a sidecar you can ignore.

> ⚠️ **This model does not fit a 16 GB Mac under an honest memory gate.** Charged for both
> artifacts it loads, at the coefficient measured on its E2B sibling, E4B asks **8.99 GiB** β€”
> and on the reference 16 GB machine the check refused it, short by 3.20 GiB. It has never been
> run under a watchdog. See [Measurements](#measurements).

> ⚠️ **Known issue β€” memory growth per generated token.** The Core AI runtime caches one
> compiled specialization per input-shape signature, and this export makes sequence length
> load-bearing: every generated token mints a new signature, retaining roughly **81 MB of GPU
> allocations per generated token** until the process exits. **Guided decoding does not protect
> you β€” shape reuse does.** Repeating a length you have already run replays from cache at no
> cost, so a fixed-schema, fixed-cap extraction workload stays bounded; long or variable-length
> generation grows without bound, and it is not fixable below a re-export. `stable/`, `ring/`
> and `ring3/` in this repo are those re-exports.

## Contents

### Production bundles

| Path | Bytes | Context | Functions | Status |
|---|---:|---|---|---|
| `gpu-pipelined/gemma4_e4b_qat_decode_int4lin_tbl_pf64` | 3,989,986,858 | 4096 | main + prefill | **QUALIFIED (guided / bounded shapes)** |
| `w4a8/gemma4_e4b_qat_decode_int4lin_a8_tbl_pf64` | 3,990,690,638 | 16384 | main + prefill | EXPERIMENTAL |

Each folder holds `<name>.aimodel/` (`main.mlirb` β‰ˆ 3.96 GB, `main.hash`, asset
`metadata.json`), a bundle-level `metadata.json`, and `tokenizer/` (`tokenizer.json`
32,169,626 B, `tokenizer_config.json`, `generation_config.json`, `chat_template.jinja`
18,569 B). The `w4a8` folder additionally ships its `calibration_corpus.jsonl` (35,045 B).

### The PLE gather-table sidecar β€” required, not optional

| Path | Files | Bytes |
|---|---|---:|
| `ios-frontend/gemma4_e4b_qat_gather_raw/` | 7 | 3,601,859,815 |

| File | Bytes |
|---|---:|
| `embed_per_layer.i8` | 2,818,572,288 |
| `embed_tokens.i8` | 671,088,640 |
| `proj.f32` | 110,100,480 |
| `embed_per_layer.scale.f32` | 1,048,576 |
| `embed_tokens.scale.f32` | 1,048,576 |
| `proj_norm.f32` | 1,024 |
| `meta.json` | 231 |

Every `_tbl` bundle binds `ple_table` (from `embed_per_layer.i8`) and `ple_scale` (from
`embed_per_layer.scale.f32`) as **static** graph inputs. **A QAT bundle must be paired with the
QAT tables.**

### Shape-stable re-exports

| Path | Bundle | Bytes | Context |
|---|---|---:|---|
| `stable/gpu-pipelined/` | `…_tbl_pf64_stable_c16384` | 3,989,939,812 | 16384 |
| `ring/gpu-pipelined/` | `…_tbl_pf64_ring_c16384` | 3,989,979,750 | 16384 |
| `ring3/gpu-pipelined/` | `…_tbl_pf64_ring_c4096_gqafull` | 3,989,929,508 | 4096 |

**Stop token:** every bundle declares `eos_token = "<turn|>"` (id 106), the turn terminator
Gemma 4 emits. `generation_config.json` independently lists `eos_token_id: [1, 106, 50]`. A
host that stops on the raw upstream `<eos>` will overrun every reply.

## Provenance

| | |
|---|---|
| Base checkpoint | `google/gemma-4-E4B-it-qat-q4_0-unquantized` (ungated) |
| Zoo recipe | `gemma-4-e4b`, `status = "verified"` β€” `int4lin` |
| Recipe (pf64) | `export_gemma4_pf_pipelined.py --pf 64` with `--tbl` and `--raw-dir` pointed at the gather table above |
| Toolchain base | `apple/coreai-models` @ `b1cb71b8522d99408059fa0b98b8742171bcb0b8` + the [coreai-model-zoo](https://github.com/john-rocky/coreai-model-zoo) python overlay |
| Toolchain | `coreai-torch 0.4.1`, `coreai-core 1.0.0b2`, `coreai-opt 0.2.1`, `torch 2.9.0` |
| Producer fingerprint | `coreai-core 1.0.0b2` on every inner asset `metadata.json` |
| Weight format | **int4, per-block-32** (`int4lin`, symmetric-with-clipping) β€” the ggml q4_0 grid the QAT checkpoint was trained on |
| Vocab | 262,144 |
| Export functions | `main` (S=1 decode) and, in `_pf64` bundles, `prefill` (S=64 chunked prefill) |

"QAT-unquantized" means QAT-*trained*, stored full width; the int4 rounding happens at export
onto the grid training already targeted.

`_tbl` = the PLE gather table is bound as a static graph input. `_pf64` = a second entrypoint,
`prefill`, with a static query width of 64 (`function_map: {"main": ["main", "prefill"]}`).

The `gpu-pipelined/` bundle mirrors the zoo's verified `gemma-4-e4b` recipe with the `pf64`
multifunction addition; the gather table follows the
[coreai-model-zoo](https://github.com/john-rocky/coreai-model-zoo) /
[mlboydaisuke](https://huggingface.co/mlboydaisuke) artifact layout.

## Requirements

- **Apple silicon Mac**, Core AI runtime. Practically, **24 GB or more** β€” see the arithmetic
  below.
- **Engine contract: 4 inputs** β€” `input_ids`, `position_ids`, plus static `ple_table` and
  `ple_scale`:
  - **Pipelined engine** β€” binds the statics zero-copy over the caller's buffer, but does not
    expose logits, so no grammar-constrained decoding. This is CoreAIKit's default for `_tbl`
    bundles.
  - **Sequential engine** β€” the only logits-capable engine, and therefore the only path for
    guided decoding. It binds everything beyond `input_ids`/`position_ids` from
    `EngineOptions.staticInputBuffers`, and **a caller that does not supply those buffers gets a
    load failure naming them**, not a fallback. It pays a **one-time copy of every constant
    input at load β€” 2.69 GB for E4B**; binding the table as a raw view instead drives peak
    footprint to 14.9 GB and faults the runtime mid-prefill.
- **States:** `keyCache` / `valueCache` `Float16, 24 Γ— 1 Γ— 2 Γ— ? Γ— 512`. Dynamic sequence dim β†’
  `GrowingKVCache` (initial 256, doubling).
- **KV cost: 98,304 bytes per token** (fp16) β€” 403 MB at 4096, 1.61 GB at 16384. **KV is not the
  ceiling at this tier; the PLE table is.**
- **Run one E4B session per process.** Wired memory is not reclaimed until the process exits
  unless the host releases the inference function; a second full run in the same process drove
  wired memory to 13.2 GB on a 16 GB machine.
- Sliding-window note: E4B interleaves sliding and full attention layers across 42 layers
  collapsed to 24 KV slots. The export models the window in the *mask*, not in memory β€” sliding
  layers ride full-length slots and zero-pad head_dim 256β†’512, so 83.3% of KV bytes sit in
  windowed slots and 41.7% is dead padding.
- The bundle manifest declares `runtime_env COREAI_CHUNK_THRESHOLD=1`. Both engines derive the
  static query width from the graph.

## Measurements

Measured on a **16 GB Apple silicon Mac (M2 Pro, macOS 27 beta)** unless stated.

### The memory verdict β€” arithmetic, not a run

**E4B has never been run under a watchdog.** What exists is its artifact inventory charged at
the coefficient measured on its E2B sibling (n=2 completed runs, `wired Γ· (blob + PLE tables)` =
**1.101**):

```
compiled blob   3,955,446,640 B  =  3.6838 GiB
PLE tables      3,601,859,584 B  =  3.3545 GiB
------------------------------------------------
artifacts                        =  7.0383 GiB
x 1.10 (measured on E2B)         =  7.742 GiB
+ 1.25 GiB in-flight floor       =  8.99 GiB   required

reclaimable at the check         =  5.79 GiB   -> REFUSED, short by 3.20 GiB
```

(The PLE figure is the binary payload the runtime binds; the published folder is 231 bytes
larger because it also carries `meta.json`.)

This is a **3.67 GiB correction against the gate this project previously shipped**, which
charged `blob Γ— 1.106 + 1.25` = 5.324 GiB and would have admitted the model. The cause is the
same one-line error in both directions across the E-series: **the runtime loads two artifacts
and the law was only ever told about the first.** (E2B carried the same defect at 2.87 GiB.)

Stated honestly: **the refusal is arithmetic on E2B's coefficient, not a measurement of E4B.**

For scale, the compile-only readings that do exist for this bundle: bundle 3.716 GiB, compiled
blob **3.684 GiB**, graph constant **βˆ’0.002 GiB** (Gemma's compiled artifact is very slightly
*smaller* than its bundle, unlike the LFMs at 1.18–1.34Γ—).

### Guided structured output

10-sample harness, guided JSON-constrained decoding against a fixed schema, greedy, **sequential
engine** with the PLE tables bound as static inputs, `reset()` between samples, 128-token cap.

| | `gpu-pipelined/…_tbl_pf64` |
|---|---|
| Load | 16.7 s |
| Guided JSON parse | **10/10** |
| Enum-clean | **10/10** |
| s/row (long samples) | **14.70** |
| s/row (short samples) | **5.93** |
| Decode | **9.5–26.4 tok/s** (mean 17.8) |
| TTFT | **1.23–6.51 s** |
| Peak footprint | 9.99 GB |
| Max RSS | 9.63 GB |
| Stop | `<turn\|>`, clean self-stop on every sample |

E4B is memory-bound at the 16 GB tier and its per-sample cost is sensitive to what else is
resident β€” expect roughly 2Γ— these figures when the model has to share. Note the spread in the
decode column: that variance *is* the memory pressure.

**The grammar is close to free; prefill is the cost.** Decode throughput is essentially
unchanged from unguided pipelined measurements; guided samples cost ~25% more because the
sequential engine's prefill is step-synchronous, not because of the constraint.

**Enum conformance is entirely the grammar's doing.** Unguided, this model emits an off-schema
enum value in all ten samples. Guided, all ten are correct, because an off-enum token is
unsamplable.

Published Mac figures of ~55.8 tok/s for E4B decode should be read as larger-machine figures.

### Unguided workload β€” memory-capped

Free-form generation from a fixed prompt. **Every generation length below is a memory cap, not a
model stop** β€” see the known issue.

| depth | prompt tokens | TTFT | decode | generated | wall | peak footprint |
|---|---|---|---|---|---|---|
| 3.4k | 3,314 | 68.46 s | 0.78 tok/s | 64 (capped) | 151.6 s | **16.91 GB** |
| 8k | 7,907 | 159.1 s | 0.52 tok/s | 32 (capped) | 218.8 s | **31.33 GB** |

Output quality up to the cap is sound; the limit is memory, not capability.

**Needle-in-haystack recall at 8k: β‰₯2/3.** Three distinctive facts planted at 10/50/90% of the
filler; the 32-token memory cap truncated the answer mid-fact-2, so fact 3 was never reachable.
15k was not attempted.

### What this model is and is not, on this stack

The shape-signature growth sets a ceiling no cap works around: at ~81 MB/token over a ~9.6 GB
intercept, a 900-token report costs ~54 GB and a 600-token insight card ~58 GB. The caps that do
survive are shorter than a single card. **Gemma 4 is an enrichment model on this stack β€”
bounded, repeating shapes β€” and it is not a report, RAG or long-insight model. The blocker is
the export, not the weights.**

## Usage

Swift Package Manager, via [CoreAIKit](https://github.com/john-rocky/coreai-kit) β€” a community
package, not affiliated with Apple, requiring macOS 27 beta:

```swift
.package(url: "https://github.com/john-rocky/coreai-kit", branch: "main")
// target dependency: .product(name: "CoreAIKit", package: "coreai-kit")
```

An E-series bundle is **two downloads**, the decoder and its paired PLE tables, addressed as two
paths inside this repo:

```swift
import CoreAIKit

let store = ModelStore.default
let decoderURL = try await store.download(
    ModelID("visible-cx/gemma-4-E4B-CoreAI",
            path: "gpu-pipelined/gemma4_e4b_qat_decode_int4lin_tbl_pf64"))
let tablesURL = try await store.download(
    ModelID("visible-cx/gemma-4-E4B-CoreAI",
            path: "ios-frontend/gemma4_e4b_qat_gather_raw"))

// engineVariant defaults to .pipelined here, which is the supported path for _tbl bundles.
let runtime = try await GemmaRuntime(
    decoderBundleAt: decoderURL, tablesAt: tablesURL, arch: .gemma4)
```

Notes that are not optional:

- **Do not pair a QAT bundle with non-QAT tables.**
- **Do not enable chunked prefill.** The `…_tbl` graph is S=1 on `main`; `GemmaRuntime` sets
  `COREAI_CHUNK_THRESHOLD=1` for you if it is unset. Leave it.
- **Guided decoding needs the sequential engine**, and the sequential engine needs
  `ple_table`/`ple_scale` supplied through `EngineOptions.staticInputBuffers` β€” it will refuse
  by name otherwise.
- **Budget 7.04 GiB of artifacts plus the in-flight floor before you start**, and release the
  inference function between sessions.
- Pass `revision:` a Hub commit hash to pin immutable bytes.

## Integrity

Core AI `.aimodel` bundles are **not byte-reproducible**: the exporter is not deterministic even
against itself. Verify by digesting the exact published bytes rather than by rebuilding. Every
bundle carries `main.hash`, the raw 32 bytes of `sha256(main.mlirb)`; on the Hub the same value
is recoverable from the LFS oid without fetching the file.

## Status

| Artifact | Status |
|---|---|
| `gpu-pipelined/…_tbl_pf64` (ctx 4096) | **QUALIFIED FOR BOUNDED-SHAPE GENERATION, ON A MACHINE THAT FITS IT** β€” measured: 10/10 parse, 10/10 enum-clean, 14.70/5.93 s/row, 9.5–26.4 tok/s, 9.99 GB peak footprint. Not qualified for free-form generation. Under the corrected two-artifact gate it is refused on 16 GB; the guided figures above were taken before that gate existed and are a real run on a machine that was over-committed. |
| `w4a8/…_a8_tbl_pf64` | **EXPERIMENTAL** β€” built, unmeasured. Same int4 weights and graph; the difference is an int8 quantize/dequantize pair on the inputs of every `F.linear`, calibrated on 128 synthetic samples (corpus ships in the folder). The compression toolchain's own documentation says the runtime fast path is `W_INT8 Γ— A_INT8` and a float weight path executes in floating point regardless of activation quantization; these weights use int4 scale-multiply dequant, so the expected TTFT gain may be zero or negative. It exists to be measured, not because a win is predicted. |
| `stable/…_stable_c16384` | **EXPERIMENTAL** β€” shape-stable decode contract, gated in torch, never run on a Mac. On the E2B sibling the memory fix is confirmed and decode falls to ~1.0 tok/s from the full-capacity cache read; the same capacity cost applies here. A working proof of the contract, not a deployable bundle. |
| `ring/…_ring_c16384`, `ring3/…_ring_c4096_gqafull` | **EXPERIMENTAL** β€” gated in torch, no on-device numbers. `ring/`'s E4B was never measured on a Mac. |

**No numerics gate has been run on device for any bundle in this repo.** The 10/10 results are
behavioural (parse rate, enum conformance, clean stop); a decode oracle against an fp32
reference has not been run.

## The shape-stable family, in detail

All three re-export the same weights and the same quantization onto a contract in which **no
input shape moves between steps**, removing the per-generated-token growth by design. They need
a host that feeds `position_ids` as the **absolute positions of the S new tokens only**; a host
that feeds the growing `0..N` prefix will write the KV cache at the wrong offset.

**Contract** (identical across all three):

```
main    IN  input_ids Int32 1x1  | position_ids Int32 1x1  | ple_table Int8 V x (L*ld) | ple_scale Float32 V
prefill IN  input_ids Int32 1x64 | position_ids Int32 1x64 | (same statics)
        ST  keyCache / valueCache Float16, literal extents
        OUT logits Float16 1 x S x 262144
```

`position_ids[0,0]` is also the cache slot the K/V for those tokens is written at. The mask β€”
causality, unwritten slots and the sliding window β€” is derived inside the graph from those
positions, so there is no mask input to bind. The context ceiling is **encoded in the graph**: a
different window needs a different export, not a manifest edit.

`stable/` gave every layer a full 16,384-slot cache and read all of it every step, which is
where the ~1 tok/s comes from. `ring/` keeps the contract byte for byte and changes what sits
behind it: sliding layers read a **576-slot ring** (the 512-token window plus one 64-token
prefill chunk), and the grouped-query head expansion is folded into the query rather than
materialised. `ring3/` folds that expansion **only in the full-attention layers**, where the key
length is long enough for the trade to pay, and keeps the stock gathered form in the short-key
sliding layers β€” E4B runs 35 sliding and 7 full attention layers, so the split matters more here
than on E2B.

| | E4B `stable/` | E4B `ring/` | E4B `ring3/` (c4096) |
|---|---|---|---|
| KV state shape | `[…, 16384, 512]` Γ—2 | `[1, 1, 2, 77056, 512]` Γ—2 | `[1, 1, 2, 27904, 512]` Γ—2 |
| KV bytes, both states | 1.61 GB | **315.6 MB** | **114.3 MB** |
| cache slots read per decode step | 393,216 | 77,056 | β€” |
| sliding / full layers | β€” | 20 Γ— 576 + 4 Γ— 16,384 | 20 Γ— 576 + 4 Γ— 4,096 |
| dynamic dimensions | none | none | none |

**One host precondition new in `ring/` and `ring3/`:** a `prefill` call's first position must be
a multiple of 64. The ring write is a fixed-width store at `p0 % 576`, and 576 is nine 64-token
chunks, so a 64-aligned chunk can never straddle the wrap. `stable/` tolerated an unaligned
chunk; the ring bundles do not.

**Equivalence gates**, all run in eager torch against the shipped graph before conversion:
`stable/` β€” greedy argmax identical on all tested decode steps including past the sliding
window, worst logits cosine 0.99999 (fp16), prefill bit-identical in fp32. `ring*/` β€” 1,600
prompt tokens as 25 chunks of 64 then 32 greedy decode steps, final position 1,631 so the ring
wraps about three times, fp32, 57 comparisons per run: **0 argmax mismatches**. No dynamic
dimension appears in either entrypoint of any of them.

## License

Google publishes the upstream QAT checkpoint under **Apache-2.0** with a `license_link` to the
[Gemma 4 license](https://ai.google.dev/gemma/docs/gemma_4_license), and this repo mirrors that
declaration. Use is governed by those terms and by the
[Gemma Prohibited Use Policy](https://ai.google.dev/gemma/prohibited_use_policy); the
obligations travel with any redistribution of these bundles, **including the gather-table
sidecar**, which is derived from the same weights. The contribution here is the conversion, not
the weights.

## levered2/ β€” gather-first embedding, int4 PLE, fp16 head

The Speed-tier build. Shape-stable windowed KV, ring sliding cache, length-bounded Metal SDPA,
cap 131,072 β€” plus three changes to what the vocab path costs, each measured on the compiled
graph rather than assumed.

**The embedding is gathered before it is dequantized.** The previous build compiled to

    mps.dequantize(si4)      -> tensor<262144x2560xf16>    1.342 GB, every pass
    mps.gather_nd(that, ids) -> tensor<1x1x2560xf16>       ONE ROW

so the whole table was reconstructed in fp16 to fetch 2,560 values and the int4 saving
evaporated at run time. This revision packs two 4-bit codes per byte and gathers the PACKED
row. Verified on the exported graph β€” every 262144-row gather is now over an integer table:

    gather_along_axis(262144x1280xsi8) -> 1x1280xsi8     embedding
    gather_along_axis(262144x80xf16)   -> 1x80xf16       its block scales
    gather_along_axis(262144x5376xsi8) -> 1x5376xsi8     PLE table
    gather_along_axis(262144x336xf16)  -> 1x336xf16      PLE scales

**The PLE table is 4-bit**, per-block-32, and gated: 42 layers, 256 teacher-forced positions,
against the checkpoint's own bf16 rows β€” **32.936 dB, 255/256 top-1, reference greedy token in
the top-5 on every position**, for 1.4766 GiB against int8's 2.6260.

**The head is fp16.** The int4 head reconstructed a full `[262144, 2560]` fp16 table into its
matmul. Element accounting across the two compiled assets: `Int4` βˆ’691,200,000 elements,
`Float16` **+671,088,640 = 262,144 x 2,560 exactly**, and `blockwise_shift_scale` **688 -> 686**
β€” one dequantize removed per entrypoint. The matmul now reads the constant out of the blob.

The embedding's own lineage gate, same method: **37.215 dB, 256/256 top-1**.

    blob 3,957,723,854 + PLE 1,585,446,912 = 5,543,170,766 artifacts

## PAIR IT WITH THE TABLES IN THIS FOLDER

`levered2/ios-frontend/gemma4_e4b_qat_gather_raw/` β€” `embed_per_layer.i4` (packed nibbles, `+8`
biased) and `embed_per_layer.scale.f16` (per-32-block scales, rows stored 352 wide for the
engine's 32-element row stride, the trailing 16 unread). **The repo's int8 `ios-frontend/`
tables will not work**: the 4-bit unpack is compiled into this graph.

**Not yet run on a Mac.** No token generated, no decode speed claimed. Requires a kit with
`coreai-kit-gemma-int4-ple.patch` and an engine with the stable-position-contract and
sequential n-state patches.

### `levered2/gpu-pipelined/gemma4_e4b_qat_decode_int4lin_tbl_pf64_ple4_pe4_hf16_msdpa_g8_wkv_stable_c131072` β€” 3,989,920,339 bytes

    export_report.json                                           2,324 B  96b34dcca50696b8b2526e99dae7afea0936dad3ecdceb24f4a3d0bb883f6f26
    gemma4_e4b_qat_decode_int4lin_tbl_pf64_ple4_pe4_hf16_msdpa_g8_wkv_stable_c131072.aimodel/main.hash            32 B  568b0927d32636eeff95623c62a9b6bb207dc71b35c1f1c72302cd68cb5f01f3
    gemma4_e4b_qat_decode_int4lin_tbl_pf64_ple4_pe4_hf16_msdpa_g8_wkv_stable_c131072.aimodel/main.mlirb 3,957,723,854 B  c60c98a02f57790a8c7a527df6b91ba6c1b625d14d7ebd78640b52280cc52447
    gemma4_e4b_qat_decode_int4lin_tbl_pf64_ple4_pe4_hf16_msdpa_g8_wkv_stable_c131072.aimodel/metadata.json           105 B  c8e32c43e985105d978b858cf336cf8d121d8d052657bbed8f1036df259f4365
    metadata.json                                                1,897 B  f374ad6d7c79dd1941769a7d9271b6b05175e7616eb03f490fa2f2fa8b377aaf
    tokenizer/chat_template.jinja                               18,569 B  0a2c8073c878ab1da004bee933a998606537bbb62016310352c7285c3f01c5b5
    tokenizer/generation_config.json                               203 B  b69207f9be617e982d13cc273cce6fd88c98dda99a4bdc5e2d52ffe0a0d9f0a9
    tokenizer/tokenizer.json                                32,169,626 B  cc8d3a0ce36466ccc1278bf987df5f71db1719b9ca6b4118264f45cb627bfe0f
    tokenizer/tokenizer_config.json                              3,729 B  3ab5c7b94dc97d65ca7064496fa69b88ff875378e1cb7ee3e43070c3a8170999

### `levered2/ios-frontend/gemma4_e4b_qat_gather_raw` β€” 1,593,838,213 bytes

    embed_per_layer.i4                                   1,409,286,144 B  5651f14b2a4cead5619da219aafceaa96e78752af9c8bb2eb11e2230310839fa
    embed_per_layer.scale.f16                              184,549,376 B  a4b621325c55f5dcc60097178491262539ccd2458ca5ae8cb8caef398806e6ba
    meta.json                                                      721 B  9f225f51dc810bbb5ec91af8c28482ecbd56229a4979ca0bd1d6ad1189909b84
    quant_report.json                                            1,972 B  bf0ca04abebe20e729fede1c7c34e4fba2b2489c2f1204b784609fbb0b8767a6