| --- |
| license: apache-2.0 |
| license_link: https://ai.google.dev/gemma/docs/gemma_4_license |
| base_model: google/gemma-4-E4B-it-qat-q4_0-unquantized |
| base_model_relation: quantized |
| library_name: coreai |
| pipeline_tag: text-generation |
| tags: |
| - core-ai |
| - aimodel |
| - apple-silicon |
| - on-device |
| - coreai-kit |
| - quantized |
| - int4 |
| - qat |
| - gemma4 |
| --- |
| |
| # Gemma 4 E4B β Core AI (.aimodel) |
|
|
| `google/gemma-4-E4B-it-qat-q4_0-unquantized` converted to Core AI `.aimodel` bundles for Apple |
| silicon by [visible-cx](https://huggingface.co/visible-cx). These are derivative artifacts: |
| Google's QAT-trained weights rounded onto the int4 grid they were trained for and re-expressed |
| as a Core AI graph. They load through Core AI on macOS and are not usable by PyTorch, GGUF or |
| MLX. |
|
|
| Gemma 4 E4B uses **Per-Layer Embeddings**, so the working bundles take a large embedding gather |
| table as a *static graph input* rather than carrying it in the graph. That table ships in |
| `ios-frontend/` and **the bundles do not load without it**; a missing table produces a bare |
| input-arity error naming `ple_table`/`ple_scale`. It is 3.60 GB and it is part of the model's |
| memory cost, not a sidecar you can ignore. |
|
|
| > β οΈ **This model does not fit a 16 GB Mac under an honest memory gate.** Charged for both |
| > artifacts it loads, at the coefficient measured on its E2B sibling, E4B asks **8.99 GiB** β |
| > and on the reference 16 GB machine the check refused it, short by 3.20 GiB. It has never been |
| > run under a watchdog. See [Measurements](#measurements). |
|
|
| > β οΈ **Known issue β memory growth per generated token.** The Core AI runtime caches one |
| > compiled specialization per input-shape signature, and this export makes sequence length |
| > load-bearing: every generated token mints a new signature, retaining roughly **81 MB of GPU |
| > allocations per generated token** until the process exits. **Guided decoding does not protect |
| > you β shape reuse does.** Repeating a length you have already run replays from cache at no |
| > cost, so a fixed-schema, fixed-cap extraction workload stays bounded; long or variable-length |
| > generation grows without bound, and it is not fixable below a re-export. `stable/`, `ring/` |
| > and `ring3/` in this repo are those re-exports. |
|
|
| ## Contents |
|
|
| ### Production bundles |
|
|
| | Path | Bytes | Context | Functions | Status | |
| |---|---:|---|---|---| |
| | `gpu-pipelined/gemma4_e4b_qat_decode_int4lin_tbl_pf64` | 3,989,986,858 | 4096 | main + prefill | **QUALIFIED (guided / bounded shapes)** | |
| | `w4a8/gemma4_e4b_qat_decode_int4lin_a8_tbl_pf64` | 3,990,690,638 | 16384 | main + prefill | EXPERIMENTAL | |
|
|
| Each folder holds `<name>.aimodel/` (`main.mlirb` β 3.96 GB, `main.hash`, asset |
| `metadata.json`), a bundle-level `metadata.json`, and `tokenizer/` (`tokenizer.json` |
| 32,169,626 B, `tokenizer_config.json`, `generation_config.json`, `chat_template.jinja` |
| 18,569 B). The `w4a8` folder additionally ships its `calibration_corpus.jsonl` (35,045 B). |
|
|
| ### The PLE gather-table sidecar β required, not optional |
|
|
| | Path | Files | Bytes | |
| |---|---|---:| |
| | `ios-frontend/gemma4_e4b_qat_gather_raw/` | 7 | 3,601,859,815 | |
|
|
| | File | Bytes | |
| |---|---:| |
| | `embed_per_layer.i8` | 2,818,572,288 | |
| | `embed_tokens.i8` | 671,088,640 | |
| | `proj.f32` | 110,100,480 | |
| | `embed_per_layer.scale.f32` | 1,048,576 | |
| | `embed_tokens.scale.f32` | 1,048,576 | |
| | `proj_norm.f32` | 1,024 | |
| | `meta.json` | 231 | |
|
|
| Every `_tbl` bundle binds `ple_table` (from `embed_per_layer.i8`) and `ple_scale` (from |
| `embed_per_layer.scale.f32`) as **static** graph inputs. **A QAT bundle must be paired with the |
| QAT tables.** |
|
|
| ### Shape-stable re-exports |
|
|
| | Path | Bundle | Bytes | Context | |
| |---|---|---:|---| |
| | `stable/gpu-pipelined/` | `β¦_tbl_pf64_stable_c16384` | 3,989,939,812 | 16384 | |
| | `ring/gpu-pipelined/` | `β¦_tbl_pf64_ring_c16384` | 3,989,979,750 | 16384 | |
| | `ring3/gpu-pipelined/` | `β¦_tbl_pf64_ring_c4096_gqafull` | 3,989,929,508 | 4096 | |
|
|
| **Stop token:** every bundle declares `eos_token = "<turn|>"` (id 106), the turn terminator |
| Gemma 4 emits. `generation_config.json` independently lists `eos_token_id: [1, 106, 50]`. A |
| host that stops on the raw upstream `<eos>` will overrun every reply. |
|
|
| ## Provenance |
|
|
| | | | |
| |---|---| |
| | Base checkpoint | `google/gemma-4-E4B-it-qat-q4_0-unquantized` (ungated) | |
| | Zoo recipe | `gemma-4-e4b`, `status = "verified"` β `int4lin` | |
| | Recipe (pf64) | `export_gemma4_pf_pipelined.py --pf 64` with `--tbl` and `--raw-dir` pointed at the gather table above | |
| | Toolchain base | `apple/coreai-models` @ `b1cb71b8522d99408059fa0b98b8742171bcb0b8` + the [coreai-model-zoo](https://github.com/john-rocky/coreai-model-zoo) python overlay | |
| | Toolchain | `coreai-torch 0.4.1`, `coreai-core 1.0.0b2`, `coreai-opt 0.2.1`, `torch 2.9.0` | |
| | Producer fingerprint | `coreai-core 1.0.0b2` on every inner asset `metadata.json` | |
| | Weight format | **int4, per-block-32** (`int4lin`, symmetric-with-clipping) β the ggml q4_0 grid the QAT checkpoint was trained on | |
| | Vocab | 262,144 | |
| | Export functions | `main` (S=1 decode) and, in `_pf64` bundles, `prefill` (S=64 chunked prefill) | |
|
|
| "QAT-unquantized" means QAT-*trained*, stored full width; the int4 rounding happens at export |
| onto the grid training already targeted. |
|
|
| `_tbl` = the PLE gather table is bound as a static graph input. `_pf64` = a second entrypoint, |
| `prefill`, with a static query width of 64 (`function_map: {"main": ["main", "prefill"]}`). |
|
|
| The `gpu-pipelined/` bundle mirrors the zoo's verified `gemma-4-e4b` recipe with the `pf64` |
| multifunction addition; the gather table follows the |
| [coreai-model-zoo](https://github.com/john-rocky/coreai-model-zoo) / |
| [mlboydaisuke](https://huggingface.co/mlboydaisuke) artifact layout. |
|
|
| ## Requirements |
|
|
| - **Apple silicon Mac**, Core AI runtime. Practically, **24 GB or more** β see the arithmetic |
| below. |
| - **Engine contract: 4 inputs** β `input_ids`, `position_ids`, plus static `ple_table` and |
| `ple_scale`: |
| - **Pipelined engine** β binds the statics zero-copy over the caller's buffer, but does not |
| expose logits, so no grammar-constrained decoding. This is CoreAIKit's default for `_tbl` |
| bundles. |
| - **Sequential engine** β the only logits-capable engine, and therefore the only path for |
| guided decoding. It binds everything beyond `input_ids`/`position_ids` from |
| `EngineOptions.staticInputBuffers`, and **a caller that does not supply those buffers gets a |
| load failure naming them**, not a fallback. It pays a **one-time copy of every constant |
| input at load β 2.69 GB for E4B**; binding the table as a raw view instead drives peak |
| footprint to 14.9 GB and faults the runtime mid-prefill. |
| - **States:** `keyCache` / `valueCache` `Float16, 24 Γ 1 Γ 2 Γ ? Γ 512`. Dynamic sequence dim β |
| `GrowingKVCache` (initial 256, doubling). |
| - **KV cost: 98,304 bytes per token** (fp16) β 403 MB at 4096, 1.61 GB at 16384. **KV is not the |
| ceiling at this tier; the PLE table is.** |
| - **Run one E4B session per process.** Wired memory is not reclaimed until the process exits |
| unless the host releases the inference function; a second full run in the same process drove |
| wired memory to 13.2 GB on a 16 GB machine. |
| - Sliding-window note: E4B interleaves sliding and full attention layers across 42 layers |
| collapsed to 24 KV slots. The export models the window in the *mask*, not in memory β sliding |
| layers ride full-length slots and zero-pad head_dim 256β512, so 83.3% of KV bytes sit in |
| windowed slots and 41.7% is dead padding. |
| - The bundle manifest declares `runtime_env COREAI_CHUNK_THRESHOLD=1`. Both engines derive the |
| static query width from the graph. |
| |
| ## Measurements |
|
|
| Measured on a **16 GB Apple silicon Mac (M2 Pro, macOS 27 beta)** unless stated. |
|
|
| ### The memory verdict β arithmetic, not a run |
|
|
| **E4B has never been run under a watchdog.** What exists is its artifact inventory charged at |
| the coefficient measured on its E2B sibling (n=2 completed runs, `wired Γ· (blob + PLE tables)` = |
| **1.101**): |
|
|
| ``` |
| compiled blob 3,955,446,640 B = 3.6838 GiB |
| PLE tables 3,601,859,584 B = 3.3545 GiB |
| ------------------------------------------------ |
| artifacts = 7.0383 GiB |
| x 1.10 (measured on E2B) = 7.742 GiB |
| + 1.25 GiB in-flight floor = 8.99 GiB required |
| |
| reclaimable at the check = 5.79 GiB -> REFUSED, short by 3.20 GiB |
| ``` |
|
|
| (The PLE figure is the binary payload the runtime binds; the published folder is 231 bytes |
| larger because it also carries `meta.json`.) |
|
|
| This is a **3.67 GiB correction against the gate this project previously shipped**, which |
| charged `blob Γ 1.106 + 1.25` = 5.324 GiB and would have admitted the model. The cause is the |
| same one-line error in both directions across the E-series: **the runtime loads two artifacts |
| and the law was only ever told about the first.** (E2B carried the same defect at 2.87 GiB.) |
|
|
| Stated honestly: **the refusal is arithmetic on E2B's coefficient, not a measurement of E4B.** |
|
|
| For scale, the compile-only readings that do exist for this bundle: bundle 3.716 GiB, compiled |
| blob **3.684 GiB**, graph constant **β0.002 GiB** (Gemma's compiled artifact is very slightly |
| *smaller* than its bundle, unlike the LFMs at 1.18β1.34Γ). |
|
|
| ### Guided structured output |
|
|
| 10-sample harness, guided JSON-constrained decoding against a fixed schema, greedy, **sequential |
| engine** with the PLE tables bound as static inputs, `reset()` between samples, 128-token cap. |
|
|
| | | `gpu-pipelined/β¦_tbl_pf64` | |
| |---|---| |
| | Load | 16.7 s | |
| | Guided JSON parse | **10/10** | |
| | Enum-clean | **10/10** | |
| | s/row (long samples) | **14.70** | |
| | s/row (short samples) | **5.93** | |
| | Decode | **9.5β26.4 tok/s** (mean 17.8) | |
| | TTFT | **1.23β6.51 s** | |
| | Peak footprint | 9.99 GB | |
| | Max RSS | 9.63 GB | |
| | Stop | `<turn\|>`, clean self-stop on every sample | |
|
|
| E4B is memory-bound at the 16 GB tier and its per-sample cost is sensitive to what else is |
| resident β expect roughly 2Γ these figures when the model has to share. Note the spread in the |
| decode column: that variance *is* the memory pressure. |
|
|
| **The grammar is close to free; prefill is the cost.** Decode throughput is essentially |
| unchanged from unguided pipelined measurements; guided samples cost ~25% more because the |
| sequential engine's prefill is step-synchronous, not because of the constraint. |
|
|
| **Enum conformance is entirely the grammar's doing.** Unguided, this model emits an off-schema |
| enum value in all ten samples. Guided, all ten are correct, because an off-enum token is |
| unsamplable. |
|
|
| Published Mac figures of ~55.8 tok/s for E4B decode should be read as larger-machine figures. |
|
|
| ### Unguided workload β memory-capped |
|
|
| Free-form generation from a fixed prompt. **Every generation length below is a memory cap, not a |
| model stop** β see the known issue. |
|
|
| | depth | prompt tokens | TTFT | decode | generated | wall | peak footprint | |
| |---|---|---|---|---|---|---| |
| | 3.4k | 3,314 | 68.46 s | 0.78 tok/s | 64 (capped) | 151.6 s | **16.91 GB** | |
| | 8k | 7,907 | 159.1 s | 0.52 tok/s | 32 (capped) | 218.8 s | **31.33 GB** | |
|
|
| Output quality up to the cap is sound; the limit is memory, not capability. |
|
|
| **Needle-in-haystack recall at 8k: β₯2/3.** Three distinctive facts planted at 10/50/90% of the |
| filler; the 32-token memory cap truncated the answer mid-fact-2, so fact 3 was never reachable. |
| 15k was not attempted. |
|
|
| ### What this model is and is not, on this stack |
|
|
| The shape-signature growth sets a ceiling no cap works around: at ~81 MB/token over a ~9.6 GB |
| intercept, a 900-token report costs ~54 GB and a 600-token insight card ~58 GB. The caps that do |
| survive are shorter than a single card. **Gemma 4 is an enrichment model on this stack β |
| bounded, repeating shapes β and it is not a report, RAG or long-insight model. The blocker is |
| the export, not the weights.** |
|
|
| ## Usage |
|
|
| Swift Package Manager, via [CoreAIKit](https://github.com/john-rocky/coreai-kit) β a community |
| package, not affiliated with Apple, requiring macOS 27 beta: |
|
|
| ```swift |
| .package(url: "https://github.com/john-rocky/coreai-kit", branch: "main") |
| // target dependency: .product(name: "CoreAIKit", package: "coreai-kit") |
| ``` |
|
|
| An E-series bundle is **two downloads**, the decoder and its paired PLE tables, addressed as two |
| paths inside this repo: |
|
|
| ```swift |
| import CoreAIKit |
| |
| let store = ModelStore.default |
| let decoderURL = try await store.download( |
| ModelID("visible-cx/gemma-4-E4B-CoreAI", |
| path: "gpu-pipelined/gemma4_e4b_qat_decode_int4lin_tbl_pf64")) |
| let tablesURL = try await store.download( |
| ModelID("visible-cx/gemma-4-E4B-CoreAI", |
| path: "ios-frontend/gemma4_e4b_qat_gather_raw")) |
| |
| // engineVariant defaults to .pipelined here, which is the supported path for _tbl bundles. |
| let runtime = try await GemmaRuntime( |
| decoderBundleAt: decoderURL, tablesAt: tablesURL, arch: .gemma4) |
| ``` |
|
|
| Notes that are not optional: |
|
|
| - **Do not pair a QAT bundle with non-QAT tables.** |
| - **Do not enable chunked prefill.** The `β¦_tbl` graph is S=1 on `main`; `GemmaRuntime` sets |
| `COREAI_CHUNK_THRESHOLD=1` for you if it is unset. Leave it. |
| - **Guided decoding needs the sequential engine**, and the sequential engine needs |
| `ple_table`/`ple_scale` supplied through `EngineOptions.staticInputBuffers` β it will refuse |
| by name otherwise. |
| - **Budget 7.04 GiB of artifacts plus the in-flight floor before you start**, and release the |
| inference function between sessions. |
| - Pass `revision:` a Hub commit hash to pin immutable bytes. |
|
|
| ## Integrity |
|
|
| Core AI `.aimodel` bundles are **not byte-reproducible**: the exporter is not deterministic even |
| against itself. Verify by digesting the exact published bytes rather than by rebuilding. Every |
| bundle carries `main.hash`, the raw 32 bytes of `sha256(main.mlirb)`; on the Hub the same value |
| is recoverable from the LFS oid without fetching the file. |
|
|
| ## Status |
|
|
| | Artifact | Status | |
| |---|---| |
| | `gpu-pipelined/β¦_tbl_pf64` (ctx 4096) | **QUALIFIED FOR BOUNDED-SHAPE GENERATION, ON A MACHINE THAT FITS IT** β measured: 10/10 parse, 10/10 enum-clean, 14.70/5.93 s/row, 9.5β26.4 tok/s, 9.99 GB peak footprint. Not qualified for free-form generation. Under the corrected two-artifact gate it is refused on 16 GB; the guided figures above were taken before that gate existed and are a real run on a machine that was over-committed. | |
| | `w4a8/β¦_a8_tbl_pf64` | **EXPERIMENTAL** β built, unmeasured. Same int4 weights and graph; the difference is an int8 quantize/dequantize pair on the inputs of every `F.linear`, calibrated on 128 synthetic samples (corpus ships in the folder). The compression toolchain's own documentation says the runtime fast path is `W_INT8 Γ A_INT8` and a float weight path executes in floating point regardless of activation quantization; these weights use int4 scale-multiply dequant, so the expected TTFT gain may be zero or negative. It exists to be measured, not because a win is predicted. | |
| | `stable/β¦_stable_c16384` | **EXPERIMENTAL** β shape-stable decode contract, gated in torch, never run on a Mac. On the E2B sibling the memory fix is confirmed and decode falls to ~1.0 tok/s from the full-capacity cache read; the same capacity cost applies here. A working proof of the contract, not a deployable bundle. | |
| | `ring/β¦_ring_c16384`, `ring3/β¦_ring_c4096_gqafull` | **EXPERIMENTAL** β gated in torch, no on-device numbers. `ring/`'s E4B was never measured on a Mac. | |
|
|
| **No numerics gate has been run on device for any bundle in this repo.** The 10/10 results are |
| behavioural (parse rate, enum conformance, clean stop); a decode oracle against an fp32 |
| reference has not been run. |
|
|
| ## The shape-stable family, in detail |
|
|
| All three re-export the same weights and the same quantization onto a contract in which **no |
| input shape moves between steps**, removing the per-generated-token growth by design. They need |
| a host that feeds `position_ids` as the **absolute positions of the S new tokens only**; a host |
| that feeds the growing `0..N` prefix will write the KV cache at the wrong offset. |
|
|
| **Contract** (identical across all three): |
|
|
| ``` |
| main IN input_ids Int32 1x1 | position_ids Int32 1x1 | ple_table Int8 V x (L*ld) | ple_scale Float32 V |
| prefill IN input_ids Int32 1x64 | position_ids Int32 1x64 | (same statics) |
| ST keyCache / valueCache Float16, literal extents |
| OUT logits Float16 1 x S x 262144 |
| ``` |
|
|
| `position_ids[0,0]` is also the cache slot the K/V for those tokens is written at. The mask β |
| causality, unwritten slots and the sliding window β is derived inside the graph from those |
| positions, so there is no mask input to bind. The context ceiling is **encoded in the graph**: a |
| different window needs a different export, not a manifest edit. |
|
|
| `stable/` gave every layer a full 16,384-slot cache and read all of it every step, which is |
| where the ~1 tok/s comes from. `ring/` keeps the contract byte for byte and changes what sits |
| behind it: sliding layers read a **576-slot ring** (the 512-token window plus one 64-token |
| prefill chunk), and the grouped-query head expansion is folded into the query rather than |
| materialised. `ring3/` folds that expansion **only in the full-attention layers**, where the key |
| length is long enough for the trade to pay, and keeps the stock gathered form in the short-key |
| sliding layers β E4B runs 35 sliding and 7 full attention layers, so the split matters more here |
| than on E2B. |
|
|
| | | E4B `stable/` | E4B `ring/` | E4B `ring3/` (c4096) | |
| |---|---|---|---| |
| | KV state shape | `[β¦, 16384, 512]` Γ2 | `[1, 1, 2, 77056, 512]` Γ2 | `[1, 1, 2, 27904, 512]` Γ2 | |
| | KV bytes, both states | 1.61 GB | **315.6 MB** | **114.3 MB** | |
| | cache slots read per decode step | 393,216 | 77,056 | β | |
| | sliding / full layers | β | 20 Γ 576 + 4 Γ 16,384 | 20 Γ 576 + 4 Γ 4,096 | |
| | dynamic dimensions | none | none | none | |
|
|
| **One host precondition new in `ring/` and `ring3/`:** a `prefill` call's first position must be |
| a multiple of 64. The ring write is a fixed-width store at `p0 % 576`, and 576 is nine 64-token |
| chunks, so a 64-aligned chunk can never straddle the wrap. `stable/` tolerated an unaligned |
| chunk; the ring bundles do not. |
|
|
| **Equivalence gates**, all run in eager torch against the shipped graph before conversion: |
| `stable/` β greedy argmax identical on all tested decode steps including past the sliding |
| window, worst logits cosine 0.99999 (fp16), prefill bit-identical in fp32. `ring*/` β 1,600 |
| prompt tokens as 25 chunks of 64 then 32 greedy decode steps, final position 1,631 so the ring |
| wraps about three times, fp32, 57 comparisons per run: **0 argmax mismatches**. No dynamic |
| dimension appears in either entrypoint of any of them. |
|
|
| ## License |
|
|
| Google publishes the upstream QAT checkpoint under **Apache-2.0** with a `license_link` to the |
| [Gemma 4 license](https://ai.google.dev/gemma/docs/gemma_4_license), and this repo mirrors that |
| declaration. Use is governed by those terms and by the |
| [Gemma Prohibited Use Policy](https://ai.google.dev/gemma/prohibited_use_policy); the |
| obligations travel with any redistribution of these bundles, **including the gather-table |
| sidecar**, which is derived from the same weights. The contribution here is the conversion, not |
| the weights. |
|
|
| ## levered2/ β gather-first embedding, int4 PLE, fp16 head |
|
|
| The Speed-tier build. Shape-stable windowed KV, ring sliding cache, length-bounded Metal SDPA, |
| cap 131,072 β plus three changes to what the vocab path costs, each measured on the compiled |
| graph rather than assumed. |
|
|
| **The embedding is gathered before it is dequantized.** The previous build compiled to |
|
|
| mps.dequantize(si4) -> tensor<262144x2560xf16> 1.342 GB, every pass |
| mps.gather_nd(that, ids) -> tensor<1x1x2560xf16> ONE ROW |
| |
| so the whole table was reconstructed in fp16 to fetch 2,560 values and the int4 saving |
| evaporated at run time. This revision packs two 4-bit codes per byte and gathers the PACKED |
| row. Verified on the exported graph β every 262144-row gather is now over an integer table: |
|
|
| gather_along_axis(262144x1280xsi8) -> 1x1280xsi8 embedding |
| gather_along_axis(262144x80xf16) -> 1x80xf16 its block scales |
| gather_along_axis(262144x5376xsi8) -> 1x5376xsi8 PLE table |
| gather_along_axis(262144x336xf16) -> 1x336xf16 PLE scales |
| |
| **The PLE table is 4-bit**, per-block-32, and gated: 42 layers, 256 teacher-forced positions, |
| against the checkpoint's own bf16 rows β **32.936 dB, 255/256 top-1, reference greedy token in |
| the top-5 on every position**, for 1.4766 GiB against int8's 2.6260. |
|
|
| **The head is fp16.** The int4 head reconstructed a full `[262144, 2560]` fp16 table into its |
| matmul. Element accounting across the two compiled assets: `Int4` β691,200,000 elements, |
| `Float16` **+671,088,640 = 262,144 x 2,560 exactly**, and `blockwise_shift_scale` **688 -> 686** |
| β one dequantize removed per entrypoint. The matmul now reads the constant out of the blob. |
|
|
| The embedding's own lineage gate, same method: **37.215 dB, 256/256 top-1**. |
|
|
| blob 3,957,723,854 + PLE 1,585,446,912 = 5,543,170,766 artifacts |
| |
| ## PAIR IT WITH THE TABLES IN THIS FOLDER |
|
|
| `levered2/ios-frontend/gemma4_e4b_qat_gather_raw/` β `embed_per_layer.i4` (packed nibbles, `+8` |
| biased) and `embed_per_layer.scale.f16` (per-32-block scales, rows stored 352 wide for the |
| engine's 32-element row stride, the trailing 16 unread). **The repo's int8 `ios-frontend/` |
| tables will not work**: the 4-bit unpack is compiled into this graph. |
|
|
| **Not yet run on a Mac.** No token generated, no decode speed claimed. Requires a kit with |
| `coreai-kit-gemma-int4-ple.patch` and an engine with the stable-position-contract and |
| sequential n-state patches. |
|
|
| ### `levered2/gpu-pipelined/gemma4_e4b_qat_decode_int4lin_tbl_pf64_ple4_pe4_hf16_msdpa_g8_wkv_stable_c131072` β 3,989,920,339 bytes |
|
|
| export_report.json 2,324 B 96b34dcca50696b8b2526e99dae7afea0936dad3ecdceb24f4a3d0bb883f6f26 |
| gemma4_e4b_qat_decode_int4lin_tbl_pf64_ple4_pe4_hf16_msdpa_g8_wkv_stable_c131072.aimodel/main.hash 32 B 568b0927d32636eeff95623c62a9b6bb207dc71b35c1f1c72302cd68cb5f01f3 |
| gemma4_e4b_qat_decode_int4lin_tbl_pf64_ple4_pe4_hf16_msdpa_g8_wkv_stable_c131072.aimodel/main.mlirb 3,957,723,854 B c60c98a02f57790a8c7a527df6b91ba6c1b625d14d7ebd78640b52280cc52447 |
| gemma4_e4b_qat_decode_int4lin_tbl_pf64_ple4_pe4_hf16_msdpa_g8_wkv_stable_c131072.aimodel/metadata.json 105 B c8e32c43e985105d978b858cf336cf8d121d8d052657bbed8f1036df259f4365 |
| metadata.json 1,897 B f374ad6d7c79dd1941769a7d9271b6b05175e7616eb03f490fa2f2fa8b377aaf |
| tokenizer/chat_template.jinja 18,569 B 0a2c8073c878ab1da004bee933a998606537bbb62016310352c7285c3f01c5b5 |
| tokenizer/generation_config.json 203 B b69207f9be617e982d13cc273cce6fd88c98dda99a4bdc5e2d52ffe0a0d9f0a9 |
| tokenizer/tokenizer.json 32,169,626 B cc8d3a0ce36466ccc1278bf987df5f71db1719b9ca6b4118264f45cb627bfe0f |
| tokenizer/tokenizer_config.json 3,729 B 3ab5c7b94dc97d65ca7064496fa69b88ff875378e1cb7ee3e43070c3a8170999 |
| |
| ### `levered2/ios-frontend/gemma4_e4b_qat_gather_raw` β 1,593,838,213 bytes |
|
|
| embed_per_layer.i4 1,409,286,144 B 5651f14b2a4cead5619da219aafceaa96e78752af9c8bb2eb11e2230310839fa |
| embed_per_layer.scale.f16 184,549,376 B a4b621325c55f5dcc60097178491262539ccd2458ca5ae8cb8caef398806e6ba |
| meta.json 721 B 9f225f51dc810bbb5ec91af8c28482ecbd56229a4979ca0bd1d6ad1189909b84 |
| quant_report.json 1,972 B bf0ca04abebe20e729fede1c7c34e4fba2b2489c2f1204b784609fbb0b8767a6 |
| |