File size: 8,107 Bytes
4d4168e 05f464d 4d4168e 05f464d 4d4168e 05f464d 4d4168e 05f464d 4d4168e 05f464d e274046 4d4168e 05f464d 4d4168e 05f464d 4d4168e 05f464d 4d4168e 05f464d 4d4168e 05f464d 4d4168e 05f464d 4d4168e 05f464d 4d4168e 05f464d 4d4168e 05f464d 4d4168e 05f464d 4d4168e 05f464d 4d4168e 05f464d 4d4168e 05f464d 4d4168e 05f464d 4d4168e e274046 4d4168e e274046 4d4168e e274046 4d4168e e274046 05f464d 4d4168e 05f464d 4d4168e 05f464d 4d4168e 05f464d | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 | ---
license: gemma
license_link: https://ai.google.dev/gemma/terms
base_model: google/embeddinggemma-300m
base_model_relation: quantized
library_name: coreml
pipeline_tag: feature-extraction
tags:
- coreml
- core-ml
- apple-silicon
- on-device
- ane
- quantized
- int8
- embeddings
- sentence-embedding
- sentence-similarity
extra_gated_heading: Access EmbeddingGemma on Hugging Face
extra_gated_description: >-
This artifact is a derivative of google/embeddinggemma-300m and is governed by
the Gemma Terms of Use, the Gemma Prohibited Use Policy and the Gemma license.
---
# embeddinggemma-300m β Core ML (int8, seq 128, ANE)
`google/embeddinggemma-300m` as a compiled **Core ML** encoder for Apple silicon, mirrored and
rebuild-verified by [visible-cx](https://huggingface.co/visible-cx). One call in, one 768-d
L2-normalised embedding out.
**This is the one artifact in this org whose weights reproduce bit-exactly from the published
recipe.** An independent rebuild on a different operating system and CPU architecture produced a
`weight.bin` identical to the published one β SHA-256 `f81f60ebβ¦`, **0 of 308,616,576 bytes
differing**. Every Core AI `.aimodel` bundle in this org has to rest on digests instead, because
that exporter is not deterministic even against itself.
> **This is Core ML, not Core AI.** It is an `.mlmodelc` loaded through `MLModel`, and the
> `CoreAIKitEmbeddings.TextEmbedder` used by the sibling
> [`visible-cx/embeddinggemma-300m-CoreAI`](https://huggingface.co/visible-cx/embeddinggemma-300m-CoreAI)
> repo will not load it β that type looks for a `*.aimodel` in the bundle directory. Pooling,
> the dense stack and L2 normalisation are inside this graph, so there is no host-side pooling
> to implement either way.
## Contents
**`3fa12f0b97b8afe23264f76800afe14af4615ca5/`** β the pinned upstream artifact, byte for byte,
under its upstream revision as the directory name. **309,358,096 B** total.
| File | Bytes |
|---|---:|
| `encoder.mlmodelc/weights/weight.bin` | 308,616,576 |
| `encoder.mlmodelc/model.mil` | 735,948 |
| `encoder.mlmodelc/metadata.json` | 2,570 |
| `encoder.mlmodelc/coremldata.bin` | 408 |
| `encoder.mlmodelc/analytics/coremldata.bin` | 243 |
| `model_config.json` | 2,351 |
**`rebuild-verification/2026-08-17/`** β an independent rebuild of the same recipe on a
different operating system and CPU architecture, published so the reproducibility claim can be
checked rather than taken on faith. **309,346,242 B** total.
| File | Bytes |
|---|---:|
| `encoder.mlpackage/Data/com.apple.CoreML/weights/weight.bin` | 308,616,576 |
| `encoder.mlpackage/Data/com.apple.CoreML/model.mlmodel` | 723,393 |
| `encoder.mlpackage/Manifest.json` | 617 |
| `model_config.json` | 2,351 |
| `SHA256SUMS` | 427 |
| `VERIFICATION.md` | 2,878 |
The shapes differ deliberately: the pinned artifact is a **compiled** `.mlmodelc`, the rebuild is
the **uncompiled** `.mlpackage` the recipe emits. Compilation (`xcrun coremlcompiler`) is a
macOS-only step and was not performed on the rebuild host.
## Provenance
| | |
|---|---|
| Base checkpoint | `google/embeddinggemma-300m` |
| Upstream bundle | [`mlboydaisuke/embeddinggemma-300m-coreml`](https://huggingface.co/mlboydaisuke/embeddinggemma-300m-coreml) @ `3fa12f0b97b8afe23264f76800afe14af4615ca5` |
| Recipe | `john-rocky/CoreML-LLM` β `conversion/build_embeddinggemma_bundle.py --max-seq-len 128 --quantize int8` |
| Format | Core ML `.mlmodelc` (compiled), int8 weights |
| Sequence length | **128** (static) |
| Output | 768-d embedding; mean pooling β dense stack β L2 normalise, all in-graph |
| ANE residency | ~99.80%, 1950/1954 ops β **upstream's published claim, not re-measured here** |
The mirror was verified byte-exact against upstream on all six files of the published artifact.
## Requirements
- **Apple silicon**, Core ML runtime. The encoder is shaped for the Neural Engine.
- **Static sequence length 128.** Inputs must be padded or truncated to 128 tokens; this is a
compile-time property of the artifact, not a runtime option.
- Weights β 0.31 GB resident. **No KV cache** β this is an encoder, so there is no per-token
memory growth and no context ladder. **Minimum practical machine memory: 8 GB.**
- `model_config.json` beside the encoder carries the pooling / dense / normalisation contract the
host must honour.
Note the sequence-length difference from the Core AI artifact in this org, which is **seq 256**.
The two are not drop-in substitutes for each other, and vectors from one should not be compared
against vectors from the other.
## Measurements
**No throughput or latency figure is published here**, and none has been taken. The ANE residency
figure above is upstream's published claim, not a measurement made here.
One cross-runtime quality datapoint is on record: a cosine of **~0.966** on short text between
this Core ML encoder and the LiteRT `.tflite` of the same base model β measured against
previously installed copies rather than against the files in this repo, so read it as an
indication that the two runtimes agree closely on short text, not as a parity gate on these
bytes.
## Verification
Weights and config reproduce bit-exactly from the recipe, across operating systems and CPU
architectures:
| File class | Verdict |
|---|---|
| `weights/weight.bin` (308,616,576 B) | **IDENTICAL** β SHA-256 `f81f60ebβ¦`, 0 differing bytes |
| `model_config.json` (2,351 B) | **IDENTICAL** β SHA-256 `0b949875β¦` |
| tokenizer / config JSON emitted by the recipe | **IDENTICAL** β all files |
| `encoder.mlmodelc/model.mil`, `coremldata.bin` Γ2, `metadata.json` | **not produced on the rebuild host** β products of the macOS-only `xcrun coremlcompiler` step |
So the claim is scoped precisely: **everything the recipe produces reproduces exactly; the
remaining four files are a macOS compile step that was not run.** Closing that gap means
compiling `rebuild-verification/2026-08-17/encoder.mlpackage` on a Mac and diffing the resulting
`encoder.mlmodelc` against the pinned artifact. `SHA256SUMS` and `VERIFICATION.md` in that folder
carry the rebuild's own receipts.
## Usage
Core ML, no third-party package required. Load the compiled `.mlmodelc` directly and honour the
contract in `model_config.json`:
```swift
import CoreML
let url = /* β¦/3fa12f0b97b8afe23264f76800afe14af4615ca5/encoder.mlmodelc */
let config = MLModelConfiguration()
config.computeUnits = .cpuAndNeuralEngine // the encoder is shaped for the ANE
let encoder = try MLModel(contentsOf: url, configuration: config)
// inputs: token ids + attention mask, padded/truncated to 128
// output: a 768-d L2-normalised embedding β cosine == dot product
```
Pooling, the dense stack and normalisation are already in the graph, so the output is directly
comparable; do not re-normalise or re-pool.
Two things to hold onto:
- **Pad or truncate to exactly 128 tokens.** The length is compiled in.
- **Apply EmbeddingGemma's query and document prompt prefixes yourself.** Unlike the Core AI
sibling β where `TextEmbedder` supplies them β nothing here does it for you, and embedding a
query with the document prefix quietly degrades retrieval.
## Status
| Artifact | Status |
|---|---|
| `3fa12f0bβ¦/encoder.mlmodelc` + `model_config.json` | **SHIP** β byte-exact mirror of the pinned upstream revision, with the `weight.bin` SHA-256 matching and independently reproduced. |
| `rebuild-verification/2026-08-17/encoder.mlpackage` | **VERIFICATION EVIDENCE, not a runtime artifact** β uncompiled and never executed. |
## License
EmbeddingGemma is Gemma-family and the upstream checkpoint is **gated** on Hugging Face. These
files are a derivative of `google/embeddinggemma-300m` and use is subject to the
[Gemma Terms of Use](https://ai.google.dev/gemma/terms) and the
[Gemma Prohibited Use Policy](https://ai.google.dev/gemma/prohibited_use_policy). Those terms
travel with the artifact and with any redistribution of it. The contribution here is the mirror
and the rebuild verification, not the weights.
|