| --- |
| license: gemma |
| license_link: https://ai.google.dev/gemma/terms |
| base_model: google/embeddinggemma-300m |
| base_model_relation: quantized |
| library_name: coreml |
| pipeline_tag: feature-extraction |
| tags: |
| - coreml |
| - core-ml |
| - apple-silicon |
| - on-device |
| - ane |
| - quantized |
| - int8 |
| - embeddings |
| - sentence-embedding |
| - sentence-similarity |
| extra_gated_heading: Access EmbeddingGemma on Hugging Face |
| extra_gated_description: >- |
| This artifact is a derivative of google/embeddinggemma-300m and is governed by |
| the Gemma Terms of Use, the Gemma Prohibited Use Policy and the Gemma license. |
| --- |
| |
| # embeddinggemma-300m β Core ML (int8, seq 128, ANE) |
|
|
| `google/embeddinggemma-300m` as a compiled **Core ML** encoder for Apple silicon, mirrored and |
| rebuild-verified by [visible-cx](https://huggingface.co/visible-cx). One call in, one 768-d |
| L2-normalised embedding out. |
|
|
| **This is the one artifact in this org whose weights reproduce bit-exactly from the published |
| recipe.** An independent rebuild on a different operating system and CPU architecture produced a |
| `weight.bin` identical to the published one β SHA-256 `f81f60ebβ¦`, **0 of 308,616,576 bytes |
| differing**. Every Core AI `.aimodel` bundle in this org has to rest on digests instead, because |
| that exporter is not deterministic even against itself. |
|
|
| > **This is Core ML, not Core AI.** It is an `.mlmodelc` loaded through `MLModel`, and the |
| > `CoreAIKitEmbeddings.TextEmbedder` used by the sibling |
| > [`visible-cx/embeddinggemma-300m-CoreAI`](https://huggingface.co/visible-cx/embeddinggemma-300m-CoreAI) |
| > repo will not load it β that type looks for a `*.aimodel` in the bundle directory. Pooling, |
| > the dense stack and L2 normalisation are inside this graph, so there is no host-side pooling |
| > to implement either way. |
| |
| ## Contents |
| |
| **`3fa12f0b97b8afe23264f76800afe14af4615ca5/`** β the pinned upstream artifact, byte for byte, |
| under its upstream revision as the directory name. **309,358,096 B** total. |
| |
| | File | Bytes | |
| |---|---:| |
| | `encoder.mlmodelc/weights/weight.bin` | 308,616,576 | |
| | `encoder.mlmodelc/model.mil` | 735,948 | |
| | `encoder.mlmodelc/metadata.json` | 2,570 | |
| | `encoder.mlmodelc/coremldata.bin` | 408 | |
| | `encoder.mlmodelc/analytics/coremldata.bin` | 243 | |
| | `model_config.json` | 2,351 | |
| |
| **`rebuild-verification/2026-08-17/`** β an independent rebuild of the same recipe on a |
| different operating system and CPU architecture, published so the reproducibility claim can be |
| checked rather than taken on faith. **309,346,242 B** total. |
| |
| | File | Bytes | |
| |---|---:| |
| | `encoder.mlpackage/Data/com.apple.CoreML/weights/weight.bin` | 308,616,576 | |
| | `encoder.mlpackage/Data/com.apple.CoreML/model.mlmodel` | 723,393 | |
| | `encoder.mlpackage/Manifest.json` | 617 | |
| | `model_config.json` | 2,351 | |
| | `SHA256SUMS` | 427 | |
| | `VERIFICATION.md` | 2,878 | |
| |
| The shapes differ deliberately: the pinned artifact is a **compiled** `.mlmodelc`, the rebuild is |
| the **uncompiled** `.mlpackage` the recipe emits. Compilation (`xcrun coremlcompiler`) is a |
| macOS-only step and was not performed on the rebuild host. |
| |
| ## Provenance |
| |
| | | | |
| |---|---| |
| | Base checkpoint | `google/embeddinggemma-300m` | |
| | Upstream bundle | [`mlboydaisuke/embeddinggemma-300m-coreml`](https://huggingface.co/mlboydaisuke/embeddinggemma-300m-coreml) @ `3fa12f0b97b8afe23264f76800afe14af4615ca5` | |
| | Recipe | `john-rocky/CoreML-LLM` β `conversion/build_embeddinggemma_bundle.py --max-seq-len 128 --quantize int8` | |
| | Format | Core ML `.mlmodelc` (compiled), int8 weights | |
| | Sequence length | **128** (static) | |
| | Output | 768-d embedding; mean pooling β dense stack β L2 normalise, all in-graph | |
| | ANE residency | ~99.80%, 1950/1954 ops β **upstream's published claim, not re-measured here** | |
| |
| The mirror was verified byte-exact against upstream on all six files of the published artifact. |
| |
| ## Requirements |
| |
| - **Apple silicon**, Core ML runtime. The encoder is shaped for the Neural Engine. |
| - **Static sequence length 128.** Inputs must be padded or truncated to 128 tokens; this is a |
| compile-time property of the artifact, not a runtime option. |
| - Weights β 0.31 GB resident. **No KV cache** β this is an encoder, so there is no per-token |
| memory growth and no context ladder. **Minimum practical machine memory: 8 GB.** |
| - `model_config.json` beside the encoder carries the pooling / dense / normalisation contract the |
| host must honour. |
| |
| Note the sequence-length difference from the Core AI artifact in this org, which is **seq 256**. |
| The two are not drop-in substitutes for each other, and vectors from one should not be compared |
| against vectors from the other. |
| |
| ## Measurements |
| |
| **No throughput or latency figure is published here**, and none has been taken. The ANE residency |
| figure above is upstream's published claim, not a measurement made here. |
| |
| One cross-runtime quality datapoint is on record: a cosine of **~0.966** on short text between |
| this Core ML encoder and the LiteRT `.tflite` of the same base model β measured against |
| previously installed copies rather than against the files in this repo, so read it as an |
| indication that the two runtimes agree closely on short text, not as a parity gate on these |
| bytes. |
| |
| ## Verification |
| |
| Weights and config reproduce bit-exactly from the recipe, across operating systems and CPU |
| architectures: |
| |
| | File class | Verdict | |
| |---|---| |
| | `weights/weight.bin` (308,616,576 B) | **IDENTICAL** β SHA-256 `f81f60ebβ¦`, 0 differing bytes | |
| | `model_config.json` (2,351 B) | **IDENTICAL** β SHA-256 `0b949875β¦` | |
| | tokenizer / config JSON emitted by the recipe | **IDENTICAL** β all files | |
| | `encoder.mlmodelc/model.mil`, `coremldata.bin` Γ2, `metadata.json` | **not produced on the rebuild host** β products of the macOS-only `xcrun coremlcompiler` step | |
| |
| So the claim is scoped precisely: **everything the recipe produces reproduces exactly; the |
| remaining four files are a macOS compile step that was not run.** Closing that gap means |
| compiling `rebuild-verification/2026-08-17/encoder.mlpackage` on a Mac and diffing the resulting |
| `encoder.mlmodelc` against the pinned artifact. `SHA256SUMS` and `VERIFICATION.md` in that folder |
| carry the rebuild's own receipts. |
| |
| ## Usage |
| |
| Core ML, no third-party package required. Load the compiled `.mlmodelc` directly and honour the |
| contract in `model_config.json`: |
| |
| ```swift |
| import CoreML |
| |
| let url = /* β¦/3fa12f0b97b8afe23264f76800afe14af4615ca5/encoder.mlmodelc */ |
| let config = MLModelConfiguration() |
| config.computeUnits = .cpuAndNeuralEngine // the encoder is shaped for the ANE |
| |
| let encoder = try MLModel(contentsOf: url, configuration: config) |
| // inputs: token ids + attention mask, padded/truncated to 128 |
| // output: a 768-d L2-normalised embedding β cosine == dot product |
| ``` |
| |
| Pooling, the dense stack and normalisation are already in the graph, so the output is directly |
| comparable; do not re-normalise or re-pool. |
| |
| Two things to hold onto: |
| |
| - **Pad or truncate to exactly 128 tokens.** The length is compiled in. |
| - **Apply EmbeddingGemma's query and document prompt prefixes yourself.** Unlike the Core AI |
| sibling β where `TextEmbedder` supplies them β nothing here does it for you, and embedding a |
| query with the document prefix quietly degrades retrieval. |
| |
| ## Status |
| |
| | Artifact | Status | |
| |---|---| |
| | `3fa12f0bβ¦/encoder.mlmodelc` + `model_config.json` | **SHIP** β byte-exact mirror of the pinned upstream revision, with the `weight.bin` SHA-256 matching and independently reproduced. | |
| | `rebuild-verification/2026-08-17/encoder.mlpackage` | **VERIFICATION EVIDENCE, not a runtime artifact** β uncompiled and never executed. | |
| |
| ## License |
| |
| EmbeddingGemma is Gemma-family and the upstream checkpoint is **gated** on Hugging Face. These |
| files are a derivative of `google/embeddinggemma-300m` and use is subject to the |
| [Gemma Terms of Use](https://ai.google.dev/gemma/terms) and the |
| [Gemma Prohibited Use Policy](https://ai.google.dev/gemma/prohibited_use_policy). Those terms |
| travel with the artifact and with any redistribution of it. The contribution here is the mirror |
| and the rebuild verification, not the weights. |
| |