license: gemma
license_link: https://ai.google.dev/gemma/terms
base_model: google/embeddinggemma-300m
base_model_relation: quantized
library_name: coreml
pipeline_tag: feature-extraction
tags:
- coreml
- core-ml
- apple-silicon
- on-device
- ane
- quantized
- int8
- embeddings
- sentence-embedding
- sentence-similarity
extra_gated_heading: Access EmbeddingGemma on Hugging Face
extra_gated_description: >-
This artifact is a derivative of google/embeddinggemma-300m and is governed by
the Gemma Terms of Use, the Gemma Prohibited Use Policy and the Gemma license.
embeddinggemma-300m β Core ML (int8, seq 128, ANE)
google/embeddinggemma-300m as a compiled Core ML encoder for Apple silicon, mirrored and
rebuild-verified by visible-cx. One call in, one 768-d
L2-normalised embedding out.
This is the one artifact in this org whose weights reproduce bit-exactly from the published
recipe. An independent rebuild on a different operating system and CPU architecture produced a
weight.bin identical to the published one β SHA-256 f81f60ebβ¦, 0 of 308,616,576 bytes
differing. Every Core AI .aimodel bundle in this org has to rest on digests instead, because
that exporter is not deterministic even against itself.
This is Core ML, not Core AI. It is an
.mlmodelcloaded throughMLModel, and theCoreAIKitEmbeddings.TextEmbedderused by the siblingvisible-cx/embeddinggemma-300m-CoreAIrepo will not load it β that type looks for a*.aimodelin the bundle directory. Pooling, the dense stack and L2 normalisation are inside this graph, so there is no host-side pooling to implement either way.
Contents
3fa12f0b97b8afe23264f76800afe14af4615ca5/ β the pinned upstream artifact, byte for byte,
under its upstream revision as the directory name. 309,358,096 B total.
| File | Bytes |
|---|---|
encoder.mlmodelc/weights/weight.bin |
308,616,576 |
encoder.mlmodelc/model.mil |
735,948 |
encoder.mlmodelc/metadata.json |
2,570 |
encoder.mlmodelc/coremldata.bin |
408 |
encoder.mlmodelc/analytics/coremldata.bin |
243 |
model_config.json |
2,351 |
rebuild-verification/2026-08-17/ β an independent rebuild of the same recipe on a
different operating system and CPU architecture, published so the reproducibility claim can be
checked rather than taken on faith. 309,346,242 B total.
| File | Bytes |
|---|---|
encoder.mlpackage/Data/com.apple.CoreML/weights/weight.bin |
308,616,576 |
encoder.mlpackage/Data/com.apple.CoreML/model.mlmodel |
723,393 |
encoder.mlpackage/Manifest.json |
617 |
model_config.json |
2,351 |
SHA256SUMS |
427 |
VERIFICATION.md |
2,878 |
The shapes differ deliberately: the pinned artifact is a compiled .mlmodelc, the rebuild is
the uncompiled .mlpackage the recipe emits. Compilation (xcrun coremlcompiler) is a
macOS-only step and was not performed on the rebuild host.
Provenance
| Base checkpoint | google/embeddinggemma-300m |
| Upstream bundle | mlboydaisuke/embeddinggemma-300m-coreml @ 3fa12f0b97b8afe23264f76800afe14af4615ca5 |
| Recipe | john-rocky/CoreML-LLM β conversion/build_embeddinggemma_bundle.py --max-seq-len 128 --quantize int8 |
| Format | Core ML .mlmodelc (compiled), int8 weights |
| Sequence length | 128 (static) |
| Output | 768-d embedding; mean pooling β dense stack β L2 normalise, all in-graph |
| ANE residency | ~99.80%, 1950/1954 ops β upstream's published claim, not re-measured here |
The mirror was verified byte-exact against upstream on all six files of the published artifact.
Requirements
- Apple silicon, Core ML runtime. The encoder is shaped for the Neural Engine.
- Static sequence length 128. Inputs must be padded or truncated to 128 tokens; this is a compile-time property of the artifact, not a runtime option.
- Weights β 0.31 GB resident. No KV cache β this is an encoder, so there is no per-token memory growth and no context ladder. Minimum practical machine memory: 8 GB.
model_config.jsonbeside the encoder carries the pooling / dense / normalisation contract the host must honour.
Note the sequence-length difference from the Core AI artifact in this org, which is seq 256. The two are not drop-in substitutes for each other, and vectors from one should not be compared against vectors from the other.
Measurements
No throughput or latency figure is published here, and none has been taken. The ANE residency figure above is upstream's published claim, not a measurement made here.
One cross-runtime quality datapoint is on record: a cosine of ~0.966 on short text between
this Core ML encoder and the LiteRT .tflite of the same base model β measured against
previously installed copies rather than against the files in this repo, so read it as an
indication that the two runtimes agree closely on short text, not as a parity gate on these
bytes.
Verification
Weights and config reproduce bit-exactly from the recipe, across operating systems and CPU architectures:
| File class | Verdict |
|---|---|
weights/weight.bin (308,616,576 B) |
IDENTICAL β SHA-256 f81f60ebβ¦, 0 differing bytes |
model_config.json (2,351 B) |
IDENTICAL β SHA-256 0b949875β¦ |
| tokenizer / config JSON emitted by the recipe | IDENTICAL β all files |
encoder.mlmodelc/model.mil, coremldata.bin Γ2, metadata.json |
not produced on the rebuild host β products of the macOS-only xcrun coremlcompiler step |
So the claim is scoped precisely: everything the recipe produces reproduces exactly; the
remaining four files are a macOS compile step that was not run. Closing that gap means
compiling rebuild-verification/2026-08-17/encoder.mlpackage on a Mac and diffing the resulting
encoder.mlmodelc against the pinned artifact. SHA256SUMS and VERIFICATION.md in that folder
carry the rebuild's own receipts.
Usage
Core ML, no third-party package required. Load the compiled .mlmodelc directly and honour the
contract in model_config.json:
import CoreML
let url = /* β¦/3fa12f0b97b8afe23264f76800afe14af4615ca5/encoder.mlmodelc */
let config = MLModelConfiguration()
config.computeUnits = .cpuAndNeuralEngine // the encoder is shaped for the ANE
let encoder = try MLModel(contentsOf: url, configuration: config)
// inputs: token ids + attention mask, padded/truncated to 128
// output: a 768-d L2-normalised embedding β cosine == dot product
Pooling, the dense stack and normalisation are already in the graph, so the output is directly comparable; do not re-normalise or re-pool.
Two things to hold onto:
- Pad or truncate to exactly 128 tokens. The length is compiled in.
- Apply EmbeddingGemma's query and document prompt prefixes yourself. Unlike the Core AI
sibling β where
TextEmbeddersupplies them β nothing here does it for you, and embedding a query with the document prefix quietly degrades retrieval.
Status
| Artifact | Status |
|---|---|
3fa12f0bβ¦/encoder.mlmodelc + model_config.json |
SHIP β byte-exact mirror of the pinned upstream revision, with the weight.bin SHA-256 matching and independently reproduced. |
rebuild-verification/2026-08-17/encoder.mlpackage |
VERIFICATION EVIDENCE, not a runtime artifact β uncompiled and never executed. |
License
EmbeddingGemma is Gemma-family and the upstream checkpoint is gated on Hugging Face. These
files are a derivative of google/embeddinggemma-300m and use is subject to the
Gemma Terms of Use and the
Gemma Prohibited Use Policy. Those terms
travel with the artifact and with any redistribution of it. The contribution here is the mirror
and the rebuild verification, not the weights.