alexwengg commited on
Commit
70fcf18
·
verified ·
1 Parent(s): 1260996

EmbeddingGemma 2 text encoder on Core ML (Neural Engine): 7 shared-weight functions, bf16 token table

Browse files
.gitattributes CHANGED
@@ -33,3 +33,5 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ embeddings.bf16 filter=lfs diff=lfs merge=lfs -text
37
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
EmbeddingGemma2Text.mlpackage/Data/com.apple.CoreML/model.mlmodel ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0df9aca02807cf490db6d4f1c68607a8c7f486ad23b376d1968fd3be5eb14677
3
+ size 7323423
EmbeddingGemma2Text.mlpackage/Data/com.apple.CoreML/weights/weight.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:75b97ab036ca87eae60560c4034a63d604201498dc4b6f6971dd0b1e4558d61f
3
+ size 276815424
EmbeddingGemma2Text.mlpackage/Manifest.json ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "fileFormatVersion": "1.0.0",
3
+ "itemInfoEntries": {
4
+ "0EA7802A-D153-497D-B221-C103BA870AB7": {
5
+ "author": "com.apple.CoreML",
6
+ "description": "CoreML Model Specification",
7
+ "name": "model.mlmodel",
8
+ "path": "com.apple.CoreML/model.mlmodel"
9
+ },
10
+ "DE6B4D44-991D-42A8-8048-7490E417887D": {
11
+ "author": "com.apple.CoreML",
12
+ "description": "CoreML Model Weights",
13
+ "name": "weights",
14
+ "path": "com.apple.CoreML/weights"
15
+ }
16
+ },
17
+ "rootModelIdentifier": "0EA7802A-D153-497D-B221-C103BA870AB7"
18
+ }
README.md ADDED
@@ -0,0 +1,81 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: google/embeddinggemma-2
4
+ library_name: coreml
5
+ pipeline_tag: feature-extraction
6
+ language:
7
+ - multilingual
8
+ tags:
9
+ - coreml
10
+ - apple-neural-engine
11
+ - on-device
12
+ - embeddings
13
+ - sentence-similarity
14
+ - gemma
15
+ ---
16
+
17
+ # EmbeddingGemma 2 (text) — Core ML, Neural Engine
18
+
19
+ The text encoder of [google/embeddinggemma-2](https://huggingface.co/google/embeddinggemma-2) converted to Core ML so it
20
+ runs entirely on the Apple Neural Engine. Weights are the original ones (stored fp16 in the model; the token table is
21
+ the original bf16). The vision and audio encoders are not included.
22
+
23
+ Output: the same 768-d, L2-normalized embedding as `SentenceTransformer("google/embeddinggemma-2").encode(...)`
24
+ (mean pooling over tokens, then normalization). Matryoshka truncation to 512/256/128 works as in the original: keep
25
+ the first N values and re-normalize.
26
+
27
+ ## Files
28
+
29
+ | File | What |
30
+ |---|---|
31
+ | `EmbeddingGemma2Text.mlpackage` | ML Program, fp16, 7 functions sharing one set of weights (271 MB) |
32
+ | `embeddings.bf16` | token embedding table, 262,144 × 512 raw bfloat16 (256 MB); look up rows, multiply by √512 |
33
+ | `tokenizer.json` | the original Gemma tokenizer, unchanged |
34
+ | `config.json` | shapes, function list, token ids, task prefixes |
35
+
36
+ ## Functions
37
+
38
+ Every function has fixed shapes, which is what keeps it on the Neural Engine. Variable-length (enumerated) shapes
39
+ put the whole graph on the CPU.
40
+
41
+ | Function | Inputs | Output |
42
+ |---|---|---|
43
+ | `embed_32` … `embed_512` | `inputs_embeds` [1, S, 512] fp16, `attention_mask` [1, S] fp16 (1 = token, right padding) | `embedding` [1, 768] |
44
+ | `pack_256` | `inputs_embeds` [1, 256, 512], `attention_bias` [1, 1, 256, 256] (0 within a text, −1e4 elsewhere), `positions` [256, 1] (restart at 0 per text), `pool` [8, 256] (1/len over each text's tokens) | `embedding` [8, 768] |
45
+
46
+ S ∈ {32, 48, 64, 128, 256, 512}. Inputs longer than 512 tokens must be truncated; on our test set capping at 512
47
+ tokens changed nothing measurable. At ≤ 512 tokens every sliding-window layer sees the whole sequence, so a padding
48
+ mask is all the model needs.
49
+
50
+ `pack_256` runs up to eight short texts in one call. Short calls are bound by streaming the weights from memory, so
51
+ one 256-token call with eight texts costs about the same as two single calls.
52
+
53
+ ## Use
54
+
55
+ 1. Prepend the task prefix (e.g. `title: none | text: ` for documents, `task: search result | query: ` for queries;
56
+ all in `config.json`).
57
+ 2. Tokenize with `tokenizer.json`: `<bos>` (2), tokens, `<eos>` (1).
58
+ 3. Look up each id's row in `embeddings.bf16` (bf16 → float: shift the 16 bits left by 16), multiply by √512 ≈
59
+ 22.627417, pass as fp16.
60
+ 4. Call the smallest `embed_S` that fits, or pack several texts into `pack_256`.
61
+
62
+ A Swift implementation (tokenizer, table lookup, packing, downloads) is in
63
+ [FluidUse](https://github.com/FluidInference/FluidUse): `EmbeddingGemma2Manager`.
64
+
65
+ Requires macOS 15 / iOS 18 (multifunction models). The first load on a device compiles all seven functions for the
66
+ Neural Engine, which takes several minutes once; Core ML caches the result.
67
+
68
+ ## Numbers (M5 Pro, macOS 27)
69
+
70
+ | | Result |
71
+ |---|---|
72
+ | Neural Engine placement | 3,862 / 3,862 ops (100%) in every function |
73
+ | Latency, one text | 32 tokens 2.6 ms · 64 tokens 3.2 ms · 128 tokens 5.6 ms · 256 tokens 11.3 ms · 512 tokens 27.4 ms |
74
+ | Throughput, short posts (`pack_256`, Swift) | 644 texts/s (one text per call: 307/s) |
75
+ | vs PyTorch fp32 (sentence-transformers) | cosine ≥ 0.9984, mean 0.99996 over 562 posts |
76
+
77
+ The model card warns against fp16: activations reach ~2,000, so squaring them inside RMSNorm overflows fp16. Every
78
+ RMSNorm here divides by the row's max-abs value first, which keeps the math in range (plain fp16 RMSNorm: cosine
79
+ 0.95 to the reference; this conversion: 0.9996). The final normalization is rescaled the same way.
80
+
81
+ License: Apache 2.0, same as the source model.
config.json ADDED
@@ -0,0 +1,27 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "base_model": "google/embeddinggemma-2",
3
+ "format": "coreml",
4
+ "modality": "text",
5
+ "precision": "fp16",
6
+ "hidden_size": 512,
7
+ "embedding_dim": 768,
8
+ "matryoshka_dims": [768, 512, 256, 128],
9
+ "embed_scale": 22.627417,
10
+ "token_table": {"file": "embeddings.bf16", "dtype": "bfloat16", "shape": [262144, 512]},
11
+ "functions": {
12
+ "embed_32": {"tokens": 32}, "embed_48": {"tokens": 48}, "embed_64": {"tokens": 64}, "embed_128": {"tokens": 128},
13
+ "embed_256": {"tokens": 256}, "embed_512": {"tokens": 512},
14
+ "pack_256": {"tokens": 256, "slots": 8}
15
+ },
16
+ "bos_token_id": 2, "eos_token_id": 1, "pad_token_id": 0,
17
+ "prompts": {
18
+ "document": "title: none | text: ",
19
+ "search_query": "task: search result | query: ",
20
+ "question_answering": "task: question answering | query: ",
21
+ "fact_checking": "task: fact checking | query: ",
22
+ "code_retrieval": "task: code retrieval | query: ",
23
+ "classification": "task: classification | query: ",
24
+ "clustering": "task: clustering | query: ",
25
+ "sentence_similarity": "task: sentence similarity | query: "
26
+ }
27
+ }
embeddings.bf16 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:56ebbdbfd706c827f3fae1d97aa39135c7ddc0f1e6d9b12f8565f84bbfbda877
3
+ size 268435456
tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4d777ef5bdc1aa36227abdfb77c3e49e7b9c892d16e1b6bda41c393504828be4
3
+ size 32170510