EmbeddingGemma 2 โ OpenVINO IR (fp16)
i made this repo because this model run faster on my cpu-only server than llama.cpp;
google/embeddinggemma-2 converted to OpenVINO IR, three towers, pure CPU inference.
| Tower | Files | Size | Output |
|---|---|---|---|
| Text | openvino/model.xml + .bin |
545 MB | sentence_embedding [b, 768] |
| Vision | openvino/vision_model.xml + .bin |
337 MB | image_features [n, 512] |
| Audio | openvino/audio_model.xml + .bin |
588 MB | audio_features [n, 512] |
The .xml holds the topology and the .bin the weights; keep them side by side with a
matching name prefix.
Usage
pip install openvino tokenizers numpy pillow
python demo.py
dog.jpeg x captions:
a dog 0.725
a panda 0.576
a penguin using ubuntu computer 0.505
- Outputs are L2-normalized, so a dot product is the cosine similarity.
- Retrieval prefixes: queries take
task: search result | query:, documents taketitle: none | text:. - Images: vision tower produces
[N, 512]soft tokens, which the text tower fuses. Put<|image|>(258880) placeholders ininput_idsand pass the features viaimage_features. Audio works the same way with<|audio|>(258881); video frames each go through the vision tower. - Image preprocessing (aspect-ratio resize to a multiple of
3*16, 16x16 patchify, pad to280*3ยฒ=2520patches withpixel_position_ids = -1) is inlined indemo.pyand only needs numpy and PIL.
Speed on 8 CPU threads: 61 ms for a short query, 172 ms for a long document, 37 ms/item at batch 8, 2.6 s for one image, 0.49 s for 11 s of audio.
Conversion
From onnx-community/embeddinggemma-2-ONNX (Apache-2.0) to OpenVINO 2026.x fp16, see
tools/convert_ov.py and tools/convert_mm.py.
The upstream export keeps contrib ops that only the ORT kernel implements correctly. The text
tower's com.microsoft.RotaryEmbedding and the vision tower's com.microsoft.MultiHeadAttention
are replaced with standard-domain explicit ops, which removes the cross-engine semantic gap.
Against the original ONNX baseline, per-token cosine is 1.000000 for text, >=0.9946 for vision and 1.000000 for audio. Cross-modal retrieval matches the official transformers.js numbers (image to "cats" 0.746, audio to "speech" 0.772).
NPU and GPU
Measured on Meteor Lake (Intel Arc iGPU + AI Boost NPU):
- NPU: the text tower works, but shapes must be static, batch must be 1, and the trailing L2 normalization is dropped so it has to be reapplied by hand. The vision tower fails to compile (VPUX pass error). No net speedup on short texts.
- GPU: enumerates 0 Level Zero devices in a containerized environment; run on the host instead.
Unofficial conversion; see this repo's acceptance numbers. Original model by Google (Apache-2.0).
- Downloads last month
- -
Model tree for wesjos/embeddinggemma-2-openvino
Base model
google/embeddinggemma-2