EmbeddingGemma 2 โ€” OpenVINO IR (fp16)

i made this repo because this model run faster on my cpu-only server than llama.cpp;

google/embeddinggemma-2 converted to OpenVINO IR, three towers, pure CPU inference.

Tower Files Size Output
Text openvino/model.xml + .bin 545 MB sentence_embedding [b, 768]
Vision openvino/vision_model.xml + .bin 337 MB image_features [n, 512]
Audio openvino/audio_model.xml + .bin 588 MB audio_features [n, 512]

The .xml holds the topology and the .bin the weights; keep them side by side with a matching name prefix.

Usage

pip install openvino tokenizers numpy pillow
python demo.py
dog.jpeg x captions:
  a dog                            0.725
  a panda                          0.576
  a penguin using ubuntu computer  0.505
  • Outputs are L2-normalized, so a dot product is the cosine similarity.
  • Retrieval prefixes: queries take task: search result | query: , documents take title: none | text: .
  • Images: vision tower produces [N, 512] soft tokens, which the text tower fuses. Put <|image|> (258880) placeholders in input_ids and pass the features via image_features. Audio works the same way with <|audio|> (258881); video frames each go through the vision tower.
  • Image preprocessing (aspect-ratio resize to a multiple of 3*16, 16x16 patchify, pad to 280*3ยฒ=2520 patches with pixel_position_ids = -1) is inlined in demo.py and only needs numpy and PIL.

Speed on 8 CPU threads: 61 ms for a short query, 172 ms for a long document, 37 ms/item at batch 8, 2.6 s for one image, 0.49 s for 11 s of audio.

Conversion

From onnx-community/embeddinggemma-2-ONNX (Apache-2.0) to OpenVINO 2026.x fp16, see tools/convert_ov.py and tools/convert_mm.py.

The upstream export keeps contrib ops that only the ORT kernel implements correctly. The text tower's com.microsoft.RotaryEmbedding and the vision tower's com.microsoft.MultiHeadAttention are replaced with standard-domain explicit ops, which removes the cross-engine semantic gap.

Against the original ONNX baseline, per-token cosine is 1.000000 for text, >=0.9946 for vision and 1.000000 for audio. Cross-modal retrieval matches the official transformers.js numbers (image to "cats" 0.746, audio to "speech" 0.772).

NPU and GPU

Measured on Meteor Lake (Intel Arc iGPU + AI Boost NPU):

  • NPU: the text tower works, but shapes must be static, batch must be 1, and the trailing L2 normalization is dropped so it has to be reapplied by hand. The vision tower fails to compile (VPUX pass error). No net speedup on short texts.
  • GPU: enumerates 0 Level Zero devices in a containerized environment; run on the host instead.

Unofficial conversion; see this repo's acceptance numbers. Original model by Google (Apache-2.0).

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for wesjos/embeddinggemma-2-openvino

Finetuned
(34)
this model