EmbeddingGemma 2 โ text-only ONNX (for Vespa)
Text-only ONNX export of google/embeddinggemma-2,
derived from onnx-community/embeddinggemma-2-ONNX
for use with Vespa's hugging-face-embedder.
The upstream graph requires image_features, video_features and audio_features inputs that are
concatenated after the text tokens. These are replaced with empty [0, 512] constants, so the
graph takes only input_ids and attention_mask. Weights are unchanged. See scripts/make_text_only.py.
| File | Precision | Size |
|---|---|---|
onnx/model.onnx (+ model.onnx_data) |
fp32 | 1.08 GB |
onnx/int8/model.onnx (+ model.onnx_data) |
int8 (dynamic quantization) | 314 MB |
Outputs: last_hidden_state [batch, seq, 768] (use mean pooling + L2 normalize) and sentence_embedding [batch, 768].
Do not use fp16. Per the model card, activations overflow float16.
Verification
scripts/verify.py compares Vespa-style inference (tokenizer.json with special tokens, mean pooling over
last_hidden_state, L2 normalization) against sentence-transformers 6.1 / transformers 5.19 in fp32,
on 11 texts including multilingual, code, and a 2,718-token document (exercises sliding-window attention).
Token IDs are identical to the reference.
| Variant | Cosine vs. reference (min, 768d / 512d / 256d / 128d) | Same ranking |
|---|---|---|
| fp32 | 1.000000 / 1.000000 / 1.000000 / 1.000000 (max abs diff 6e-7) | yes |
| int8 | 0.999896 / 0.999901 / 0.999910 / 0.999940 | yes |
Vespa usage
<component id="embeddinggemma2" type="hugging-face-embedder">
<transformer-model url="https://huggingface.co/vespa-engine/embeddinggemma-2-ONNX/resolve/main/onnx/model.onnx"/>
<tokenizer-model url="https://huggingface.co/vespa-engine/embeddinggemma-2-ONNX/resolve/main/tokenizer.json"/>
<max-tokens>8192</max-tokens>
<pooling-strategy>mean</pooling-strategy>
<normalize>true</normalize>
<prepend>
<query>task: search result | query: </query>
<document>title: none | text: </document>
</prepend>
</component>
Supports Matryoshka truncation to 512, 256 and 128 dimensions (re-normalize after truncating).
Model tree for vespa-engine/embeddinggemma-2-ONNX
Base model
google/embeddinggemma-2