Google EmbeddingGemma 2 audio encoder exported to ONNX.

Filename Quantization File size Output
audio_embedding.onnx (4.24 MB)
audio_embedding.data (2.31 GB)
FP32 2.31 GB total Final embedding; weights are split across both files
audio_embedding_int8.onnx INT8 dynamic 583.1 MB Final embedding
audio_embedding_int4.onnx INT4 weight-only 777.1 MB Final embedding
audio_tower.onnx FP32 1.22 GB Intermediate states
audio_tower_int8.onnx INT8 dynamic 307.7 MB Intermediate states
audio_tower_int4.onnx INT4 weight-only 164.5 MB Intermediate states

Input and output

  • Input: input_ids and attention_mask int64 [1, text_tokens]; input_features float32 [1, feature_frames, 128]; input_features_mask bool [1, feature_frames]. From the pinned AutoProcessor for mono 16 kHz audio; decoding and preprocessing are external.
  • Output: audio_embedding*.onnx returns an L2-normalized embedding, float32 [1, 512]. audio_tower*.onnx returns intermediate audio states.
  • Runtime: Opset 18; INT4 requires ONNX Runtime support for com.microsoft::MatMulNBits.

SHA-256 hashes

Filename SHA-256
audio_embedding.onnx 974b4fdbbc0b23fbd63a7ec06c36c3aa405d2563f65f1fdffa45e953a88a32a3
audio_embedding.data a9e6f1aee2489d02eb7d4953e9bc2bb2e03cc557bf4f01654f1af9dcf5d7d4dd
audio_embedding_int8.onnx 41441cf3cccd9c755ca7efbdb6b69c726632a516d200d8230daa030e9fb840bd
audio_embedding_int4.onnx d4c269c7a024d1d2d495d121a2b608c4612833fa52e5eda5d7a9345ee493a3fe
audio_tower.onnx 08f2eb80d794d4cdef5e8d16224ac1685893cad2e6cb7e68d8bed614a7c3a023
audio_tower_int8.onnx 0fee17fc7ba12788122634d1780ae15bcce48b6c2e3943e47d98e92b8cace6ae
audio_tower_int4.onnx b0a3d2e2a0f3177322bab1947577c2e7abd408e063e3c6d7619f8aef8a96eded
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for bdrgon/embeddinggemma2-audio-tower-onnx

Quantized
(62)
this model