Instructions to use lethalbeats/embeddinggemma-2-0.8bpw-audio with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use lethalbeats/embeddinggemma-2-0.8bpw-audio with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("lethalbeats/embeddinggemma-2-0.8bpw-audio") sentences = [ "That is a happy person", "That is a happy dog", "That is a very happy person", "Today is a sunny day" ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
EmbeddingGemma-2 0.8 BPW Audio (LittleBit Quantized)
EmbeddingGemma-2 0.8 BPW Audio is an ultra-compact, sub-1-bit quantized runtime derivative of google/embeddinggemma-2 developed for extreme memory efficiency in cross-modal audio-text retrieval, sound semantic search, audio classification, and speech embedding.
Following Google's modular architecture, this model packages the 270M Text Backbone and 300M Audio Encoder (570M total parameters), with the Audio Encoder natively processing mono audio at 16 kHz. By applying Samsung Research's LittleBit sub-1-bit quantization framework ([arXiv:2506.13771]) with an asymmetric 65/35 latent factorization, encoder weights are compressed to 0.8 Bits-Per-Weight (BPW).
This slashes runtime RAM from ~2,280 MB (BF16/FP32 base) down to ~549.0 MB (a 75.9% reduction in model memory, or $\sim 4.2\times$ smaller).
Base Model & Derivative Notice
- Base Model:
google/embeddinggemma-2 - Modality: Text + Audio (Audio Convolutional Transformer + Gemma 2 Text Transformer)
- Original Model Developers: Google DeepMind
- Quantized by: lethalbeats (LethalBeats)
- Distribution Notice: This repository is a community-contributed quantized runtime artifact and is not an official Google distribution. It is subject to the Gemma Terms of Use.
Memory Footprint & Runtime Execution
Model weights reside in RAM strictly in packed bitstream format (~549.0 MB). Unpacking occurs on-the-fly inside CPU SIMD registers (ymm0, ymm1 in AVX2) in chunks of 8/16 elements per clock cycle with $< 64\text{ KB}$ working cache overhead.
| Model Configuration | Active Modalities | Model Weights RAM | Reduction vs Base |
|---|---|---|---|
| EmbeddingGemma-2 (Text + Audio Base) | Text, Audio (16 kHz mono) | ~2,280 MB | Baseline (0%) |
| EmbeddingGemma-2 0.8 BPW Audio (This Model) | Text, Audio (16 kHz mono) | ~549.0 MB | -75.9% |
Note: In-register streaming dequantization prevents in-memory RAM spikes, allowing audio-text semantic retrieval inside memory-constrained edge hardware and voice-enabled appliances.
Quickstart & Usage
Adheres to Google's official SentenceTransformers dictionary input convention:
import numpy as np
from modeling_littlebit import EmbeddingGemma2AudioLittleBit
# Load model from local directory or Hugging Face Hub
model = EmbeddingGemma2AudioLittleBit.from_pretrained("lethalbeats/embeddinggemma-2-0.8bpw-audio")
# 1. Text search query
query_emb = model.encode("A recording of thunder during a heavy rainstorm.")
# 2. Audio embedding (accepts file path to .wav/.mp3 or 16 kHz mono float32 waveform array)
audio_emb = model.encode({"audio": "thunderstorm.wav"})
# Cross-modal cosine similarity
similarity = float(query_emb[0] @ audio_emb[0])
print(f"Cross-Modal Similarity (Text vs Audio): {similarity:.4f}")
Matryoshka Dimension Slicing (MRL)
EmbeddingGemma-2 natively supports Matryoshka Representation Learning (MRL). You can slice audio and text embeddings to 512 or 256 dimensions and re-normalize for additional storage savings in vector databases:
# Truncate to 256 dimensions and re-normalize L2
emb_256 = query_emb[:, :256]
emb_256 = emb_256 / np.linalg.norm(emb_256, axis=-1, keepdims=True)
print("Sliced embedding shape:", emb_256.shape)
Model Specifications
| Parameter | Value |
|---|---|
| Base Architecture | Gemma 2 Multimodal Text & Audio Transformer |
| Base Model | google/embeddinggemma-2 |
| Supported Modalities | Text, Audio (Waveform / Spectrogram) |
| Output Dimension | 768 (supports MRL slicing to 512, 256) |
| Quantization Method | LittleBit 0.8 BPW |
| Capacity Distribution | Asymmetric 65% Primary / 35% Secondary |
| Model Weights RAM Footprint | ~549.0 MB |
| Max Text Sequence Length | 8,192 tokens |
| Audio Sample Rate | 16,000 Hz mono |
| Similarity Function | Cosine Similarity (Dot product on unit $\mathbb{S}^{767}$) |
Theoretical Foundation: LittleBit (Samsung Research / ICML)
This model's quantization relies directly on the breakthroughs introduced by Samsung Research in the paper:
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
Samsung Research (ICML)
arXiv:2506.13771 | HTML Full Paper (v5)
Citations & References
If you use this model in your research or applications, please cite:
@article{littlebit_2025,
title={LittleBit: Ultra Low-Bit Quantization via Latent Factorization},
author={Samsung Research},
journal={arXiv preprint arXiv:2506.13771},
year={2025},
url={https://arxiv.org/abs/2506.13771}
}
@article{embedding_gemma_2025,
title={EmbeddingGemma: Powerful and Lightweight Text Representations},
author={Schechter Vera, Henrique and Dua, Sahil and Zhang, Biao and Salz, Daniel and Mullins, Ryan and Raghuram Panyam, Sindhu and Smoot, Sara and Naim, Iftekhar and Zou, Joe and Chen, Feiyang and Cer, Daniel and Lisak, Alice and Choi, Min and Gonzalez, Lucas and Sanseviero, Omar and Cameron, Glenn and Ballantyne, Ian and Black, Kat and Chen, Kaifeng and Wang, Weiyi and Li, Zhe and Martin, Scott},
journal={arXiv preprint},
year={2025}
}
@misc{embeddinggemma2_google,
title={EmbeddingGemma-2: Multimodal Representation Models},
author={Google DeepMind},
year={2025},
publisher={Hugging Face},
howpublished={\url{https://huggingface.co/google/embeddinggemma-2}}
}
License
This model inherits the Gemma Terms of Use from Google.
- Downloads last month
- -
Model tree for lethalbeats/embeddinggemma-2-0.8bpw-audio
Base model
google/embeddinggemma-2