Instructions to use lethalbeats/embeddinggemma-2-0.8bpw-text with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use lethalbeats/embeddinggemma-2-0.8bpw-text with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("lethalbeats/embeddinggemma-2-0.8bpw-text") sentences = [ "That is a happy person", "That is a happy dog", "That is a very happy person", "Today is a sunny day" ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
EmbeddingGemma-2 0.8 BPW Text (LittleBit Quantized)
EmbeddingGemma-2 0.8 BPW Text is an ultra-compact, sub-1-bit quantized runtime derivative of google/embeddinggemma-2 developed for extreme memory efficiency in dense representation, semantic search, clustering, and sentence similarity.
By leveraging the LittleBit sub-1-bit quantization framework introduced by Samsung Research in [2506.13771] LittleBit: Ultra Low-Bit Quantization via Latent Factorization (HTML version) presented at ICML, and applying an asymmetric 65/35 latent capacity allocation directly over the continuous weights of google/embeddinggemma-2, this model compresses the encoder weights down to 0.8 Bits-Per-Weight (BPW).
This slashes the runtime model footprint in RAM from ~1,500 MB (Original BF16/FP32 base) down to 274.5 MB (an 81.7% reduction in model memory, or $\sim 5.5\times$ smaller).
Base Model & Derivative Notice
- Base Model:
google/embeddinggemma-2 - Modality: Text Only
- Original Model Developers: Google DeepMind
- Quantized by: lethalbeats (LethalBeats)
- Distribution Notice: This repository is a community-contributed quantized runtime artifact and is not an official Google distribution. It is subject to the Gemma Terms of Use.
Memory Footprint & Runtime Execution
Model weights reside in RAM strictly in packed bitstream format (274.5 MB). Unpacking occurs on-the-fly inside CPU SIMD registers (ymm0, ymm1 in AVX2) in chunks of 8/16 elements per clock cycle with $< 64\text{ KB}$ working cache overhead.
| Model Configuration | Quantization Format | Model Weights RAM | Reduction vs Base |
|---|---|---|---|
| EmbeddingGemma-2 Base | BF16 / FP32 | ~1,500 MB | Baseline (0%) |
| EmbeddingGemma-2 ONNX | INT8 | 705.7 MB | -53.0% |
| EmbeddingGemma-2 0.8 BPW (This Model) | LittleBit (0.8 BPW) | 274.5 MB | -81.7% |
Note: Streaming in-register dequantization prevents in-memory RAM spikes, allowing the model to operate continuously inside low-memory instances and edge devices.
Quickstart & Usage
1. Python Inference via Streaming Loader
import numpy as np
from modeling_littlebit import EmbeddingGemma2LittleBit
# Load model from local directory or Hugging Face Hub
model = EmbeddingGemma2LittleBit.from_pretrained("lethalbeats/embeddinggemma-2-0.8bpw-text")
sentences = [
"What causes the northern lights?",
"The northern lights are caused by charged particles from the sun."
]
# Generate normalized 768-dimensional embeddings
embeddings = model.encode(sentences)
print("Embeddings shape:", embeddings.shape)
similarity = float(embeddings[0] @ embeddings[1])
print(f"Cosine Similarity: {similarity:.4f}")
2. Sentence Similarity & Semantic Search
# Compute pairwise similarity
query = model.encode(["What causes the northern lights?"])
corpus = model.encode([
"The northern lights are caused by charged particles from the sun.",
"Photosynthesis is a process used by plants and other organisms to convert light energy into chemical energy."
])
scores = (corpus @ query.T).flatten()
best_match_idx = int(np.argmax(scores))
print(f"Top Match Index: {best_match_idx}, Score: {scores[best_match_idx]:.4f}")
3. Matryoshka Dimension Slicing (MRL)
EmbeddingGemma-2 natively supports Matryoshka Representation Learning (MRL). You can slice embeddings to 512 or 256 dimensions and re-normalize for additional storage savings in vector databases:
# Truncate to 256 dimensions and re-normalize L2
embeddings_256 = embeddings[:, :256]
embeddings_256 = embeddings_256 / np.linalg.norm(embeddings_256, axis=-1, keepdims=True)
print("Sliced embeddings shape:", embeddings_256.shape)
Model Specifications
| Parameter | Value |
|---|---|
| Base Architecture | Gemma 2 Multimodal Text Transformer |
| Base Model | google/embeddinggemma-2 |
| Output Dimension | 768 (supports MRL slicing to 512, 256) |
| Quantization Method | LittleBit 0.8 BPW |
| Capacity Distribution | Asymmetric 65% Primary / 35% Secondary |
| Model Weights RAM Footprint | 274.5 MB |
| Max Sequence Length | 8,192 tokens |
| Similarity Function | Cosine Similarity (Dot product on unit $\mathbb{S}^{767}$) |
Theoretical Foundation: LittleBit (Samsung Research / ICML)
This model's quantization relies directly on the breakthroughs introduced by Samsung Research in the paper:
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
Samsung Research (ICML)
arXiv:2506.13771 | HTML Full Paper (v5)
Core Theoretical Pillars from LittleBit:
- Low-Rank Latent Factorization: Decomposes continuous weight matrices $W \in \mathbb{R}^{d_{out} \times d_{in}}$ into low-rank latent factors $U$ and $V$, enabling stable binarization down to sub-1-bit budgets (0.1 to 1.0 BPW).
- Multi-Scale Compensation: Learns importance parameters across row, column, and latent dimensions to prevent catastrophic representation collapse in deep Transformer blocks.
- Dual-SVID (Sign-Value-Independent Decomposition): Decouples structural orientation from scale, facilitating robust Quantization-Aware Training (QAT) and knowledge distillation.
- Asymmetric 65/35 Allocation: Dedicates 65% of bit budget capacity to dominant semantic subspace factors and 35% to secondary residual compensation.
Citations & References
If you use this model in your research or applications, please cite:
@article{littlebit_2025,
title={LittleBit: Ultra Low-Bit Quantization via Latent Factorization},
author={Samsung Research},
journal={arXiv preprint arXiv:2506.13771},
year={2025},
url={https://arxiv.org/abs/2506.13771}
}
@article{embedding_gemma_2025,
title={EmbeddingGemma: Powerful and Lightweight Text Representations},
author={Schechter Vera, Henrique and Dua, Sahil and Zhang, Biao and Salz, Daniel and Mullins, Ryan and Raghuram Panyam, Sindhu and Smoot, Sara and Naim, Iftekhar and Zou, Joe and Chen, Feiyang and Cer, Daniel and Lisak, Alice and Choi, Min and Gonzalez, Lucas and Sanseviero, Omar and Cameron, Glenn and Ballantyne, Ian and Black, Kat and Chen, Kaifeng and Wang, Weiyi and Li, Zhe and Martin, Scott},
journal={arXiv preprint},
year={2025}
}
@misc{embeddinggemma2_google,
title={EmbeddingGemma-2: Multimodal Representation Models},
author={Google DeepMind},
year={2025},
publisher={Hugging Face},
howpublished={\url{https://huggingface.co/google/embeddinggemma-2}}
}
License
This model inherits the Gemma Terms of Use from Google.
- Downloads last month
- -
Model tree for lethalbeats/embeddinggemma-2-0.8bpw-text
Base model
google/embeddinggemma-2