EmbeddingGemma-2 0.8 BPW Text (LittleBit Quantized)

EmbeddingGemma-2 0.8 BPW Text is an ultra-compact, sub-1-bit quantized runtime derivative of google/embeddinggemma-2 developed for extreme memory efficiency in dense representation, semantic search, clustering, and sentence similarity.

By leveraging the LittleBit sub-1-bit quantization framework introduced by Samsung Research in [2506.13771] LittleBit: Ultra Low-Bit Quantization via Latent Factorization (HTML version) presented at ICML, and applying an asymmetric 65/35 latent capacity allocation directly over the continuous weights of google/embeddinggemma-2, this model compresses the encoder weights down to 0.8 Bits-Per-Weight (BPW).

This slashes the runtime model footprint in RAM from ~1,500 MB (Original BF16/FP32 base) down to 274.5 MB (an 81.7% reduction in model memory, or $\sim 5.5\times$ smaller).

Base Model & Derivative Notice


Memory Footprint & Runtime Execution

Model weights reside in RAM strictly in packed bitstream format (274.5 MB). Unpacking occurs on-the-fly inside CPU SIMD registers (ymm0, ymm1 in AVX2) in chunks of 8/16 elements per clock cycle with $< 64\text{ KB}$ working cache overhead.

Model Configuration Quantization Format Model Weights RAM Reduction vs Base
EmbeddingGemma-2 Base BF16 / FP32 ~1,500 MB Baseline (0%)
EmbeddingGemma-2 ONNX INT8 705.7 MB -53.0%
EmbeddingGemma-2 0.8 BPW (This Model) LittleBit (0.8 BPW) 274.5 MB -81.7%

Note: Streaming in-register dequantization prevents in-memory RAM spikes, allowing the model to operate continuously inside low-memory instances and edge devices.


Quickstart & Usage

1. Python Inference via Streaming Loader

import numpy as np
from modeling_littlebit import EmbeddingGemma2LittleBit

# Load model from local directory or Hugging Face Hub
model = EmbeddingGemma2LittleBit.from_pretrained("lethalbeats/embeddinggemma-2-0.8bpw-text")

sentences = [
    "What causes the northern lights?",
    "The northern lights are caused by charged particles from the sun."
]

# Generate normalized 768-dimensional embeddings
embeddings = model.encode(sentences)

print("Embeddings shape:", embeddings.shape)
similarity = float(embeddings[0] @ embeddings[1])
print(f"Cosine Similarity: {similarity:.4f}")

2. Sentence Similarity & Semantic Search

# Compute pairwise similarity
query = model.encode(["What causes the northern lights?"])
corpus = model.encode([
    "The northern lights are caused by charged particles from the sun.",
    "Photosynthesis is a process used by plants and other organisms to convert light energy into chemical energy."
])

scores = (corpus @ query.T).flatten()
best_match_idx = int(np.argmax(scores))
print(f"Top Match Index: {best_match_idx}, Score: {scores[best_match_idx]:.4f}")

3. Matryoshka Dimension Slicing (MRL)

EmbeddingGemma-2 natively supports Matryoshka Representation Learning (MRL). You can slice embeddings to 512 or 256 dimensions and re-normalize for additional storage savings in vector databases:

# Truncate to 256 dimensions and re-normalize L2
embeddings_256 = embeddings[:, :256]
embeddings_256 = embeddings_256 / np.linalg.norm(embeddings_256, axis=-1, keepdims=True)
print("Sliced embeddings shape:", embeddings_256.shape)

Model Specifications

Parameter Value
Base Architecture Gemma 2 Multimodal Text Transformer
Base Model google/embeddinggemma-2
Output Dimension 768 (supports MRL slicing to 512, 256)
Quantization Method LittleBit 0.8 BPW
Capacity Distribution Asymmetric 65% Primary / 35% Secondary
Model Weights RAM Footprint 274.5 MB
Max Sequence Length 8,192 tokens
Similarity Function Cosine Similarity (Dot product on unit $\mathbb{S}^{767}$)

Theoretical Foundation: LittleBit (Samsung Research / ICML)

This model's quantization relies directly on the breakthroughs introduced by Samsung Research in the paper:

LittleBit: Ultra Low-Bit Quantization via Latent Factorization
Samsung Research (ICML)
arXiv:2506.13771 | HTML Full Paper (v5)

Core Theoretical Pillars from LittleBit:

  1. Low-Rank Latent Factorization: Decomposes continuous weight matrices $W \in \mathbb{R}^{d_{out} \times d_{in}}$ into low-rank latent factors $U$ and $V$, enabling stable binarization down to sub-1-bit budgets (0.1 to 1.0 BPW).
  2. Multi-Scale Compensation: Learns importance parameters across row, column, and latent dimensions to prevent catastrophic representation collapse in deep Transformer blocks.
  3. Dual-SVID (Sign-Value-Independent Decomposition): Decouples structural orientation from scale, facilitating robust Quantization-Aware Training (QAT) and knowledge distillation.
  4. Asymmetric 65/35 Allocation: Dedicates 65% of bit budget capacity to dominant semantic subspace factors and 35% to secondary residual compensation.

Citations & References

If you use this model in your research or applications, please cite:

@article{littlebit_2025,
  title={LittleBit: Ultra Low-Bit Quantization via Latent Factorization},
  author={Samsung Research},
  journal={arXiv preprint arXiv:2506.13771},
  year={2025},
  url={https://arxiv.org/abs/2506.13771}
}

@article{embedding_gemma_2025,
  title={EmbeddingGemma: Powerful and Lightweight Text Representations},
  author={Schechter Vera, Henrique and Dua, Sahil and Zhang, Biao and Salz, Daniel and Mullins, Ryan and Raghuram Panyam, Sindhu and Smoot, Sara and Naim, Iftekhar and Zou, Joe and Chen, Feiyang and Cer, Daniel and Lisak, Alice and Choi, Min and Gonzalez, Lucas and Sanseviero, Omar and Cameron, Glenn and Ballantyne, Ian and Black, Kat and Chen, Kaifeng and Wang, Weiyi and Li, Zhe and Martin, Scott},
  journal={arXiv preprint},
  year={2025}
}

@misc{embeddinggemma2_google,
  title={EmbeddingGemma-2: Multimodal Representation Models},
  author={Google DeepMind},
  year={2025},
  publisher={Hugging Face},
  howpublished={\url{https://huggingface.co/google/embeddinggemma-2}}
}

License

This model inherits the Gemma Terms of Use from Google.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for lethalbeats/embeddinggemma-2-0.8bpw-text

Quantized
(56)
this model

Paper for lethalbeats/embeddinggemma-2-0.8bpw-text