EmbeddingGemma-2 0.8 BPW Vision (LittleBit Quantized)

EmbeddingGemma-2 0.8 BPW Vision is an ultra-compact, sub-1-bit quantized runtime derivative of google/embeddinggemma-2 developed for extreme memory efficiency in cross-modal image-text retrieval, video moment retrieval, visual semantic search, and zero-shot image classification.

Following Google's modular architecture, this model packages the 270M Text Backbone and 170M Vision Encoder (440M total parameters), with the Vision Encoder natively handling both images and video (sampled at 1 fps). By applying Samsung Research's LittleBit sub-1-bit quantization framework ([arXiv:2506.13771]) with an asymmetric 65/35 latent factorization, encoder weights are compressed to 0.8 Bits-Per-Weight (BPW).

This slashes runtime RAM from ~1,760 MB (BF16/FP32 base) down to ~429.5 MB (a 75.6% reduction in model memory, or $\sim 4.1\times$ smaller).

Base Model & Derivative Notice

  • Base Model: google/embeddinggemma-2
  • Modality: Text + Vision + Video (Gemma 2 Text Transformer + SigLIP-style Vision Transformer)
  • Original Model Developers: Google DeepMind
  • Quantized by: lethalbeats (LethalBeats)
  • Distribution Notice: This repository is a community-contributed quantized runtime artifact and is not an official Google distribution. It is subject to the Gemma Terms of Use.

Memory Footprint & Runtime Execution

Model weights reside in RAM strictly in packed bitstream format (~429.5 MB). Unpacking occurs on-the-fly inside CPU SIMD registers (ymm0, ymm1 in AVX2) in chunks of 8/16 elements per clock cycle with $< 64\text{ KB}$ working cache overhead.

Model Configuration Active Modalities Model Weights RAM Reduction vs Base
EmbeddingGemma-2 (Text + Vision Base) Text, Image, Video ~1,760 MB Baseline (0%)
EmbeddingGemma-2 0.8 BPW Vision (This Model) Text, Image, Video ~429.5 MB -75.6%

Note: In-register streaming dequantization prevents in-memory RAM spikes, allowing multimodal image and video retrieval inside memory-constrained edge hardware and CPU servers.


Quickstart & Usage

Adheres to Google's official SentenceTransformers dictionary input convention:

import numpy as np
from PIL import Image
from modeling_littlebit import EmbeddingGemma2VisionLittleBit

# Load model from local directory or Hugging Face Hub
model = EmbeddingGemma2VisionLittleBit.from_pretrained("lethalbeats/embeddinggemma-2-0.8bpw-vision")

# 1. Text Query
query_emb = model.encode("What causes the northern lights?")

# 2. Image Embedding (accepts file path or PIL.Image)
image_emb = model.encode({"image": "aurora_borealis.jpg"})

# 3. Video Embedding (sampled at 1 fps via vision encoder)
video_emb = model.encode({"video": "aurora_timelapse.mp4"})

# Cross-modal cosine similarity
similarity_img = float(query_emb[0] @ image_emb[0])
similarity_vid = float(query_emb[0] @ video_emb[0])

print(f"Similarity (Text vs Image): {similarity_img:.4f}")
print(f"Similarity (Text vs Video): {similarity_vid:.4f}")

Matryoshka Dimension Slicing (MRL)

EmbeddingGemma-2 natively supports Matryoshka Representation Learning (MRL). You can slice image and text embeddings to 512 or 256 dimensions and re-normalize for additional storage savings in vector databases:

# Truncate to 256 dimensions and re-normalize L2
emb_256 = query_emb[:, :256]
emb_256 = emb_256 / np.linalg.norm(emb_256, axis=-1, keepdims=True)
print("Sliced embedding shape:", emb_256.shape)

Model Specifications

Parameter Value
Base Architecture Gemma 2 Multimodal Text & Vision Transformer
Base Model google/embeddinggemma-2
Supported Modalities Text, Image, Video
Output Dimension 768 (supports MRL slicing to 512, 256)
Quantization Method LittleBit 0.8 BPW
Capacity Distribution Asymmetric 65% Primary / 35% Secondary
Model Weights RAM Footprint ~429.5 MB
Max Sequence Length 8,192 tokens
Video Processing Sampled at 1 fps via Vision Encoder
Similarity Function Cosine Similarity (Dot product on unit $\mathbb{S}^{767}$)

Theoretical Foundation: LittleBit (Samsung Research / ICML)

This model's quantization relies directly on the breakthroughs introduced by Samsung Research in the paper:

LittleBit: Ultra Low-Bit Quantization via Latent Factorization
Samsung Research (ICML)
arXiv:2506.13771 | HTML Full Paper (v5)


Citations & References

If you use this model in your research or applications, please cite:

@article{littlebit_2025,
  title={LittleBit: Ultra Low-Bit Quantization via Latent Factorization},
  author={Samsung Research},
  journal={arXiv preprint arXiv:2506.13771},
  year={2025},
  url={https://arxiv.org/abs/2506.13771}
}

@article{embedding_gemma_2025,
  title={EmbeddingGemma: Powerful and Lightweight Text Representations},
  author={Schechter Vera, Henrique and Dua, Sahil and Zhang, Biao and Salz, Daniel and Mullins, Ryan and Raghuram Panyam, Sindhu and Smoot, Sara and Naim, Iftekhar and Zou, Joe and Chen, Feiyang and Cer, Daniel and Lisak, Alice and Choi, Min and Gonzalez, Lucas and Sanseviero, Omar and Cameron, Glenn and Ballantyne, Ian and Black, Kat and Chen, Kaifeng and Wang, Weiyi and Li, Zhe and Martin, Scott},
  journal={arXiv preprint},
  year={2025}
}

@misc{embeddinggemma2_google,
  title={EmbeddingGemma-2: Multimodal Representation Models},
  author={Google DeepMind},
  year={2025},
  publisher={Hugging Face},
  howpublished={\url{https://huggingface.co/google/embeddinggemma-2}}
}

License

This model inherits the Gemma Terms of Use from Google.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for lethalbeats/embeddinggemma-2-0.8bpw-vision

Quantized
(56)
this model

Paper for lethalbeats/embeddinggemma-2-0.8bpw-vision