EmbeddingGemma 2 Edge — Quantization
Quantized EmbeddingGemma 2 text towers that query an existing BF16 index without re-embedding the corpus.
Feature Extraction • 0.1B • Updated • 45Note 4-bit text tower, 152,669,232 B of weights. Keeps 0.936 of the BF16 top-10 neighborhood (BF16 = 1.000); margin 0.1712 vs BF16 0.1817. Internally measured on a private frozen set; serving not validated.
ThakiCloud/eg2-text-hybrid-129m
Feature Extraction • 0.1B • Updated • 28Note Mixed-precision text tower (GPTQ 4-bit transformer, 2-bit vocabulary with 30% of rows at 4-bit), 129,532,796 B — 85% of eg2-text-q4. Beats uniform 3-bit on every held-out slice; within 25% of the 3→4-bit gap of the Q4 recipe. Internally measured; serving not validated.
ThakiCloud/eg2-text-hybrid-119m
Feature Extraction • 0.1B • Updated • 15Note Size-oriented mixed-precision text tower (GPTQ 4-bit transformer, 2-bit vocabulary, no row rescue), 119,152,312 B. At the uniform-3-bit byte budget it wins all 8 held-out cells over uniform 3-bit, and it gives up most of the 4-bit advantage on multilingual and rare-identifier retrieval. Internally measured; serving not validated.
Where Post-Training Quantization Breaks Text Embedders: A Measured Map Across Four Embedder Families
Paper • 2609.16391 • PublishedNote Why this line allocates bits the way it does. The embedding table never emerges as the dominant isolated protection priority in any of the four families measured, and at INT4/g16 the spread between modules is too small to allocate against. Measures EmbeddingGemma-300M, not EmbeddingGemma 2.
Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders
Paper • 2610.09227 • PublishedNote The label-free drift signal this line's mixed-precision allocation descends from, measured on five development embedders including EmbeddingGemma-300M. EmbeddingGemma 2 itself is not in that paper's grid.