technotic commited on
Commit
949a07a
·
verified ·
1 Parent(s): 5287ad8

Upload folder using huggingface_hub

Browse files
Files changed (3) hide show
  1. .gitattributes +1 -0
  2. README.md +70 -0
  3. embeddinggemma-300M.litertlm +3 -0
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ embeddinggemma-300M.litertlm filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,70 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: gemma
3
+ pipeline_tag: sentence-similarity
4
+ library_name: litert
5
+ tags:
6
+ - sentence-transformers
7
+ - sentence-similarity
8
+ - feature-extraction
9
+ - text-embeddings-inference
10
+ - litert
11
+ - litertlm
12
+ - quantized
13
+ - int8
14
+ extra_gated_heading: Access EmbeddingGemma on Hugging Face
15
+ extra_gated_prompt: To access EmbeddingGemma on Hugging Face, you’re required to review and
16
+ agree to Google’s usage license. To do this, please ensure you’re logged in to Hugging
17
+ Face and click below. Requests are processed immediately.
18
+ extra_gated_button_content: Acknowledge license
19
+ ---
20
+
21
+ # EmbeddingGemma-300M (INT8 LiteRT-LM)
22
+
23
+ > **Note:** This repository contains a **quantized INT8 LiteRT-LM (`litertlm`)** build of **EmbeddingGemma-300M** created by **technotic** for efficient on-device inference using Google LiteRT / LiteRT-LM runtime.
24
+
25
+ ---
26
+
27
+ # EmbeddingGemma model card
28
+
29
+ **Model Page**: [EmbeddingGemma](https://ai.google.dev/gemma/docs/embeddinggemma)
30
+
31
+ **Resources and Technical Documentation**:
32
+
33
+ * [Responsible Generative AI Toolkit](https://ai.google.dev/responsible)
34
+ * [EmbeddingGemma on Kaggle](https://www.kaggle.com/models/google/embeddinggemma/)
35
+ * [EmbeddingGemma on Vertex Model Garden](https://console.cloud.google.com/vertex-ai/publishers/google/model-garden/embeddinggemma)
36
+
37
+ **Terms of Use**: [Terms](https://ai.google.dev/gemma/terms)
38
+
39
+ **Authors**: Google DeepMind
40
+
41
+ ## Model Information
42
+
43
+ ### Description
44
+
45
+ EmbeddingGemma is a 300M parameter, state-of-the-art for its size, open embedding model from Google, built from Gemma 3 (with T5Gemma initialization) and the same research and technology used to create Gemini models. EmbeddingGemma produces vector representations of text, making it well-suited for search and retrieval tasks, including classification, clustering, and semantic similarity search. This model was trained with data in 100+ spoken languages.
46
+
47
+ The small size and on-device focus makes it possible to deploy in environments with limited resources such as mobile phones, laptops, or desktops, democratizing access to state of the art AI models and helping foster innovation for everyone.
48
+
49
+ For more technical details, refer to our paper: [EmbeddingGemma: Powerful and Lightweight Text Representations](https://arxiv.org/abs/2509.20354).
50
+
51
+ ### Inputs and outputs
52
+
53
+ - **Input:**
54
+ - Text string, such as a question, a prompt, or a document to be embedded
55
+ - Maximum input context length of 2048 tokens
56
+
57
+ - **Output:**
58
+ - Numerical vector representations of input text data
59
+ - Output embedding dimension size of 768, with smaller options available (512, 256, or 128) via Matryoshka Representation Learning (MRL). MRL allows users to truncate the output embedding of size 768 to their desired size and then re-normalize for efficient and accurate representation.
60
+
61
+ ### Citation
62
+
63
+ ```none
64
+ @article{embedding_gemma_2025,
65
+ title={EmbeddingGemma: Powerful and Lightweight Text Representations},
66
+ author={Schechter Vera, Henrique* and Dua, Sahil* and Zhang, Biao and Salz, Daniel and Mullins, Ryan and Raghuram Panyam, Sindhu and Smoot, Sara and Naim, Iftekhar and Zou, Joe and Chen, Feiyang and Cer, Daniel and Lisak, Alice and Choi, Min and Gonzalez, Lucas and Sanseviero, Omar and Cameron, Glenn and Ballantyne, Ian and Black, Kat and Chen, Kaifeng and Wang, Weiyi and Li, Zhe and Martins, Gus and Lee, Jinhyuk and Sherwood, Mark and Ji, Juyeong and Wu, Renjie and Zheng, Jingxiao and Singh, Jyotinder and Sharma, Abheesht and Sreepat, Divya and Jain, Aashi and Elarabawy, Adham and Co, AJ and Doumanoglou, Andreas and Samari, Babak and Hora, Ben and Potetz, Brian and Kim, Dahun and Alfonseca, Enrique and Moiseev, Fedor and Han, Feng and Palma Gomez, Frank and Hernández Ábrego, Gustavo and Zhang, Hesen and Hui, Hui and Han, Jay and Gill, Karan and Chen, Ke and Chen, Koert and Shanbhogue, Madhuri and Boratko, Michael and Suganthan, Paul and Duddu, Sai Meher Karthik and Mariserla, Sandeep and Ariafar, Setareh and Zhang, Shanfeng and Zhang, Shijie and Baumgartner, Simon and Goenka, Sonam and Qiu, Steve and Dabral, Tanmaya and Walker, Trevor and Rao, Vikram and Khawaja, Waleed and Zhou, Wenlei and Ren, Xiaoqi and Xia, Ye and Chen, Yichang and Chen, Yi-Ting and Dong, Zhe and Ding, Zhongli and Visin, Francesco and Liu, Gaël and Zhang, Jiageng and Kenealy, Kathleen and Casbon, Michelle and Kumar, Ravin and Mesnard, Thomas and Gleicher, Zach and Brick, Cormac and Lacombe, Olivier and Roberts, Adam and Sung, Yunhsuan and Hoffmann, Raphael and Warkentin, Tris and Joulin, Armand and Duerig, Tom and Seyedhosseini, Mojtaba},
67
+ publisher={Google DeepMind},
68
+ year={2025},
69
+ url={[https://arxiv.org/abs/2509.20354](https://arxiv.org/abs/2509.20354)}
70
+ }
embeddinggemma-300M.litertlm ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:cc9b0e9c3fb3c4f3109f16ead0a72372a6515f5d1fada8b2facb995797e49323
3
+ size 525899056