Feature Extraction
sentence-transformers
Safetensors
Transformers
embedding_gemma2
sentence-similarity
autoround
4-bit precision
quantization
embeddinggemma
embedding
mrl
matryoshka
auto-round
Instructions to use webmp3/Sakura-EmbeddingGemma-2-AutoRound with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use webmp3/Sakura-EmbeddingGemma-2-AutoRound with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("webmp3/Sakura-EmbeddingGemma-2-AutoRound") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Transformers
How to use webmp3/Sakura-EmbeddingGemma-2-AutoRound with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="webmp3/Sakura-EmbeddingGemma-2-AutoRound")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModel processor = AutoProcessor.from_pretrained("webmp3/Sakura-EmbeddingGemma-2-AutoRound") model = AutoModel.from_pretrained("webmp3/Sakura-EmbeddingGemma-2-AutoRound", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Download quantization_config.json from webmp3/Sakura-EmbeddingGemma-2-AutoRound: direct link, hf CLI and curl.
- Browser
- Download file 291 Bytes
-
https://huggingface.co/webmp3/Sakura-EmbeddingGemma-2-AutoRound/resolve/main/quantization_config.json
- Command line
-
hf download hf://webmp3/Sakura-EmbeddingGemma-2-AutoRound/quantization_config.json
-
curl -L -o quantization_config.json https://huggingface.co/webmp3/Sakura-EmbeddingGemma-2-AutoRound/resolve/main/quantization_config.json
291 Bytes
| { | |
| "bits": 4, | |
| "data_type": "int", | |
| "group_size": 128, | |
| "sym": true, | |
| "iters": 100, | |
| "low_gpu_mem_usage": true, | |
| "autoround_version": "0.16.0", | |
| "block_name_to_quantize": "language_model.layers", | |
| "quant_method": "auto-round", | |
| "packing_format": "auto_round:auto_gptq" | |
| } |