Instructions to use opensearch-project/opensearch-neural-sparse-encoding-v2-distill with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use opensearch-project/opensearch-neural-sparse-encoding-v2-distill with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("opensearch-project/opensearch-neural-sparse-encoding-v2-distill") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Transformers
How to use opensearch-project/opensearch-neural-sparse-encoding-v2-distill with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="opensearch-project/opensearch-neural-sparse-encoding-v2-distill")# Load model directly from transformers import AutoTokenizer, AutoModelForMaskedLM tokenizer = AutoTokenizer.from_pretrained("opensearch-project/opensearch-neural-sparse-encoding-v2-distill") model = AutoModelForMaskedLM.from_pretrained("opensearch-project/opensearch-neural-sparse-encoding-v2-distill", device_map="auto") - Inference
- Notebooks
- Google Colab
- Kaggle
How did you evaluate BEIR?
I am trying to reproduce your BEIR results. Some results match exactly but others are lower than the numbers in your model card. What tool did you use to evaluate BEIR? I am currently using pyserini & the beir package.
Hi @freethenation , we are using OpenSearch as the evaluate engine. The max input length is 512 tokens. Please note that for some BEIR dataset, we need to filter out the query id from the search results, because for these datasets queries and documents are from the same space.
Evaluation code is available here: https://github.com/zhichao-aws/opensearch-sparse-model-tuning-sample/tree/main