Sentence Similarity
sentence-transformers
Safetensors
English
modernbert
colbert
late-interaction
retrieval
pylate
multi-vector
text-embeddings-inference
Instructions to use chungimungi/GLInt with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use chungimungi/GLInt with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("chungimungi/GLInt") sentences = [ "That is a happy person", "That is a happy dog", "That is a very happy person", "Today is a sunny day" ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
Correct BEIR best-value formatting
Browse files
README.md
CHANGED
|
@@ -21,6 +21,8 @@ Training has two stages:
|
|
| 21 |
`jinaai/jina-reranker-v3.5` scores, temperature sharpening, false-negative masking, and an
|
| 22 |
InfoNCE anchor.
|
| 23 |
|
|
|
|
|
|
|
| 24 |
## Usage
|
| 25 |
|
| 26 |
```python
|
|
@@ -43,14 +45,14 @@ computed by summing, over query tokens, the maximum similarity to a document tok
|
|
| 43 |
|
| 44 |
| Model | Average | Size (M) | Embed dim | ArguAna | CQADupstackRetrieval | ClimateFEVER | DBPedia | FEVER | FiQA2018 | HotpotQA | MSMARCO | NFCorpus | NQ | QuoraRetrieval | SCIDOCS | SciFact | TRECCOVID | Touche2020 |
|
| 45 |
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 46 |
-
| [ColBERTv2](https://huggingface.co/colbert-ir/colbertv2.0) | 48.63 | 110 | 128 | 46.50 |
|
| 47 |
-
| [Jina-ColBERT-v2](https://huggingface.co/jinaai/jina-colbert-v2) | 51.85 | 600 | 128 | 36.60 |
|
| 48 |
| [ColBERT-small](https://huggingface.co/answerdotai/answerai-colbert-small-v1) | 53.79 | 33 | 96 | 50.09 | 38.75 | 33.07 | 45.58 | 90.96 | 41.15 | 76.11 | 43.50 | 37.30 | 59.10 | 87.72 | 18.42 | 74.77 | 84.59 | 25.69 |
|
| 49 |
-
| [GTE-ModernColBERT-v1](https://huggingface.co/lightonai/GTE-ModernColBERT-v1) | 54.75 | 149 | 128 | 47.52 | 41.08 | 31.33 | 47.56 | 87.67 | 45.25 | 77.48 | 45.60 | 37.83 | 61.62 | 86.71 | 19.22 | 76.33 | **84.84** | 31.25 |
|
| 50 |
| [ColBERT-Zero](https://huggingface.co/lightonai/ColBERT-Zero) | 55.39 | 149 | 128 | **52.82** | 41.41 | 35.90 | 47.43 | 90.52 | 42.50 | 79.45 | 45.95 | 37.21 | 61.82 | 85.19 | 19.84 | 76.33 | 78.27 | **36.24** |
|
| 51 |
-
| [LateOn-unsupervised](https://huggingface.co/lightonai/LateOn-unsupervised) | 50.11 | 149 | 128 | 43.12 | **47.71** | 18.76 | 43.36 | 65.74 | 51.94 | 68.17 | 37.51 | 37.15 | 58.41 | 89.48 | 21.13 |
|
| 52 |
-
| [LateOn](https://huggingface.co/lightonai/LateOn) | 57.22 | 149 | 128 | 50.52 | 47.36 | **39.67** | 45.99 |
|
| 53 |
-
|
|
| 54 |
|
| 55 |
GLINT-base was evaluated with the project BEIR protocol: corpus IDs are excluded from their own
|
| 56 |
query results for ArguAna and Quora, and ArguAna uses 64 query tokens. The 57.43 average is one
|
|
|
|
| 21 |
`jinaai/jina-reranker-v3.5` scores, temperature sharpening, false-negative masking, and an
|
| 22 |
InfoNCE anchor.
|
| 23 |
|
| 24 |
+
This is a private research release. It is a single checkpoint, not an ensemble or a re-ranker.
|
| 25 |
+
|
| 26 |
## Usage
|
| 27 |
|
| 28 |
```python
|
|
|
|
| 45 |
|
| 46 |
| Model | Average | Size (M) | Embed dim | ArguAna | CQADupstackRetrieval | ClimateFEVER | DBPedia | FEVER | FiQA2018 | HotpotQA | MSMARCO | NFCorpus | NQ | QuoraRetrieval | SCIDOCS | SciFact | TRECCOVID | Touche2020 |
|
| 47 |
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 48 |
+
| [ColBERTv2](https://huggingface.co/colbert-ir/colbertv2.0) | 48.63 | 110 | 128 | 46.50 | 38.30 | 17.60 | 45.20 | 78.50 | 35.40 | 67.50 | 46.00 | 33.70 | 52.40 | 85.50 | 15.40 | 68.90 | 72.60 | 26.00 |
|
| 49 |
+
| [Jina-ColBERT-v2](https://huggingface.co/jinaai/jina-colbert-v2) | 51.85 | 600 | 128 | 36.60 | 40.80 | 23.90 | 47.10 | 80.50 | 40.80 | 76.60 | **46.90** | 34.60 | 64.00 | 88.70 | 18.60 | 67.80 | 83.40 | 27.40 |
|
| 50 |
| [ColBERT-small](https://huggingface.co/answerdotai/answerai-colbert-small-v1) | 53.79 | 33 | 96 | 50.09 | 38.75 | 33.07 | 45.58 | 90.96 | 41.15 | 76.11 | 43.50 | 37.30 | 59.10 | 87.72 | 18.42 | 74.77 | 84.59 | 25.69 |
|
| 51 |
+
| [GTE-ModernColBERT-v1](https://huggingface.co/lightonai/GTE-ModernColBERT-v1) | 54.75 | 149 | 128 | 47.52 | 41.08 | 31.33 | 47.56 | 87.67 | 45.25 | 77.48 | 45.60 | **37.83** | 61.62 | 86.71 | 19.22 | 76.33 | **84.84** | 31.25 |
|
| 52 |
| [ColBERT-Zero](https://huggingface.co/lightonai/ColBERT-Zero) | 55.39 | 149 | 128 | **52.82** | 41.41 | 35.90 | 47.43 | 90.52 | 42.50 | 79.45 | 45.95 | 37.21 | 61.82 | 85.19 | 19.84 | 76.33 | 78.27 | **36.24** |
|
| 53 |
+
| [LateOn-unsupervised](https://huggingface.co/lightonai/LateOn-unsupervised) | 50.11 | 149 | 128 | 43.12 | **47.71** | 18.76 | 43.36 | 65.74 | 51.94 | 68.17 | 37.51 | 37.15 | 58.41 | 89.48 | 21.13 | 76.89 | 69.81 | 22.53 |
|
| 54 |
+
| [LateOn](https://huggingface.co/lightonai/LateOn) | 57.22 | 149 | 128 | 50.52 | 47.36 | **39.67** | 45.99 | 92.02 | **53.12** | 79.98 | 45.67 | 37.79 | 63.91 | 89.67 | **21.90** | 76.61 | 83.60 | 30.52 |
|
| 55 |
+
| GLINT-base | **57.43** | 149 | 128 | 52.38 | 46.49 | 34.17 | **47.68** | **92.45** | 50.85 | **82.54** | 46.38 | 37.51 | **68.03** | **90.08** | 20.65 | **77.13** | 84.78 | 30.26 |
|
| 56 |
|
| 57 |
GLINT-base was evaluated with the project BEIR protocol: corpus IDs are excluded from their own
|
| 58 |
query results for ArguAna and Quora, and ArguAna uses 64 query tokens. The 57.43 average is one
|