Sentence Similarity
sentence-transformers
Safetensors
English
modernbert
colbert
late-interaction
retrieval
pylate
text-embeddings-inference
Instructions to use chungimungi/GLInt with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use chungimungi/GLInt with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("chungimungi/GLInt") sentences = [ "That is a happy person", "That is a happy dog", "That is a very happy person", "Today is a sunny day" ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
File size: 6,806 Bytes
4e1b7b3 07f195b 4e1b7b3 aebf33a a043934 4e1b7b3 2d208fb 4e1b7b3 7bd207d 07f195b 4e1b7b3 2d208fb 4e1b7b3 e582b00 4e1b7b3 07f195b 4e1b7b3 2d208fb 07f195b 4e1b7b3 07f195b 4e1b7b3 07f195b 4e1b7b3 07f195b 4e1b7b3 07f195b c6633f4 07f195b c6633f4 07f195b c6633f4 2d208fb 4e1b7b3 bb0f8f2 349c238 bb0f8f2 07f195b 4e1b7b3 07f195b b7408db e582b00 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 | ---
library_name: sentence-transformers
tags:
- colbert
- late-interaction
- retrieval
- pylate
language:
- en
base_model:
- lightonai/LateOn-unsupervised
license: apache-2.0
pipeline_tag: sentence-similarity
---
# GLInt
GLINT is a SOTA 149M-parameter English late-interaction retriever built from
[LateOn-unsupervised](https://huggingface.co/lightonai/LateOn-unsupervised). It retains
128-dimensional token embeddings and uses MaxSim retrieval with 32 query tokens and 300
document tokens.
## What is new in GLInt?
GLINT is designed around the mismatch between ordinary dense hard-negative mining and a
late-interaction retriever. Dense mining selects documents that are difficult under one pooled
vector; GLINT instead mines negatives under the same token-level MaxSim geometry used at
retrieval time. This exposes lexical, compositional, and localized token matches that a
single-vector miner can miss.
The training recipe has two stages:
1. supervised fine-tuning with multi-vector (MaxSim) hard negatives;
2. mixed listwise knowledge distillation over a diverse seven-source hard-negative mixture.
For the second stage, a frozen listwise teacher (`jinaai/jina-reranker-v3.5`) scores each
32-document candidate set jointly. GLINT distils that ordering with a sharpened listwise KL
objective, while a false-negative-masked InfoNCE term preserves a direct retrieval signal.
## Usage
```python
from pylate import models
model = models.ColBERT("chungimungi/GLInt")
query_embeddings = model.encode(["what causes a lunar eclipse?"], is_query=True)
document_embeddings = model.encode(
["A lunar eclipse happens when Earth passes between the Sun and the Moon."],
is_query=False,
)
```
Use a late-interaction backend such as PyLate/PLAID for corpus-scale retrieval. Scores are
computed by summing, over query tokens, the maximum similarity to a document token.
## Results
### BEIR (15 datasets, NDCG@10)
| Model | Average | Size (M) | Embed dim | ArguAna | CQADupstackRetrieval | ClimateFEVER | DBPedia | FEVER | FiQA2018 | HotpotQA | MSMARCO | NFCorpus | NQ | QuoraRetrieval | SCIDOCS | SciFact | TRECCOVID | Touche2020 |
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
| [ColBERTv2](https://huggingface.co/colbert-ir/colbertv2.0) | 48.63 | 110 | 128 | 46.50 | 38.30 | 17.60 | 45.20 | 78.50 | 35.40 | 67.50 | 46.00 | 33.70 | 52.40 | 85.50 | 15.40 | 68.90 | 72.60 | 26.00 |
| [Jina-ColBERT-v2](https://huggingface.co/jinaai/jina-colbert-v2) | 51.85 | 600 | 128 | 36.60 | 40.80 | 23.90 | 47.10 | 80.50 | 40.80 | 76.60 | **46.90** | 34.60 | 64.00 | 88.70 | 18.60 | 67.80 | 83.40 | 27.40 |
| [ColBERT-small](https://huggingface.co/answerdotai/answerai-colbert-small-v1) | 53.79 | 33 | 96 | 50.09 | 38.75 | 33.07 | 45.58 | 90.96 | 41.15 | 76.11 | 43.50 | 37.30 | 59.10 | 87.72 | 18.42 | 74.77 | 84.59 | 25.69 |
| [GTE-ModernColBERT-v1](https://huggingface.co/lightonai/GTE-ModernColBERT-v1) | 54.75 | 149 | 128 | 47.52 | 41.08 | 31.33 | 47.56 | 87.67 | 45.25 | 77.48 | 45.60 | **37.83** | 61.62 | 86.71 | 19.22 | 76.33 | **84.84** | 31.25 |
| [ColBERT-Zero](https://huggingface.co/lightonai/ColBERT-Zero) | 55.39 | 149 | 128 | **52.82** | 41.41 | 35.90 | 47.43 | 90.52 | 42.50 | 79.45 | 45.95 | 37.21 | 61.82 | 85.19 | 19.84 | 76.33 | 78.27 | **36.24** |
| [LateOn-unsupervised](https://huggingface.co/lightonai/LateOn-unsupervised) | 50.11 | 149 | 128 | 43.12 | **47.71** | 18.76 | 43.36 | 65.74 | 51.94 | 68.17 | 37.51 | 37.15 | 58.41 | 89.48 | 21.13 | 76.89 | 69.81 | 22.53 |
| [LateOn](https://huggingface.co/lightonai/LateOn) | 57.22 | 149 | 128 | 50.52 | 47.36 | **39.67** | 45.99 | 92.02 | **53.12** | 79.98 | 45.67 | 37.79 | 63.91 | 89.67 | **21.90** | 76.61 | 83.60 | 30.52 |
| GLInt | **57.43** | 149 | 128 | 52.38 | 46.49 | 34.17 | **47.68** | **92.45** | 50.85 | **82.54** | 46.38 | 37.51 | **68.03** | **90.08** | 20.65 | **77.13** | 84.78 | 30.26 |
### BEIR-Decontaminated (14 datasets, NDCG@10)
| Model | Average | ArguAna | ClimateFEVER | DBPedia | FEVER | FiQA2018 | HotpotQA | MS MARCO | NFCorpus | Natural Questions | Quora | SciDocs | SciFact | TREC-COVID | Touché-2020 |
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
| **GLInt** | **62.50** | 51.67 | 36.35 | 42.50 | 92.89 | 56.88 | **81.16** | 72.70 | 26.21 | **94.97** | **92.06** | **22.02** | **89.07** | 81.51 | 34.97 |
| LateOn | 61.4 | 52.2 | 42.1 | 31.7 | 92.7 | **57.9** | 78.9 | 70.3 | 27.0 | 93.1 | 91.5 | 15.1 | 88.9 | 80.9 | 36.8 |
| DenseOn | 58.8 | 40.0 | 39.5 | 28.8 | 91.2 | 55.9 | 73.7 | 68.9 | 28.5 | 92.1 | 91.1 | 14.7 | 85.4 | 82.5 | 31.0 |
| pplx-embed-v1-0.6b | 59.7 | 43.7 | 42.4 | 28.4 | 91.1 | 55.2 | 73.5 | 71.9 | 28.0 | 91.6 | 91.5 | 15.4 | 89.0 | 83.7 | 30.0 |
| jina-v5-text-nano | 58.8 | 47.2 | 41.6 | 30.2 | 90.0 | 51.5 | 67.5 | 68.6 | 29.4 | 92.3 | 91.3 | 14.9 | 89.4 | 76.8 | 33.2 |
| harrier-oss-v1-0.6b | 58.0 | 47.4 | 25.7 | 31.3 | 80.7 | 50.1 | 71.4 | **73.4** | 27.9 | 90.0 | 90.9 | 17.1 | **90.7** | 81.8 | 33.3 |
| arctic-embed-l-v2 | 57.9 | 43.1 | **45.7** | **45.7** | 92.2 | 50.4 | 63.1 | 71.0 | 26.0 | 90.7 | 91.3 | 13.9 | 87.4 | 81.4 | 26.8 |
| bge-large-en-v1.5 | 57.3 | 46.0 | 39.0 | 28.9 | 87.6 | 49.3 | 75.2 | 68.9 | **29.8** | 85.9 | 91.3 | 14.0 | 86.5 | 72.7 | 26.9 |
| Qwen3-Embedding-0.6B | 57.0 | 48.4 | 38.0 | 25.3 | 86.4 | 49.1 | 62.2 | 63.6 | 25.8 | 88.3 | 90.0 | 15.3 | 85.5 | **87.9** | 31.8 |
| GTE-ModernBERT | 56.6 | 52.5 | 47.5 | 25.9 | **94.1** | 55.5 | 65.5 | 64.8 | 26.1 | 84.5 | 90.8 | 11.6 | 88.6 | 62.4 | 23.1 |
| bge-base-en-v1.5 | 56.2 | 45.6 | 32.9 | 26.7 | 86.8 | 44.5 | 72.7 | 66.8 | 27.4 | 85.6 | 91.1 | 13.8 | 87.6 | 76.6 | 28.1 |
| Nomic v1.5 | 55.9 | 35.8 | 43.5 | 28.8 | 86.8 | 44.7 | 72.7 | 67.4 | 24.4 | 85.1 | 87.2 | 12.7 | 83.3 | 80.7 | 29.4 |
| modernbert-embed-base | 55.6 | 36.5 | 37.8 | 24.7 | 87.8 | 46.0 | 62.7 | 65.3 | 24.3 | 89.3 | 89.9 | 12.9 | 85.5 | 82.7 | 33.1 |
| ColBERT-Zero | 60.0 | 54.5 | 36.8 | 33.0 | 90.5 | 46.6 | 77.8 | **74.2** | 26.6 | 91.1 | 88.3 | 14.2 | 89.5 | 75.3 | **40.9** |
| pplx-embed-v1-late-0.6b | 59.8 | **60.9** | 36.4 | 29.9 | 89.7 | 50.9 | 78.6 | 69.2 | 27.9 | 92.8 | 83.8 | 13.5 | 89.3 | 80.2 | 34.7 |
| GTE-ModernColBERT | 59.3 | 48.8 | 33.5 | 33.2 | 88.1 | 50.2 | 77.3 | 71.6 | 27.3 | 93.1 | 89.1 | 13.6 | 87.7 | 81.4 | 35.3 |
| colbert-small | 58.1 | 47.7 | 35.7 | 31.7 | 89.3 | 45.6 | 77.1 | 71.4 | 25.0 | 86.2 | 90.1 | 13.1 | 89.2 | 81.5 | 29.0 |
## Training data and reproducibility
The corresponding private training artifacts are in
[GLINT-data](https://huggingface.co/datasets/chungimungi/GLINT-data). It contains the complete
prepared SFT data, the 1,046,009-row seven-source KD mixture, and Jina teacher-score parquet
shards. The repository contains no BEIR evaluation corpus or evaluation labels.
|