LexSPLADE (bert-large-uncased, silver)

A sparse contextual word encoder. Given a sentence with one word marked, it returns a sparse vector over the 30,524-token BERT vocabulary describing that word in that context โ€” a word embedding, not a sentence embedding. Because the dimensions are vocabulary entries, a vector can be read directly as a list of words.

The target word must be wrapped in <t> โ€ฆ </t>. Pooling runs over the tokens between those markers only.

Usage

from sentence_transformers import SparseEncoder
import torch

model = SparseEncoder("jadermcs/lexsplade-bert-large-silver", trust_remote_code=True)

sents = [
    "He sat on the <t> bank </t> of the river and watched the water.",
    "She withdrew the money from her <t> bank </t> on Monday morning.",
]
v = model.encode(sents, convert_to_tensor=True).to_dense().float()

print(torch.nn.functional.cosine_similarity(v[0:1], v[1:2], dim=1))  # ~0.15

# Read a vector as words
top = v[0].topk(8).indices
print([model.tokenizer.convert_ids_to_tokens(int(i)) for i in top])
# ['bank', 'banks', 'side', 'water', ',', 'bed', 'shore', 'edge']

trust_remote_code=True is required: the target-span pooling is a custom module shipped in this repo as modeling_lexsplade.py.

Keep the marked word inside max_seq_length. If the markers are truncated away the pooling falls back to the whole sentence and silently returns a sentence embedding.

Training

base model google-bert/bert-large-uncased
training data 31,914 same-lemma word-in-context pairs (LLM-verified silver + MCL-WiC train)
objective AnglE loss on the sparse target-word vectors, with a SPLADE FLOPS regularizer and a Barlow Twins term on the dense span representation
seed 42

Evaluation

MCL-WiC, English. AP is average precision; accuracy is at the best threshold.

split accuracy AP
test 91.20 96.30
dev 90.40 95.24

Mean nonzero dimensions per usage: 307 of 30,524.

Variants

  • jadermcs/lexsplade-bert-large-silver (this model) โ€” full 30,524-dim output.
  • jadermcs/lexsplade-bert-large-silver-ldv โ€” identical training, but the output is restricted to 1,897 dimensions drawn from the Longman Defining Vocabulary, so every vector reads as a definition in basic English.
Downloads last month
7
Safetensors
Model size
0.3B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for jadermcs/lexsplade-bert-large-silver

Finetuned
(160)
this model