LexSPLADE (bert-large-uncased, silver, Longman Defining Vocabulary)

A sparse contextual word encoder whose output vocabulary is restricted to the Longman Defining Vocabulary. Given a sentence with one word marked, it returns a sparse vector describing that word in that context โ€” but only 1,897 of the 30,524 vocabulary dimensions (6.2%) can ever be non-zero, and each is a basic-English defining word.

The effect is that a vector reads as a definition:

"He sat on the <t> bank </t> of the river"      -> bank shore bed edge water wall river face
"She withdrew the money from her <t> bank </t>" -> bank money account pay check payment trust debt
"The soldiers began their <t> charge </t>"      -> charge attack march offensive advance defense battle
"There is no <t> charge </t> for delivery"      -> charge cost tax payment pay rate price ticket

The restriction is trained, not applied afterwards: the mask is in place during training, so the gradients, the regularizers and the similarity all live in the restricted subspace.

The target word must be wrapped in <t> โ€ฆ </t>. Pooling runs over the tokens between those markers only.

Usage

from sentence_transformers import SparseEncoder
import torch

model = SparseEncoder("jadermcs/lexsplade-bert-large-silver-ldv", trust_remote_code=True)

sents = [
    "He sat on the <t> bank </t> of the river and watched the water.",
    "She withdrew the money from her <t> bank </t> on Monday morning.",
]
v = model.encode(sents, convert_to_tensor=True).to_dense().float()

print(torch.nn.functional.cosine_similarity(v[0:1], v[1:2], dim=1))  # ~0.17

top = v[0].topk(8).indices
print([model.tokenizer.convert_ids_to_tokens(int(i)) for i in top])
# ['bank', 'shore', 'bed', 'edge', 'water', 'wall', 'river', 'face']

trust_remote_code=True is required. The restriction lives in the custom pooling module (modeling_lexsplade.py, ids in 1_TargetWordSpladePooling/config.json). Loading the weights as a plain AutoModelForMaskedLM and pooling yourself will produce unrestricted vectors, which is not what this model was trained to emit.

Keep the marked word inside max_seq_length; if the markers are truncated away the pooling falls back to the whole sentence.

The vocabulary

Built from the Longman Dictionary of Contemporary English defining vocabulary โ€” the ~2,000 basic words in which every LDOCE definition is written. Three filters:

step dropped remaining
published list, single-word entries โ€” 2,188
must be one whole-word token in the BERT vocabulary (advertise โ†’ advert ##ise has no dimension of its own) 63 2,125
no stopwords or punctuation 228 1,897

Shipped in this repo as longman_defining_vocabulary.txt (the word list) and ldv_vocabulary.json (words plus the token ids the mask uses). The word list is reproduced from the published LDOCE defining vocabulary for research use; the dictionary and its defining vocabulary are the work of Longman/Pearson.

Training

base model google-bert/bert-large-uncased
training data 31,914 same-lemma word-in-context pairs (LLM-verified silver + MCL-WiC train)
objective AnglE loss on the sparse target-word vectors, with a SPLADE FLOPS regularizer and a Barlow Twins term on the dense span representation
output space 1,897 LDV dimensions, masked throughout training
seed 42

Evaluation

MCL-WiC, English. AP is average precision; accuracy is at the best threshold.

split accuracy AP
test 90.70 96.16
dev 90.90 95.36

Mean nonzero dimensions per usage: 118 of 30,524 โ€” roughly 60% fewer than the unrestricted model, at the same MCL-WiC AP (96.16 vs 96.30).

Variants

  • jadermcs/lexsplade-bert-large-silver-ldv (this model) โ€” 1,897 LDV dimensions.
  • jadermcs/lexsplade-bert-large-silver โ€” identical training, full 30,524-dim output.
Downloads last month
10
Safetensors
Model size
0.3B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for jadermcs/lexsplade-bert-large-silver-ldv

Finetuned
(160)
this model