Instructions to use jadermcs/lexsplade-bert-large-silver-ldv with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use jadermcs/lexsplade-bert-large-silver-ldv with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("jadermcs/lexsplade-bert-large-silver-ldv") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
LexSPLADE (bert-large-uncased, silver, Longman Defining Vocabulary)
A sparse contextual word encoder whose output vocabulary is restricted to the Longman Defining Vocabulary. Given a sentence with one word marked, it returns a sparse vector describing that word in that context โ but only 1,897 of the 30,524 vocabulary dimensions (6.2%) can ever be non-zero, and each is a basic-English defining word.
The effect is that a vector reads as a definition:
"He sat on the <t> bank </t> of the river" -> bank shore bed edge water wall river face
"She withdrew the money from her <t> bank </t>" -> bank money account pay check payment trust debt
"The soldiers began their <t> charge </t>" -> charge attack march offensive advance defense battle
"There is no <t> charge </t> for delivery" -> charge cost tax payment pay rate price ticket
The restriction is trained, not applied afterwards: the mask is in place during training, so the gradients, the regularizers and the similarity all live in the restricted subspace.
The target word must be wrapped in <t> โฆ </t>. Pooling runs over the tokens
between those markers only.
Usage
from sentence_transformers import SparseEncoder
import torch
model = SparseEncoder("jadermcs/lexsplade-bert-large-silver-ldv", trust_remote_code=True)
sents = [
"He sat on the <t> bank </t> of the river and watched the water.",
"She withdrew the money from her <t> bank </t> on Monday morning.",
]
v = model.encode(sents, convert_to_tensor=True).to_dense().float()
print(torch.nn.functional.cosine_similarity(v[0:1], v[1:2], dim=1)) # ~0.17
top = v[0].topk(8).indices
print([model.tokenizer.convert_ids_to_tokens(int(i)) for i in top])
# ['bank', 'shore', 'bed', 'edge', 'water', 'wall', 'river', 'face']
trust_remote_code=True is required. The restriction lives in the custom
pooling module (modeling_lexsplade.py, ids in
1_TargetWordSpladePooling/config.json). Loading the weights as a plain
AutoModelForMaskedLM and pooling yourself will produce unrestricted vectors,
which is not what this model was trained to emit.
Keep the marked word inside max_seq_length; if the markers are truncated away
the pooling falls back to the whole sentence.
The vocabulary
Built from the Longman Dictionary of Contemporary English defining vocabulary โ the ~2,000 basic words in which every LDOCE definition is written. Three filters:
| step | dropped | remaining |
|---|---|---|
| published list, single-word entries | โ | 2,188 |
must be one whole-word token in the BERT vocabulary (advertise โ advert ##ise has no dimension of its own) |
63 | 2,125 |
| no stopwords or punctuation | 228 | 1,897 |
Shipped in this repo as longman_defining_vocabulary.txt (the word list) and
ldv_vocabulary.json (words plus the token ids the mask uses). The word list is
reproduced from the published LDOCE defining vocabulary for research use;
the dictionary and its defining vocabulary are the work of Longman/Pearson.
Training
| base model | google-bert/bert-large-uncased |
| training data | 31,914 same-lemma word-in-context pairs (LLM-verified silver + MCL-WiC train) |
| objective | AnglE loss on the sparse target-word vectors, with a SPLADE FLOPS regularizer and a Barlow Twins term on the dense span representation |
| output space | 1,897 LDV dimensions, masked throughout training |
| seed | 42 |
Evaluation
MCL-WiC, English. AP is average precision; accuracy is at the best threshold.
| split | accuracy | AP |
|---|---|---|
| test | 90.70 | 96.16 |
| dev | 90.90 | 95.36 |
Mean nonzero dimensions per usage: 118 of 30,524 โ roughly 60% fewer than the unrestricted model, at the same MCL-WiC AP (96.16 vs 96.30).
Variants
jadermcs/lexsplade-bert-large-silver-ldv(this model) โ 1,897 LDV dimensions.jadermcs/lexsplade-bert-large-silverโ identical training, full 30,524-dim output.
- Downloads last month
- 10
Model tree for jadermcs/lexsplade-bert-large-silver-ldv
Base model
google-bert/bert-large-uncased