BERT Legal (Topic-masked, continual pretraining)
bert-base-uncased continually pretrained (MLM) on a legal corpus, using
topic-based selective masking (tf-idf weighted) instead of random masking:
during training, content words that distinguish a document within the domain
(e.g. "contract", "asylum", "tax") are masked preferentially.
Training data
A subset of lexlms/lex_files
covering ECtHR case law, EU case law, and Indian case law.
Training
- Base model:
bert-base-uncased - Objective: masked language modeling, 15% masking rate, words sampled proportionally to their tf-idf weight
- 10 epochs, chunk size 512, effective batch size 1024, learning rate 2e-5
Method
The masking strategy follows:
Belfathi, A., Gallina, Y., Hernandez, N., Monceaux, L., & Dufour, R. (2025). Is Selective Masking A Key to Improving Domain Adaptation for Masked Language Model? ICAIL 2025. doi:10.1145/3769126.3769216
Usage
from transformers import AutoModelForMaskedLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("AnasBelfathi/bert-legal-topic")
model = AutoModelForMaskedLM.from_pretrained("AnasBelfathi/bert-legal-topic")
- Downloads last month
- 27
Model tree for AnasBelfathi/bert-legal-topic
Base model
google-bert/bert-base-uncased