BERT Legal (Topic-masked, continual pretraining)

Selective Masking

bert-base-uncased continually pretrained (MLM) on a legal corpus, using topic-based selective masking (tf-idf weighted) instead of random masking: during training, content words that distinguish a document within the domain (e.g. "contract", "asylum", "tax") are masked preferentially.

Training data

A subset of lexlms/lex_files covering ECtHR case law, EU case law, and Indian case law.

Training

  • Base model: bert-base-uncased
  • Objective: masked language modeling, 15% masking rate, words sampled proportionally to their tf-idf weight
  • 10 epochs, chunk size 512, effective batch size 1024, learning rate 2e-5

Method

The masking strategy follows:

Belfathi, A., Gallina, Y., Hernandez, N., Monceaux, L., & Dufour, R. (2025). Is Selective Masking A Key to Improving Domain Adaptation for Masked Language Model? ICAIL 2025. doi:10.1145/3769126.3769216

Usage

from transformers import AutoModelForMaskedLM, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("AnasBelfathi/bert-legal-topic")
model = AutoModelForMaskedLM.from_pretrained("AnasBelfathi/bert-legal-topic")
Downloads last month
27
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AnasBelfathi/bert-legal-topic

Finetuned
(7016)
this model