gliner2-PII-basque

A Basque-adapted version of fastino/gliner2-privacy-filter-PII-multi, a 205M-parameter multilingual PII (personally identifiable information) detection model built on GLiNER2 with an mDeBERTa-v3-base encoder.

The base model recognizes 42 PII entity types, but underperforms on Basque text: person names, declension suffixes, and agglutinative forms are frequently missed. This model fine-tunes the base on Basque named-entity data while preserving the original multilingual PII detection capabilities.

Training

Two data sources were combined (experience replay):

  • Basque NER: the nerc_id split of orai-nlp/basqueGLUE (EIEC + naiz.eus, manually annotated PER/LOC/ORG/MISC; 2,842 sentences), mapped to the base model's existing types: person name, location, organization, miscellaneous.
  • Synthetic PII: 6,000 synthetic sentences in Basque, Spanish and English covering 14 common PII types, with checksum-valid values (Spanish national IDs mod-23, IBANs mod-97, Luhn-valid cards, phone formats). This follows the synthetic-data methodology of the GLiNER2-PII paper, since the original training corpus is not public.

Training used LoRA (r=16, alpha=32) on the boundary/task heads, merged into the base model after training; the resulting adapter is only ~13 MB. The model card repository contains the merged standalone model.

Results

Evaluation on BasqueGLUE validation sets (500 sentences per split, exact match on type + mention string):

Metric Base (zero-shot) This model
nerc_id/val person name F1 0.664 0.850
nerc_id/val micro F1 0.464 0.787
nerc_od/val (Wikipedia) person name F1 0.721 0.813

On a PII probe sentence combining several entity types:

PII type Base (zero-shot) This model
person name βœ“ βœ“
email βœ“ βœ“
phone number βœ“ βœ“
bank account number βœ“ βœ“
date of birth βœ“ βœ“
zip code βœ“ βœ“
national id βœ— βœ“
address βœ— βœ“

Fine-tuning a PII model on new-language NER data alone typically causes catastrophic forgetting of the original entity types. The LoRA + experience replay combination used here preserves the base model's PII detection behavior (and even improves on national id and address) while achieving strong Basque NER results.

Usage

from gliner2 import AutoExtractor

model = AutoExtractor.from_pretrained("tknika/gliner2-PII-basque")

text = ("Kaixo, Ane Azkune naiz, Bilboko La Casilla kalean bizi naiz. "
        "Posta: ane.azkune@euskaltel.eus, telefonoa 688 123 456.")
result = model.extract_entities(
    text,
    ["person name", "location", "address", "email", "phone number",
     "national id", "bank account number", "credit card number",
     "date of birth", "city", "zip code", "age", "date"],
)

Limitations

  • Research-quality evaluation (500 sentences per split); not production-tested.
  • The synthetic replay data covers 14 of the 42 PII types of the base model; types not covered may still degrade after fine-tuning.
  • The license of the BasqueGLUE/EIEC training corpus should be verified before redistribution of derivatives beyond this model.
  • The training corpus text is reconstructed by joining tokens with whitespace, so punctuation spacing is slightly non-standard.
Downloads last month
-
Safetensors
Model size
0.3B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for tknika/gliner2-PII-basque

Finetuned
(4)
this model

Paper for tknika/gliner2-PII-basque