Instructions to use tknika/gliner2-PII-basque with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- GLiNER2
How to use tknika/gliner2-PII-basque with GLiNER2:
from gliner2 import AutoExtractor extractor = AutoExtractor.from_pretrained("tknika/gliner2-PII-basque") # Extract entities text = "Apple CEO Tim Cook announced iPhone 15 in Cupertino yesterday." result = extractor.extract_entities(text, ["company", "person", "product", "location"]) print(result) - Notebooks
- Google Colab
- Kaggle
gliner2-PII-basque
A Basque-adapted version of fastino/gliner2-privacy-filter-PII-multi, a 205M-parameter multilingual PII (personally identifiable information) detection model built on GLiNER2 with an mDeBERTa-v3-base encoder.
The base model recognizes 42 PII entity types, but underperforms on Basque text: person names, declension suffixes, and agglutinative forms are frequently missed. This model fine-tunes the base on Basque named-entity data while preserving the original multilingual PII detection capabilities.
Training
Two data sources were combined (experience replay):
- Basque NER: the
nerc_idsplit of orai-nlp/basqueGLUE (EIEC + naiz.eus, manually annotated PER/LOC/ORG/MISC; 2,842 sentences), mapped to the base model's existing types:person name,location,organization,miscellaneous. - Synthetic PII: 6,000 synthetic sentences in Basque, Spanish and English covering 14 common PII types, with checksum-valid values (Spanish national IDs mod-23, IBANs mod-97, Luhn-valid cards, phone formats). This follows the synthetic-data methodology of the GLiNER2-PII paper, since the original training corpus is not public.
Training used LoRA (r=16, alpha=32) on the boundary/task heads, merged into the base model after training; the resulting adapter is only ~13 MB. The model card repository contains the merged standalone model.
Results
Evaluation on BasqueGLUE validation sets (500 sentences per split, exact match on type + mention string):
| Metric | Base (zero-shot) | This model |
|---|---|---|
| nerc_id/val person name F1 | 0.664 | 0.850 |
| nerc_id/val micro F1 | 0.464 | 0.787 |
| nerc_od/val (Wikipedia) person name F1 | 0.721 | 0.813 |
On a PII probe sentence combining several entity types:
| PII type | Base (zero-shot) | This model |
|---|---|---|
| person name | β | β |
| β | β | |
| phone number | β | β |
| bank account number | β | β |
| date of birth | β | β |
| zip code | β | β |
| national id | β | β |
| address | β | β |
Fine-tuning a PII model on new-language NER data alone typically causes catastrophic forgetting of the original entity types. The LoRA + experience replay combination used here preserves the base model's PII detection behavior (and even improves on national id and address) while achieving strong Basque NER results.
Usage
from gliner2 import AutoExtractor
model = AutoExtractor.from_pretrained("tknika/gliner2-PII-basque")
text = ("Kaixo, Ane Azkune naiz, Bilboko La Casilla kalean bizi naiz. "
"Posta: ane.azkune@euskaltel.eus, telefonoa 688 123 456.")
result = model.extract_entities(
text,
["person name", "location", "address", "email", "phone number",
"national id", "bank account number", "credit card number",
"date of birth", "city", "zip code", "age", "date"],
)
Limitations
- Research-quality evaluation (500 sentences per split); not production-tested.
- The synthetic replay data covers 14 of the 42 PII types of the base model; types not covered may still degrade after fine-tuning.
- The license of the BasqueGLUE/EIEC training corpus should be verified before redistribution of derivatives beyond this model.
- The training corpus text is reconstructed by joining tokens with whitespace, so punctuation spacing is slightly non-standard.
- Downloads last month
- -
Model tree for tknika/gliner2-PII-basque
Base model
fastino/gliner2-privacy-filter-PII-multi