Instructions to use tknika/gliner2-PII-basque-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- GLiNER2
How to use tknika/gliner2-PII-basque-v2 with GLiNER2:
from gliner2 import AutoExtractor extractor = AutoExtractor.from_pretrained("tknika/gliner2-PII-basque-v2") # Extract entities text = "Apple CEO Tim Cook announced iPhone 15 in Cupertino yesterday." result = extractor.extract_entities(text, ["company", "person", "product", "location"]) print(result) - Notebooks
- Google Colab
- Kaggle
gliner2-PII-basque-v2
Version 2 of tknika/gliner2-PII-basque, a Basque-adapted fine-tune of fastino/gliner2-privacy-filter-PII-multi (205M-parameter multilingual PII detection model, GLiNER2 architecture, mDeBERTa-v3-base encoder).
Compared to v1, this version is trained with a v2 synthetic replay dataset with realistic value distributions (name frequencies from public statistics, coherent street/town/postcode triples from OpenStreetMap) and three entity types specific to educational contexts: user name (LMS/forum handles, including @-mentions), personal url (personal blogs and profiles) and student id.
Training
Two data sources were combined (experience replay, to avoid catastrophic forgetting of the base model's PII capabilities):
- Basque NER: the
nerc_idsplit of orai-nlp/basqueGLUE (2,842 sentences), mapped to the base model's types:person name,location,organization,miscellaneous. - Synthetic PII: the train split (5,000 sentences) of tknika/pii-synthetic-basque-v2 — Basque, Spanish and French sentences covering the whole Basque Country (Araba, Bizkaia, Gipuzkoa, Nafarroa, Iparralde), 17 PII types, checksum-valid values (mod-23 national IDs, mod-97 IBANs, Luhn-valid cards, real phone formats).
Training used LoRA (r=16, alpha=32) on the boundary/task heads, merged into the base model after training (~13 MB adapter). This repository contains the merged standalone model, the original adapter (lora-adapter/) and the full training recipe (errezeta/).
Results
PII detection
On the eval split of pii-synthetic-basque-v2 (1,000 sentences, 17 types, disjoint from training by construction):
| Model | Precision | Recall | Micro F1 |
|---|---|---|---|
| Base (zero-shot) | 0.647 | 0.742 | 0.691 |
| v1 (gliner2-PII-basque) | 0.761 | 0.862 | 0.808 |
| This model (v2) | 0.949 | 0.919 | 0.934 |
Per-type F1 of this model (types sorted by frequency; base / v1 shown for the educational types):
| Type | F1 | Type | F1 | |
|---|---|---|---|---|
| person name | 0.957 | personal url | 1.000 (base 0.175, v1 0.727) | |
| 1.000 | organization | 0.726 | ||
| phone number | 0.992 | location | 0.000 âš | |
| user name | 0.996 (base 0.864, v1 0.671) | student id | 0.894 (v1 zero-shot 0.972) | |
| date | 1.000 | bank account number | 1.000 | |
| address | 0.990 | national id | 0.771 | |
| city | 0.723 | age | 1.000 | |
| date of birth | 0.997 | credit card number | 1.000 | |
| zip code | 0.997 |
Basque NER
BasqueGLUE validation sets (500 sentences per split), showing that the PII fine-tuning does not cause forgetting of the Basque NER capabilities:
| Metric | Base (zero-shot) | This model |
|---|---|---|
| nerc_id/val person name F1 | 0.664 | 0.843 |
| nerc_id/val micro F1 | 0.464 | 0.769 |
| nerc_od/val (Wikipedia) person name F1 | 0.721 | 0.808 |
| nerc_od/val micro F1 | 0.547 | 0.725 |
Usage
from gliner2 import AutoExtractor
model = AutoExtractor.from_pretrained("tknika/gliner2-PII-basque-v2")
text = ("Kaixo, @mikel_etxeberria naiz ikaslea, ikasle zk 123456B dut. "
"Nire bloga https://mikeletxeberria.wordpress.com da eta "
"helbidea Barakaldoko Nagusia kalea 42 da, 48901.")
result = model.extract_entities(
text,
["person name", "email", "phone number", "national id",
"bank account number", "credit card number", "date of birth",
"date", "address", "city", "location", "zip code", "age",
"organization", "user name", "personal url", "student id"],
)
Limitations
- The PII evaluation set is synthetic; real-world performance (messier text, ambiguous contexts) will be lower.
- The model systematically labels town names as
cityeven wherelocationwould be expected (locationF1 is 0 on the eval set). national idprecision is 0.63 (over-prediction on other numeric identifiers).- The synthetic replay data covers 17 of the 42 PII types of the base model; other types may degrade after fine-tuning.
- Research-quality evaluation; not production-tested.
- The license of the BasqueGLUE/EIEC training corpus should be verified before redistribution of derivatives beyond this model.
- Downloads last month
- -
Model tree for tknika/gliner2-PII-basque-v2
Base model
fastino/gliner2-privacy-filter-PII-multi