Buckets:
| license: cc-by-4.0 | |
| language: | |
| - en | |
| - pt | |
| - es | |
| - fr | |
| - de | |
| - ar | |
| tags: | |
| - pii | |
| - privacy | |
| - ner | |
| - token-classification | |
| - gdpr | |
| - lgpd | |
| - multilingual | |
| - redaction | |
| - data-governance | |
| - compliance | |
| size_categories: | |
| - 1M<n<10M | |
| task_categories: | |
| - token-classification | |
| pretty_name: "PII Shield — Multilingual PII Detection Dataset by Auren Research" | |
| # PII Shield: Multilingual PII Detection Dataset | |
| <p align="center"> | |
| <img src="https://img.shields.io/badge/Languages-6-blue" alt="6 Languages"> | |
| <img src="https://img.shields.io/badge/Examples-2.6M+-green" alt="2.6M+ Examples"> | |
| <img src="https://img.shields.io/badge/PII_Types-40+-orange" alt="40+ PII Types"> | |
| <img src="https://img.shields.io/badge/License-CC--BY--4.0-lightgrey" alt="CC BY 4.0"> | |
| <img src="https://img.shields.io/badge/Domains-5+-purple" alt="5+ Domains"> | |
| </p> | |
| **PII Shield** is a large-scale, multilingual dataset for training and evaluating Personally Identifiable Information (PII) detection models. Built by [Auren Research](https://huggingface.co/auren-research), it combines real-world documents from diverse domains with high-quality span-level PII annotations produced by [fastino/gliner2-privacy-filter-PII-multi](https://huggingface.co/fastino/gliner2-privacy-filter-PII-multi) — achieving the highest F1 on the SPY benchmark among open-source PII detectors. | |
| The dataset is designed to support production-grade compliance workflows under **GDPR**, **LGPD**, **CCPA**, and related privacy regulations. | |
| --- | |
| ## Dataset Overview | |
| | Attribute | Value | | |
| |---|---| | |
| | **Total examples** | ~2.6M | | |
| | **English examples** | ~531k | | |
| | **Translated examples** | ~2.1M (5 languages) | | |
| | **Languages** | English, Portuguese, Spanish, French, German, Arabic | | |
| | **PII entity types** | 40+ (see full list below) | | |
| | **Domains** | Corporate email, Legal, Financial, Healthcare, General | | |
| | **Annotation model** | fastino/gliner2-privacy-filter-PII-multi | | |
| | **Translation model** | tencent/HY-MT1.5-1.8B | | |
| | **License** | CC BY 4.0 | | |
| --- | |
| ## Motivation | |
| Despite growing regulatory pressure around data privacy, high-quality **multilingual** PII detection datasets remain scarce. Most existing resources are: | |
| - English-only or cover only 2–3 languages | |
| - Limited to a single domain (e.g. clinical notes or emails) | |
| - Fully synthetic, with limited coverage of real-world PII distribution | |
| - Too small for training production-grade models | |
| **PII Shield** addresses all four gaps: it is built from real documents across five domains, annotated with a state-of-the-art open-source PII detector, and expanded to six languages through a placeholder-preserving translation pipeline that ensures PII entities are never corrupted during translation. | |
| --- | |
| ## Data Sources | |
| The English portion of the dataset was collected and deduplicated from the following public sources: | |
| | Source | HuggingFace ID | Domain | Documents | | |
| |---|---|---|---| | |
| | Enron Email Dataset | `corbt/enron-emails` | Corporate email | ~495k | | |
| | AI4Privacy | `ai4privacy/pii-masking-200k` | General | ~43k | | |
| | Gretel PII Masking | `gretelai/gretel-pii-masking-en-v1` | Healthcare / Financial | ~48k | | |
| | Nemotron-PII | `nvidia/Nemotron-PII` | 50+ industries | ~100k | | |
| After deduplication via exact-match on the `text` column, the final English set contains **531,729 documents**. | |
| --- | |
| ## Annotation Pipeline | |
| All English documents were annotated using **[fastino/gliner2-privacy-filter-PII-multi](https://huggingface.co/fastino/gliner2-privacy-filter-PII-multi)**, a GLiNER2-based multilingual PII detector that achieves the highest span-level F1 (0.477) on the SPY benchmark — outperforming OpenAI Privacy Filter, NVIDIA GLiNER-PII, and urchade/gliner_multi_pii-v1. | |
| ### Annotation settings | |
| - **Detection threshold:** 0.5 | |
| - **Max document length:** 8,000 characters (longer documents truncated) | |
| - **Hardware:** AMD Instinct MI300X (192 GB) | |
| - **Throughput:** ~8,000–8,500 documents/second | |
| ### PII entity types detected (40 types) | |
| | Group | Labels | | |
| |---|---| | |
| | **Person / names** | `person`, `full_name`, `first_name`, `middle_name`, `last_name`, `date_of_birth` | | |
| | **Contact / address** | `email`, `phone_number`, `address`, `street_address`, `city`, `state_or_region`, `postal_code`, `country` | | |
| | **Government / tax IDs** | `government_id`, `national_id_number`, `passport_number`, `drivers_license_number`, `license_number`, `tax_id`, `tax_number` | | |
| | **Banking / payment** | `bank_account`, `account_number`, `routing_number`, `iban`, `payment_card`, `card_number`, `card_expiry`, `card_cvv` | | |
| | **Digital identity** | `username`, `ip_address`, `account_id`, `sensitive_account_id` | | |
| | **Secrets / credentials** | `password`, `secret`, `api_key`, `access_token`, `recovery_code` | | |
| | **Sensitive dates** | `sensitive_date`, `document_date`, `expiration_date`, `transaction_date` | | |
| ### Annotation results | |
| | Metric | Value | | |
| |---|---| | |
| | Documents with PII detected | 516,880 (97.2%) | | |
| | Documents without PII | 14,849 (2.8%) | | |
| | Total documents | 531,729 | | |
| --- | |
| ## Translation Pipeline | |
| The 516,880 documents containing PII were translated into 5 additional languages using **[tencent/HY-MT1.5-1.8B](https://huggingface.co/tencent/HY-MT1.5-1.8B)**, a state-of-the-art multilingual machine translation model. | |
| ### Placeholder-preserving translation | |
| To prevent PII corruption during translation, we implemented a **placeholder injection** strategy: | |
| 1. Before translation, all detected PII spans are replaced with positional tokens (`__X0__`, `__X1__`, ...) | |
| 2. The document with placeholders is translated | |
| 3. After translation, placeholders are replaced with the original PII values | |
| This ensures that names, emails, phone numbers, and other identifiers are never mistranslated or altered. | |
| ### Translation settings | |
| - **Inference engine:** vLLM 0.21.0 | |
| - **Hardware:** AMD Instinct MI300X (192 GB) | |
| - **Batch size:** 512 | |
| - **Throughput:** ~100–450 documents/second (varies by document length) | |
| - **Max document length:** 1,500 characters | |
| ### Language distribution | |
| | Language | Code | Examples | | |
| |---|---|---| | |
| | English | `en` | ~531k | | |
| | Portuguese | `pt` | ~516k | | |
| | Spanish | `es` | ~516k | | |
| | French | `fr` | ~516k | | |
| | German | `de` | ~516k | | |
| | Arabic | `ar` | ~516k | | |
| | **Total** | | **~2.6M** | | |
| --- | |
| ## Dataset Structure | |
| ### Files | |
| ``` | |
| auren-research/pii-shield/ | |
| ├── en/ | |
| │ └── auren_privacy_annotated.parquet # English, annotated | |
| ├── pt/ | |
| │ └── auren_privacy_pt.parquet # Portuguese translations | |
| ├── es/ | |
| │ └── auren_privacy_es.parquet # Spanish translations | |
| ├── fr/ | |
| │ └── auren_privacy_fr.parquet # French translations | |
| ├── de/ | |
| │ └── auren_privacy_de.parquet # German translations | |
| └── ar/ | |
| └── auren_privacy_ar.parquet # Arabic translations | |
| ``` | |
| ### Schema | |
| **English file (`auren_privacy_annotated.parquet`):** | |
| | Column | Type | Description | | |
| |---|---|---| | |
| | `text` | `string` | Document text (up to 8,000 characters) | | |
| | `has_pii` | `bool` | Whether any PII was detected | | |
| | `pii_count` | `int` | Number of PII spans detected | | |
| | `pii_types` | `string` (JSON list) | List of unique PII label types found | | |
| | `spans` | `string` (JSON list) | List of `{"label": str, "text": str}` objects | | |
| **Translated files (`auren_privacy_{lang}.parquet`):** | |
| | Column | Type | Description | | |
| |---|---|---| | |
| | `text` | `string` | Translated document text | | |
| | `original_text` | `string` | Original English document | | |
| | `language` | `string` | Language code (e.g. `pt`) | | |
| | `has_pii` | `bool` | Inherited from English annotation | | |
| | `pii_count` | `int` | Inherited from English annotation | | |
| | `pii_types` | `string` (JSON list) | Inherited from English annotation | | |
| | `spans` | `string` (JSON list) | Original English PII spans (values preserved via placeholder injection) | | |
| --- | |
| ## Usage | |
| ```python | |
| from datasets import load_dataset | |
| import json | |
| # Load English annotated data | |
| ds = load_dataset("auren-research/pii-shield", "en", split="train") | |
| # Load a specific language | |
| ds_pt = load_dataset("auren-research/pii-shield", "pt", split="train") | |
| # Filter only documents with PII | |
| pii_docs = ds.filter(lambda x: x["has_pii"]) | |
| # Parse spans | |
| example = ds[0] | |
| spans = json.loads(example["spans"]) | |
| for span in spans: | |
| print(f"[{span['label']}] {span['text']}") | |
| ``` | |
| --- | |
| ## Intended Use Cases | |
| - **Training PII detection / NER models** for GDPR and LGPD compliance pipelines | |
| - **Benchmarking** multilingual PII detectors across real-world document types | |
| - **Data governance tooling** — redaction, pseudonymization, de-identification | |
| - **Research** on cross-lingual transfer of privacy-related information extraction | |
| - **Fine-tuning** encoder models (BERT, DeBERTa, ModernBERT) for span-level PII classification | |
| --- | |
| ## Limitations | |
| The following limitations should be considered before using this dataset in production: | |
| **Annotation limitations:** | |
| - The annotation model (GLiNER2-PII) has known precision issues with `person` and `full_name` labels — it can confuse common nouns and organization names with personal names | |
| - IPv6 addresses, URLs, and device identifiers (e.g. IMEI numbers) may be misclassified as `phone_number` or missed entirely | |
| - Annotation confidence scores are not stored — all spans at threshold ≥ 0.5 are included | |
| **Translation limitations:** | |
| - Machine translation quality varies by language and domain — Arabic and German may contain more errors than Portuguese and Spanish | |
| - Translations were not validated by human annotators | |
| - Documents were truncated to 1,500 characters before translation — very long documents may have incomplete translations | |
| **Coverage limitations:** | |
| - All source documents are in English — the multilingual examples are translations, not natively authored text | |
| - Medical domain coverage is limited (Nemotron-PII provides some coverage, but MIMIC-III was not included) | |
| - No Asian languages (Chinese, Japanese, Korean, Hindi) in this release | |
| --- | |
| ## Ethical Considerations | |
| The English source documents contain **real PII from real people**, collected from publicly available datasets (Enron emails, legal proceedings, SEC filings). While these datasets are in the public domain or released under open licenses, users of PII Shield should: | |
| - **Not use this dataset** to build surveillance, tracking, or profiling systems | |
| - **Not attempt to identify** individuals from the PII spans | |
| - **Follow the upstream licenses** of each source dataset | |
| - **Comply with applicable privacy laws** in their jurisdiction when deploying models trained on this data | |
| This dataset is intended for **privacy-enhancing** use cases only — building tools that detect and redact PII, not tools that exploit it. | |
| --- | |
| ## Source Dataset Licenses | |
| | Source | License | | |
| |---|---| | |
| | Enron Email Dataset | Public domain | | |
| | AI4Privacy | CC BY 4.0 | | |
| | Gretel PII Masking | Apache 2.0 | | |
| | Nemotron-PII | CC BY 4.0 | | |
| The compiled dataset is released under **CC BY 4.0**. Please cite the original sources if you use this dataset in academic work. | |
| --- | |
| ## Citation | |
| If you use PII Shield in your research or products, please cite: | |
| ```bibtex | |
| @dataset{auren2026piishield, | |
| title = {PII Shield: A Multilingual PII Detection Dataset}, | |
| author = {Barroso, Francisco Antonio Da Costa}, | |
| year = {2026}, | |
| publisher = {Auren Research}, | |
| url = {https://huggingface.co/datasets/auren-research/pii-shield} | |
| } | |
| ``` | |
| Please also cite the annotation and translation models: | |
| ```bibtex | |
| @misc{fastino2026gliner2pii, | |
| title = {GLiNER2-PII: Multilingual PII Extraction via Synthetic Fine-Tuning}, | |
| author = {{Fastino AI Team}}, | |
| year = {2026}, | |
| url = {https://huggingface.co/fastino/gliner2-pii-v1} | |
| } | |
| @misc{tencent2025hymt, | |
| title = {HY-MT: Tencent Hunyuan Machine Translation}, | |
| author = {{Tencent Hunyuan Team}}, | |
| year = {2025}, | |
| url = {https://huggingface.co/tencent/HY-MT1.5-1.8B} | |
| } | |
| ``` | |
| --- | |
| ## About Auren Research | |
| [Auren Research](https://github.com/Auren-Research) is an independent AI research lab focused on AI safety, privacy-preserving machine learning, and open-source model development. Our work includes: | |
| - **Lunaris MoC** — a novel sparse Transformer with mediator-based expert collaboration ([technical report](https://github.com/Auren-Research/lunaris)) | |
| - **Lunaris Guard** — a multilingual dual-head safety classifier (ROC-AUC 0.979 on prompt injection, 0.928 on content safety) | |
| - **PII Shield** — this dataset | |
| All research is open-source and reproducible. | |
| --- | |
| *Built with ❤️ by [Francisco Antonio Da Costa Barroso](https://github.com/MeryylleA) · Auren Research · 2026* |
Xet Storage Details
- Size:
- 12.6 kB
- Xet hash:
- c3ff040475c6804fc07cabc0926293c094372fccc21fbfd286401350430cdf67
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.