Instructions to use LiquidAI/LFM2.5-Encoder-350M-PII-Detector with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use LiquidAI/LFM2.5-Encoder-350M-PII-Detector with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="LiquidAI/LFM2.5-Encoder-350M-PII-Detector", trust_remote_code=True)# Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("LiquidAI/LFM2.5-Encoder-350M-PII-Detector", trust_remote_code=True) model = AutoModelForTokenClassification.from_pretrained("LiquidAI/LFM2.5-Encoder-350M-PII-Detector", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
language:
- en
- de
- fr
- es
- pt
- it
- pl
- ru
- zh
- ja
- ko
- ar
- hi
- id
- vi
- th
tags:
- liquid
- lfm2
- lfm2.5
- bidirectional
- masked-lm
- encoder
- pii
- ner
- privacy
- multilingual
- token-classification
library_name: transformers
license: other
license_name: lfm1.0
license_link: LICENSE
pipeline_tag: token-classification
base_model:
- LiquidAI/LFM2.5-Encoder-350M
LFM2.5-Encoder-350-PII-Detector
A full fine-tune of LFM2.5-Encoder-350M with a token-classification head, covering 40 PII types across 16 languages (en, de, fr, es, pt, it, pl, ru, zh, ja, ko, ar, hi, id, vi, th).
Ships with an inference-timemhybrid regex decode (pii_hybrid_decode.py) that adds validator-gated formats (email/IBAN/credit-card/IP/JWT/…) and cue-gated IDs on top of the model.
Trained on a persona-driven, gemma-generated synthetic corpus (coherent locale-personas × scenarios × cue/inline/structured embedding × ID-contrastive disambiguation), LLM-judge-filtered and contamination-cleaned against all evaluation sets.
Find more details about our encoders in our blog post.
💻 Demos: Try this fine-tuned model running in a CPU-only Hugging Face space: PII detection** — spot and remove 40 kinds of personal information across 16 languages.
Entity types (40 PII types across 11 domains)
| Domain | Types |
|---|---|
| Identity | identity.person_name, identity.ssn, identity.national_id, identity.passport, identity.drivers_license, identity.date_of_birth, identity.tax_id |
| Contact | contact.email, contact.phone, contact.address, contact.postal_code, contact.ip_address |
| Financial | financial.credit_card, financial.iban, financial.bank_account, financial.swift_bic, financial.crypto_wallet, financial.amount |
| Credentials | credential.api_key, credential.password, credential.private_key, credential.jwt, credential.connection_string, developer.login_credentials |
| Online | online.username, online.url |
| Device | device.mac_address, device.imei, developer.device_id |
| Location | location.gps_coordinates |
| Healthcare | healthcare.medical_record, healthcare.condition, healthcare.medication, healthcare.health_plan_id |
| Organization | org.company_name |
| Special-category | special.religion, special.political, special.orientation, special.health_status |
| Legal | legal.case_number |
Benchmarks (18-locale-filtered, partial-F1, hybrid decode)
| Benchmark | this model | detection-tier | prev (v8) | GLiNER | LFM-demo-q4 |
|---|---|---|---|---|---|
| SPY | 0.428 | 0.509 | 0.351 | 0.280 | 0.192 |
| Gretel | 0.880 | 0.885 | 0.758 | 0.663 | 0.804 |
| TAB | 0.867 | 0.888 | 0.749 | 0.685 | 0.490 |
| ai4privacy | 0.715 | 0.774 | 0.643 | 0.488 | 0.500 |
| Nemotron | 0.855 | 0.863 | 0.773 | 0.639 | 0.656 |
| MAPA | 0.236 | 0.267 | 0.486 | 0.416 | 0.250 |
| Internal (40-type) | 0.720 | 0.829 | 0.616 | 0.479 | 0.466 |
| ShieldFlow | 0.901 | 0.911 | 0.847 | 0.646 | 0.839 |
| ShieldFlow-xl | 0.859 | 0.871 | 0.797 | 0.658 | 0.842 |
- Best overall across general/multilingual benchmarks and the ShieldFlow product gate; beats SauerkrautLM-GLiNER and the LFM demo on every benchmark except MAPA's idiosyncratic date-as-
date_of_birthlabeling convention. - Detection-tier (did it find the PII span, ignoring fine type — the metric that matters for redaction) is markedly higher than exact-type, e.g. Internal 0.83 / ShieldFlow 0.91.
Usage
⚠️ Loads custom code via
trust_remote_code=True(the model wraps atrust_remote_codeencoder).
Install the required packages:
pip install torch transformers huggingface_hub
Run PII detection:
import importlib.util
import sys
from huggingface_hub import hf_hub_download
from transformers import AutoModelForTokenClassification, AutoTokenizer
model_id = "LiquidAI/LFM2.5-Encoder-350-PII-Detector"
helper_path = hf_hub_download(model_id, "pii_hybrid_decode.py")
hf_hub_download(model_id, "context_cued.py")
sys.path.insert(0, helper_path.rsplit("/", 1)[0])
spec = importlib.util.spec_from_file_location("pii_hybrid_decode", helper_path)
hd = importlib.util.module_from_spec(spec)
spec.loader.exec_module(hd)
tok = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForTokenClassification.from_pretrained(model_id, trust_remote_code=True).eval()
spans = hd.predict("Email Dr. Laura Schmidt at laura@charite.de.", tok, model)
print(spans)
📬 Contact
- Got questions or want to connect? Join our Discord community
- If you are interested in custom solutions with edge deployment, please contact our sales team.
Citation
@article{liquidAI2026Encoders,
author = {Liquid AI},
title = {LFM2.5-Encoders: Fast at Long Context, Even on CPU},
journal = {Liquid AI Blog},
year = {2026},
note = {www.liquid.ai/blog/lfm2-5-encoders},
}
