--- license: apache-2.0 library_name: transformers pipeline_tag: text-classification base_model: FacebookAI/xlm-roberta-base tags: - ner - on-device - privacy - flowx - openner - cross - de-identification - text-classification metrics: - f1 --- # PrivacyFilter **PrivacyFilter** is a small, on-device cross text classifier from the FlowX **OpenNER** family. Developed by **FlowX.AI**. Runs 100% on-premise / air-gapped, so no data leaves your boundary. ## What it does - **Task:** text-classification - **Base model:** `FacebookAI/xlm-roberta-base` - **Classes (4):** NONE, PERSONAL, FINANCIAL, HEALTH - **Held-out F1:** 1.0000 - **Runtime:** CPU, Apple Silicon, one GPU, or browser/edge via ONNX (INT8). ~100-160 ms/doc on CPU. ## Why a small model Fine-tuned encoders match or beat frontier LLMs on structured, convention-bound extraction, at a fraction of the latency and cost, with **zero data egress**. Identifiers are validated by checksum (IBAN mod-97, card Luhn, ISIN/LEI, container ISO-6346, VIN, national IDs), a correctness guarantee general LLMs lack. See the FlowX OpenNER benchmark for measured results. ## Usage ```python from transformers import AutoTokenizer, AutoModelForTokenClassification tok = AutoTokenizer.from_pretrained("flowxai/privacyfilter") model = AutoModelForTokenClassification.from_pretrained("flowxai/privacyfilter") ``` ## License & attribution Licensed under the **Apache License 2.0**. Copyright 2026 **FlowX.AI** (https://flowx.ai). See the `NOTICE` file. Trained on synthetic, checksum-validated data. _Part of the FlowX OpenNER model family. Synthetic-data F1 reflects an in-distribution synthetic distribution; validate on real documents before production use._