ds-EkaCare's picture
Wording: faster, not fastest
c088fc4 verified
|
Raw
History Blame Contribute Delete
8.54 kB
---
license: other
license_name: per-model-licenses
license_link: https://huggingface.co/ekacare/document-pii-redactor#license-and-attribution
tags: [pii, redaction, document-ai]
---
# document-pii-redactor model weights
**Try it live:** [medai.eka.care/pii-redactor](https://medai.eka.care/pii-redactor)
(faster), or the [Hugging Face Space](https://huggingface.co/spaces/ekacare/document-pii-redactor)
— upload a document image or paste text and see detection + redaction run
in your browser. Background and design notes are in the
[launch blog](https://info.eka.care/services/releasing-parrotlet-open-models-open-weights-for-medical-document-intelligence).
Most PII redactors stop at plain text. These models power
[document-pii-redactor](https://github.com/eka-care/document-pii-redactor),
which also redacts **document images** and is light enough to deploy on
**CPU**. They are trained to understand **Indian names, documents, and
contexts**, and the text model works across **Indian languages**. The main
contribution is the **PII token classifier** — OCR is just the pluggable
input stage in front of it. It defaults to lightweight **Tesseract**, which
keeps memory low and works well for PDFs and good-quality images; for more
difficult or blurred images, **Bring-your-own OCR** lets a model like [**Nemotron OCR**](https://huggingface.co/nvidia/nemotron-ocr-v2) (or
Textract, Google Vision, etc) plug straight in for better results
([example notebook](https://github.com/eka-care/document-pii-redactor/blob/main/examples/byo_ocr_nemotron.ipynb)).
Single repo holding the models used by the document-pii-redactor GitHub repo
(https://github.com/eka-care/document-pii-redactor), organized by modality. You
can use the document-pii-redactor repo directly to use these models — it has
the library, a Docker image, and a FastAPI server.
- `image/layoutlmv3/` — a text-in-image classifier for **text PII in
images** (47 categories), run on Tesseract OCR words.
- `image/yolo/best.pt` — a detector for **visual entities** (signature,
seal/stamp, QR/barcode, face photo, fingerprint, logo).
- `text/minilm/` — a lightweight multilingual classifier for **PII in
plain text** (no image, no OCR).
## Install
```bash
pip install "document-pii-redactor[visual]" # full pipeline, used by the examples below (AGPL-3.0 — see License below)
pip install document-pii-redactor # text pipeline + built-in OCR only, no visual entities (permissive licenses)
```
`ImagePIIRedactor` detects **visual entities** by default — PII that is an
image region rather than readable text: signatures, seals/stamps,
QR codes/barcodes, face photos, fingerprints, and logos. That detector
needs the `[visual]` extra, so the image examples below assume it — or
pass `detect_visual=False` to detect only text PII on the core install.
System dependency: **Tesseract OCR** — used by the **image** modality's
built-in OCR step. Not needed for the text-only modality, nor if you bring
your own OCR (`detect(..., words=..., boxes=...)`).
```bash
# Debian/Ubuntu
sudo apt-get install -y tesseract-ocr
# macOS
brew install tesseract
```
## Usage
**`detect()` is the core primitive** — it finds every PII entity with its
location, category, and confidence, and runs the models exactly once. The
transforms (redact / anonymize / de-identify) take its result as a required
argument: detect once, feed the result to any transform.
```python
from document_pii_redactor import ImagePIIRedactor, TextPIIRedactor
image_redactor = ImagePIIRedactor("ekacare/document-pii-redactor")
entities = image_redactor.detect("page.jpg") # built-in Tesseract OCR
# each entity: kind ("text"/"visual"), category, bbox (pixels), text, score
# …or bring your own OCR — pass words + pixel boxes, Tesseract is skipped
# and your exact boxes come back on the detected entities:
entities = image_redactor.detect("page.jpg", words=["John", "Doe"],
boxes=[[100, 20, 140, 40], [145, 20, 180, 40]])
text_redactor = TextPIIRedactor("ekacare/document-pii-redactor")
text = "Mr. John Doe, 45 yrs, DOB 12-03-1979, Indiranagar, Bangalore. Contact: +91 98765 43210."
spans = text_redactor.detect(text) # char-offset spans
```
**Redact** — destroy:
```python
image_redactor.redact("page.jpg", entities, mode="blur").save("redacted.png") # or "solid" / "pixelate"
text_redactor.redact(text, spans)
# '[REDACTED], [REDACTED] yrs, DOB [REDACTED], [REDACTED], [REDACTED]. Contact: [REDACTED].'
```
**Anonymize** — generalize, one-way, no mapping kept. Ages become 10-year
buckets, dates keep only the year, fine geography collapses to `[LOCATION]`
(state and country survive), everything else becomes an unnumbered token;
faces/signatures are filled solid:
```python
image_redactor.anonymize("page.jpg", entities).save("anonymized.png")
text_redactor.anonymize(text, spans)
# '[PERSON], 40–49 yrs, DOB 1979, [LOCATION], [LOCATION]. Contact: [PHONE].'
```
**De-identify** — pseudonymize. Same value → same pseudonym throughout the
document (rendered in place in images; faces/signatures become neutral
placeholders), and the entity→pseudonym mapping comes back for authorized
re-linking — yours to store securely, never persisted by the library.
`strategy="hash"` gives globally deterministic tokens that stay stable across
documents with no mapping to thread (`secret=` salts the hash so guessable
values can't be dictionary-reversed):
```python
deid = image_redactor.deidentify("page.jpg", entities) # .image + .mapping
deid.image.save("deidentified.png")
text_redactor.deidentify(text, spans).text
# 'Person_1, Age_1 yrs, DOB Date_1, City_1, City_2. Contact: Phone_1.'
text_redactor.deidentify(text, spans, strategy="hash").text
# 'Person_539681, Age_6c8349 yrs, DOB Date_7f19c4, City_d12704, City_60c7d5. Contact: Phone_d57003.'
```
Good to know:
- `categories=[...]` on `detect()` limits which of the 53 PII categories are
found (default: all); `detect_visual=False` on `ImagePIIRedactor` skips the
visual-entity detector entirely.
- Sequential pseudonyms are scoped to the returned `mapping` — pass
`mapping=result.mapping` on the next page of the same record to keep
numbering consistent. Hash tokens need no threading.
- Anonymization is best-effort removal/generalization of detected
identifiers — not a k-anonymity guarantee or a compliance determination.
See the [document-pii-redactor GitHub repo](https://github.com/eka-care/document-pii-redactor)
for runnable notebook walkthroughs —
[`quickstart.ipynb`](https://github.com/eka-care/document-pii-redactor/blob/main/examples/quickstart.ipynb) and
[`byo_ocr_nemotron.ipynb`](https://github.com/eka-care/document-pii-redactor/blob/main/examples/byo_ocr_nemotron.ipynb) — the full
API reference, the category taxonomy, and the Docker/FastAPI deployment setup
(the same setup behind the demo Space above).
## License and attribution
The weights are licensed **per model**, following each base model's
license (a fine-tune is a derivative of its base — the base license
flows through):
| weights | fine-tuned from | license |
|---|---|---|
| `text/minilm/` | Multilingual MiniLM (MIT) | **CC-BY-4.0** — free use incl. commercial; credit Eka Care with a link back |
| `image/layoutlmv3/` | [microsoft/layoutlmv3-base](https://huggingface.co/microsoft/layoutlmv3-base) (CC-BY-NC-SA-4.0) | **CC-BY-NC-SA-4.0****non-commercial use only**, ShareAlike |
| `image/yolo/best.pt` | YOLO11m ([Ultralytics](https://github.com/ultralytics/ultralytics), AGPL-3.0) | **AGPL-3.0** |
Practical summary: the **plain-text pipeline (`TextPIIRedactor`) has a
fully permissive lineage** and may be used commercially with attribution.
The image pipeline currently inherits its bases' restrictions — no
commercial use of the LayoutLMv3 fine-tune, and AGPL obligations for the
visual detector (the `ultralytics` runtime it needs is also AGPL-3.0 and
is an optional `[visual]` extra of the pip package).
The library code is Apache-2.0 at
https://github.com/eka-care/document-pii-redactor.
If you use the models or the library, please cite:
```bibtex
@software{document_pii_redactor,
author = {{Eka Care}},
title = {document-pii-redactor: detect, redact, de-identify, or anonymize
PII in document images and plain text},
year = {2026},
url = {https://github.com/eka-care/document-pii-redactor},
note = {Model weights: https://huggingface.co/ekacare/document-pii-redactor}
}
```