ACE-privacy-filter-zhtw
Model Description
ACE-privacy-filter-zhtw is a PII detection model from APMIC's ACE family, engineered to detect, classify, and enable the neutralization of personally identifiable information (PII) within Traditional Chinese (zh-TW) text. It is the enterprise sibling of an internal research lineage, hardened for production and aligned to the realities of Taiwanese data — government records, financial documents, healthcare notes, and customer correspondence.
The model treats privacy not as a post-processing step, but as a native behavior: given free-form text, it labels every sensitive span in a single forward pass, so identifiers can be masked or removed while the surrounding meaning is preserved.
It is fine-tuned from OpenAI's open-weight openai/privacy-filter — a bidirectional token-classification PII detector (1.5B total / 50M active parameters, MoE backbone with banded attention) — extending its 8 original PII labels with 11 Taiwan-specific ones. The corpora behind its zh-TW alignment and the full methodology of its training recipe, however, remain proprietary to APMIC. What is shared here is what it does — not entirely how it came to do it.
Model Details
- Developed by: APMIC
- Funded by: APMIC, led by CEO Jerry Wu
- Model type: Token classification (BIOES tagging, 77 classes = 1 background + 19 PII labels × 4 boundary tags) with constrained Viterbi decoding; not a generative / chat model
- Language(s): Traditional Chinese (zh-TW) & English
- License: APMIC proprietary (enterprise use; contact APMIC for terms)
- Base model:
openai/privacy-filter(Apache 2.0) — OpenAI's open-weight PII detector, fine-tuned by APMIC for zh-TW privacy filtering. (The training recipe and zh-TW alignment corpora remain proprietary.) - Format: OPF-native checkpoint, loaded with the
opfruntime
What It Does
Given Traditional Chinese text, ACE-privacy-filter-zhtw:
- Detects personally identifiable information embedded in natural, conversational, and document-style language.
- Classifies each identifier into a privacy category.
- Locates it precisely — returning character offsets for each span, so downstream code can redact, mask, or replace it while keeping the text readable and semantically intact.
It is designed to operate on the messy, real-world text where regex and rule engines fail: mixed Chinese-English content, inconsistent formatting, OCR-derived noise, and the idiomatic phrasing of Taiwanese business and government communication.
Privacy Entity Coverage
The filter is tuned toward identifiers that matter in a Taiwanese context, including (but not limited to):
- 身分證字號 (National ID numbers)
- 健保卡號 / 病歷號 (NHI card & medical record numbers)
- 手機與市話號碼 (Mobile & landline numbers)
- 地址 (Residential & mailing addresses)
- 銀行帳號與信用卡號 (Bank account & card numbers)
- 姓名 (Personal names)
- Email 與帳號識別碼 (Email & account identifiers)
- 車牌號碼 (Vehicle plate numbers)
- 公司統一編號 (Business registration numbers)
Data Foundation
The structural backbone of ACE-privacy-filter-zhtw's privacy understanding draws on nvidia/Nemotron-PII — NVIDIA's large-scale synthetic corpus of 100,000 records spanning 55+ PII/PHI categories across 50+ industries, covering both structured documents (forms, invoices) and unstructured content (emails, notes).
This foundation gave the model a broad, industry-spanning prior over what privacy looks like — across healthcare, finance, legal, and enterprise scenarios. APMIC then carried that prior across the language boundary, re-grounding it in the entity types, formats, and cultural conventions specific to Traditional Chinese and Taiwan. The bridge from Nemotron-PII's English foundation to native zh-TW behavior is where APMIC's proprietary work lives.
NVIDIA Ecosystem
ACE-privacy-filter-zhtw is part of APMIC's broader collaboration with NVIDIA's data and platform ecosystem. It builds on NVIDIA-originated privacy data, is optimized for inference on modern NVIDIA GPU architectures, and is designed to slot into enterprise deployment pipelines alongside other models in the ACE family.
Intended Use
- De-identification of Traditional Chinese documents prior to storage, analytics, or LLM ingestion.
- Privacy guardrails in conversational AI and RAG pipelines handling Taiwanese user data.
- Compliance support for organizations operating under Taiwan's 個人資料保護法 (Personal Data Protection Act) and adjacent regulatory regimes.
Out of Scope
- The model is an assistive control, not a legal guarantee. It does not certify compliance, and its output should be reviewed in high-stakes settings.
- It is not a general-purpose chat assistant.
- Performance on languages or locales outside Traditional Chinese / Taiwan is not a design target.
Usage
📘 完整使用指南(繁體中文): USAGE.md — 安裝、CLI、Python API、批次處理、輸出格式與類別說明。
The model is an OPF-native token-classification checkpoint. Load it with the
opfruntime from openai/privacy-filter, not withtransformers(AutoModelForCausalLM/pipeline) — doing so skips the constrained Viterbi decoder and badly fragments Chinese spans.
pip install -e git+https://github.com/openai/privacy-filter#egg=opf
hf download APMIC/ACE-privacy-filter-zhtw --local-dir ./ACE-privacy-filter-zhtw
opf --checkpoint ./ACE-privacy-filter-zhtw "您好,我是王小明,身分證字號 A123456789,手機 0912-345-678,住台北市信義區市府路1號。"
The output is a list of labelled character spans, compatible with openai/privacy-filter:
[
{"entity": "private_person", "start": 5, "end": 8},
{"entity": "tw_national_id", "start": 15, "end": 25},
{"entity": "private_phone", "start": 29, "end": 41},
{"entity": "private_address", "start": 42, "end": 54}
]
Label space (tw_pii_v1, see tw_label_space.json): the 8 original labels (account_number, private_address, private_date, private_email, private_person, private_phone, private_url, secret) plus 11 Taiwan-specific labels (tw_national_id, tw_nhi_card, tw_company_id, tw_passport, tw_line_id, tw_license_plate, tw_driver_license, tw_household_no, tw_ptt_id, tw_medical_license, tw_military_id).
Positioning
ACE-privacy-filter-zhtw demonstrates APMIC's capacity to take a foundation of NVIDIA privacy data and forge it into a Traditional-Chinese-native, enterprise-ready privacy layer — for organizations that need their data protected before it is ever processed, and who would rather not know exactly how the lock was made.
Disclaimer
This model is provided for enterprise privacy-filtering use. No PII detection system is perfect; APMIC makes no warranty that all sensitive information will be identified or removed. Operators remain responsible for validating outputs and meeting their own regulatory obligations.
© APMIC. Part of the ACE model family.
- Downloads last month
- 7
Model tree for APMIC/ACE-privacy-filter-zhtw
Base model
openai/privacy-filter
