Instructions to use atonlee/Qoamo-PII-Decision with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use atonlee/Qoamo-PII-Decision with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="atonlee/Qoamo-PII-Decision")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("atonlee/Qoamo-PII-Decision", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Qoamo-PII-Decision
A routing model that catches personal data in Korean requests before they leave the device. Built on skt/A.X-Encoder-base (ModernBERT, 149M parameters), it was fine-tuned as a decision model that answers questions about a sentence with a probability for every option.
It gives four decisions for a sentence in one call: whether it contains personal data, whether it contains a strong identifier, its tier and its types. The confidence of two of them decides whether the sentence is sent as is, masked and then sent, or blocked. Locating the values is not its job. For a sentence it marks as containing personal data, atonlee/koelectra-ko-pii-ner, built on the same tier1 and tier2 criteria, finds and masks the values. This keeps the fast filtering decision separate from precise masking by NER.
On the official KDPII test split, it had the highest F1 (0.896) among the compared models under 1B parameters at telling whether a sentence contains personal data, and it also scored above the 1.4B models. The weights are a single safetensors file. ONNX exports in four precisions are in onnx/ (float32 ยท FP16 ยท INT8 ยท UINT8).
Decisions
The model can be asked four things about a sentence. Ask the ones you need in one call and get a probability for every option.
pii_anyโ contains personal data? (์ yes ยท ์๋์ค no): on "no", send as is; on "yes", find the values with NER and mask them.pii_t1โ contains a strong identifier? (์ ยท ์๋์ค): on "yes", do not send it off-device.pii_tierโ tier (tier1 ยท tier2 ยท none): picks the single highest tier in the sentence. Decide the handling for each tier.type_*โ 23 types (์ ยท ์๋์ค for each type): set the policy by the kinds of values present.
Tiers and types
| tier | meaning | recommended handling | types |
|---|---|---|---|
| tier1 | the value alone identifies a person | do not send off-device | resident registration number, alien registration number, passport number, driver's licence number, i-PIN, other national ID numbers, card number, card expiry date, CVC, account number, password or authentication key |
| tier2 | identifies a person when combined with other information | find and mask it with NER, then send | name, email, phone number, address, postal code, date of birth, age or gender, numbers assigned to a person (employee, member, student numbers and the like), user ID, URL, IP address, vehicle number |
| none | no personal data | send as is | โ |
The same sentences, four decisions
| sentence | contains personal data? (yes) | strong identifier (yes) | tier | types (yes > 0.5) |
|---|---|---|---|---|
| ๋ฐ์งํ ๋๋ฆฌํํ 010-4821-3307 ๋ก ๋ฌธ์ ๋ณด๋ด ์ค | 0.998 | 0.000 | tier2 (0.998) | name 0.993 ยท phone number 0.993 |
| ๋ด ์ฃผ๋ฏผ๋ฒํธ 901023-1234567 ๋ก ๋ฑ๋ก | 0.995 | 0.985 | tier1 (0.983) | resident registration number 0.901 |
| ์ด๋ฒ ์ฃผ๋ง์ ๋น ์ค๋์ง ์๋ ค ์ค | 0.000 | 0.000 | none (0.998) | none |
Approach โ borrowing the Jev format ยท NLI questions
The decision format borrows Jev's approach. A question and its options are attached to one text to judge, and the model gives a score for every option.
25 of the 26 questions are yes/no questions in NLI (natural language inference) form. They ask "์ด ๊ธ์ ๋ค์์ ํด๋นํ๋๊ฐ: {criterion}" ("Does this text match the following: {criterion}"), with the text to judge as the premise and the criterion as the hypothesis. Only the tier question (pii_tier) instead picks one of tier1 ยท tier2 ยท none. The text the model actually reads looks like this:
๋ฌธ์ฅ: ๋ฐ์งํ ๋๋ฆฌํํ
010-4821-3307 ๋ก ๋ฌธ์ ๋ณด๋ด ์ค
์ง๋ฌธ: ์ด ๊ธ์ ๋ค์์ ํด๋นํ๋๊ฐ: ๊ฐ์ธ์ ๋ณด ํ์ โ ์ด๋ฆยท์ ํ๋ฒํธยท์ด๋ฉ์ผยท์ฃผ์(โฆ)ยทโฆ ๊ฐ ์ ํ ์๋ค. โฆ
[OPT] ์
[OPT] ์๋์ค
Questions must be taken from the 26 in pii_config.json. Do not write new question texts; use the texts in that file as they are.
Usage
Ask the two questions (pii_any ยท pii_t1) in one call and pick the route (๋ณด๋ send ยท ํ์ธ check ยท ๋ง์ block).
pip install "transformers>=4.57" torch safetensors huggingface_hub
import json
import torch
from huggingface_hub import snapshot_download
from safetensors.torch import load_file
from transformers import AutoConfig, AutoModel, AutoTokenizer
path = snapshot_download("atonlee/Qoamo-PII-Decision")
cfg = json.load(open(f"{path}/pii_config.json", encoding="utf-8"))
tok = AutoTokenizer.from_pretrained(path)
state = load_file(f"{path}/model.safetensors")
head_w, head_b = state.pop("head.weight"), state.pop("head.bias")
model = AutoModel.from_config(AutoConfig.from_pretrained(path), dtype=torch.float32)
model.load_state_dict(state)
model.eval()
marker = tok.convert_tokens_to_ids(cfg["marker"])
def ask(sentence, questions=("pii_any", "pii_t1")):
texts = []
for q in questions:
c = cfg["questions"][q]
options = "\n".join(f"{cfg['marker']} {line}" for line in c["option_lines"])
texts.append(f"๋ฌธ์ฅ: {sentence}\n์ง๋ฌธ: {c['instructions']}\n{options}")
enc = tok(texts, padding=True, return_tensors="pt")
with torch.inference_mode():
h = model(input_ids=enc["input_ids"], attention_mask=enc["attention_mask"]).last_hidden_state
scores = (h @ head_w.T + head_b).squeeze(-1)
out = {}
for i, q in enumerate(questions):
z = scores[i][enc["input_ids"][i] == marker] / cfg["temperature"]
out[q] = dict(zip(cfg["questions"][q]["options"], torch.softmax(z, -1).tolist()))
return out
def route(p, tau=0.97):
conf = lambda d: (len(d) * max(d.values()) - 1) / (len(d) - 1)
a, t = p["pii_any"], p["pii_t1"]
if a["์๋์ค"] > a["์"] and conf(a) > tau:
return "๋ณด๋"
if t["์"] >= t["์๋์ค"] and conf(t) > 0.9:
return "๋ง์"
return "ํ์ธ"
for s in ["๋ด์ผ ์คํ ์ธ ์์ ํ ํ์ ์ก์ ์ค",
"๋ฐ์งํ ๋๋ฆฌํํ
010-4821-3307 ๋ก ๋ฌธ์ ๋ณด๋ด ์ค",
"๋ด ์ฃผ๋ฏผ๋ฒํธ 901023-1234567 ๋ก ๋ฑ๋ก"]:
p = ask(s)
print(route(p), {q: {k: round(v, 4) for k, v in d.items()} for q, d in p.items()})
Output:
๋ณด๋ {'pii_any': {'์': 0.0008, '์๋์ค': 0.9992}, 'pii_t1': {'์': 0.0002, '์๋์ค': 0.9998}}
ํ์ธ {'pii_any': {'์': 0.9978, '์๋์ค': 0.0022}, 'pii_t1': {'์': 0.0002, '์๋์ค': 0.9998}}
๋ง์ {'pii_any': {'์': 0.9948, '์๋์ค': 0.0052}, 'pii_t1': {'์': 0.9849, '์๋์ค': 0.0151}}
The tier and the types are asked the same way.
questions = ["pii_tier"] + [q for q in cfg["questions"] if q.startswith("type_")]
for s in ["๋ด์ผ ์คํ ์ธ ์์ ํ ํ์ ์ก์ ์ค",
"๋ฐ์งํ ๋๋ฆฌํํ
010-4821-3307 ๋ก ๋ฌธ์ ๋ณด๋ด ์ค",
"๋ด ์ฃผ๋ฏผ๋ฒํธ 901023-1234567 ๋ก ๋ฑ๋ก"]:
p = ask(s, questions)
tier = max(p["pii_tier"], key=p["pii_tier"].get)
types = [q[5:] for q, d in p.items() if q.startswith("type_") and d["์"] > 0.5]
print(tier, types)
Output:
none []
tier2 ['์ด๋ฆ', '์ ํ๋ฒํธ']
tier1 ['์ฃผ๋ฏผ๋ฑ๋ก๋ฒํธ']
- For a question with K options, confidence is (K ร largest probability โ 1) / (K โ 1).
- When the route is check, find and mask the personal data with NER, then send.
- To judge several sentences, build the text for each one and pass them to the model in one batch.
- For a long text, split it into lines or sentences, judge each piece, and give the whole text the strictest route (block, then check, then send).
Evaluation
Whether a sentence contains personal data was measured on the official test split of KDPII (Yonsei University HamSaeM Kim's Lab ยท TSCIENTIFIC, CC BY 4.0). Sentences were selected based on our types and the 10 common types (3,998 sentences for our types, 3,737 for the 10 types). KDPII was not used to train this model. Widely used PII detection, decision and guard models on Hugging Face were run on the same sentences and scored the same way.
| model | size | precision | recall | F1 (our types) | F1 (10 types) |
|---|---|---|---|---|---|
| Qoamo-PII-Decision | 149M | 0.895 | 0.897 | 0.896 | 0.918 |
| perplexity-ai/PII-Tracer | about 600M | 0.665 | 0.783 | 0.719 | 0.739 |
| LiquidAI/LFM2.5-Encoder-350M-PII-Detector | 350M | 0.940 | 0.503 | 0.655 | 0.738 |
| convaiinnovations/laya (Korean question) | 322M | 0.205 | 0.998 | 0.340 | 0.259 |
| Qwen/Qwen3Guard-Gen-0.6B | 0.6B | 0.556 | 0.231 | 0.326 | 0.357 |
| SupersonicLabs/Julia-1 (English question) | 144M | 0.178 | 0.523 | 0.266 | 0.205 |
| fastino/GLiNER2.5-multi-Decide (English question) | 287M | 0.279 | 0.217 | 0.244 | 0.214 |
| openai/privacy-filter (reference) | 1.4B | 0.813 | 0.612 | 0.698 | 0.775 |
| OpenMed/privacy-filter-multilingual (reference) | 1.4B | 0.530 | 0.869 | 0.658 | 0.602 |
- F1 (our types): name ยท phone ยท mobile ยท email ยท vehicle number ยท IP ยท user ID ยท date of birth ยท age ยท gender ยท resident registration number ยท alien registration number ยท passport number ยท driver's licence number ยท card number ยท account number
- F1 (10 types): name ยท phone ยท mobile ยท email ยท resident registration number ยท alien registration number ยท passport number ยท driver's licence number ยท card number ยท account number
Dataset
Training data includes the following (CC BY 4.0).
License
Apache-2.0
- Downloads last month
- 30
Model tree for atonlee/Qoamo-PII-Decision
Base model
skt/A.X-Encoder-base