File size: 2,566 Bytes
97959d0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
---
base_model: microsoft/Phi-3-mini-4k-instruct
library_name: peft
pipeline_tag: text-classification
tags:
- base_model:adapter:microsoft/Phi-3-mini-4k-instruct
- lora
- transformers
- text-classification
- data-governance
- pii-detection
license: mit
language:
- en
---

# phi3-mini-sensitive-lora

A LoRA adapter that fine-tunes **Phi-3-mini** to classify text as **sensitive
(1)** or **not sensitive (0)** — for detecting PII, financial data, credentials,
HR/medical info, and government IDs mixed into company databases.

Part of the **AI Data Governance Platform** project.

## Results (held-out test set, n = 900)

| Metric | Baseline (untrained head) | Fine-tuned (this adapter) |
|--------|---------------------------|---------------------------|
| Accuracy | 0.529 | **1.000** |
| Precision | 0.645 | **1.000** |
| Recall | 0.091 | **1.000** |
| F1 | 0.159 | **1.000** |

*Note: 1.00 reflects strong learning of a controlled synthetic dataset;
real-world data would need a human-labeled test set to confirm generalization.*

## Training

- **Base model:** microsoft/Phi-3-mini-4k-instruct (loaded in 4-bit / QLoRA)
- **Method:** LoRA (rank 16, α 32, dropout 0.05) on the sequence-classification head
- **Trainable params:** 25.2M (0.67%)
- **Data:** 6,000 synthetic supermarket records (Indian locale), 70/15/15 split
- **Hardware:** RTX 3060 12GB, ~38 min, 3 epochs

## Usage

```python
import torch
from peft import PeftModel
from transformers import AutoModelForSequenceClassification, AutoTokenizer, BitsAndBytesConfig

BASE = "microsoft/Phi-3-mini-4k-instruct"
ADAPTER = "shivam14245/phi3-mini-sensitive-lora"

tok = AutoTokenizer.from_pretrained(BASE)
tok.pad_token = tok.pad_token or tok.eos_token

bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
                         bnb_4bit_compute_dtype=torch.bfloat16, bnb_4bit_use_double_quant=True)
model = AutoModelForSequenceClassification.from_pretrained(BASE, num_labels=2,
                                                           quantization_config=bnb, device_map="cuda")
model.config.pad_token_id = tok.pad_token_id
model = PeftModel.from_pretrained(model, ADAPTER).eval()

text = "employee salary 85000 bank_account 9876543210 ifsc HDFC0001234"
inputs = tok(text, return_tensors="pt", truncation=True, max_length=256).to("cuda")
label = model(**inputs).logits.argmax(-1).item()   # 1 = sensitive, 0 = not
print("sensitive" if label == 1 else "not sensitive")
```

## Labels
`0` = not sensitive · `1` = sensitive (pii / financial / credentials / hr_medical / govt_id)