File size: 6,256 Bytes
2cfb4cd
 
 
 
 
 
48f5407
2cfb4cd
 
 
 
 
 
 
48f5407
 
2cfb4cd
48f5407
 
 
2cfb4cd
48f5407
2cfb4cd
48f5407
2cfb4cd
48f5407
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2cfb4cd
 
 
48f5407
 
2cfb4cd
 
 
 
 
48f5407
2cfb4cd
48f5407
 
 
2cfb4cd
48f5407
 
 
 
 
 
 
 
 
2cfb4cd
 
48f5407
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2cfb4cd
 
48f5407
2cfb4cd
 
 
 
 
48f5407
 
 
 
 
 
 
 
 
 
 
 
 
 
2cfb4cd
 
48f5407
 
 
 
2cfb4cd
48f5407
 
2cfb4cd
 
 
48f5407
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
---
license: unknown
tags:
- security
- sql-injection
- scikit-learn
- pytorch
- anomaly-detection
- text-classification
pipeline_tag: text-classification
---

# VNU SQLi Detection Models

Trained model artifacts for a **3-branch** AI-based SQL Injection detection system, deployed at
the Database Proxy layer:

- **Branch 1** — supervised multi-class classifier (Normal + SQLi variants), per query.
- **Branch 2** — anomaly detector trained on benign traffic only (generalises to unseen syntax).
- **Branch 3** — session-level sequence model over a stream of requests (the main contribution).

Training data: [Jason-42195/VNU-SQLi-Detection](https://huggingface.co/datasets/Jason-42195/VNU-SQLi-Detection).

## Contents

| Folder | Branch | Model | Headline metric |
|---|---|---|---|
| `branch1_v1/` | 1 | TF-IDF (char) + Logistic Regression | F1-macro **0.9822** (5 classes) |
| `branch1_comparison/` | 1 | 4 candidate architectures (5-class) | see table below |
| `branch1_no_*/` | 1 | Leave-one-class-out ablations (zero-day coverage) | F1 ~0.98 on remaining |
| `branch2_v1/` | 2 | One-Class SVM (4 structural features) | avg-precision **0.982** |
| `branch2_zeroday/` | 2 | One-Class SVM (zero-day coverage variant) | — |
| `branch3_v1/` | 3 | GRU session-sequence classifier (4 classes) | see project repo |

Download everything: `hf download Jason-42195/VNU-SQLi-Detection-Models --local-dir models/`

---

## `branch1_v1/` — Branch 1 (supervised multiclass) — **production**

TF-IDF (`char_wb`, 2–4 gram, 50k features) + Logistic Regression. Classifies a query into one of
**5 classes**: `normal`, `union_based`, `error_based`, `boolean_blind`, `time_blind`.

- **F1-macro: 0.9822** on the held-out test set (13,560 rows).
- p50 latency ~0.8 ms, size ~3.5 MB.
- Files: `vectorizer.joblib` (TfidfVectorizer), `model.joblib` (LogisticRegression), `metadata.json`.

```python
import joblib
vectorizer = joblib.load("branch1_v1/vectorizer.joblib")
clf = joblib.load("branch1_v1/model.joblib")

X = vectorizer.transform(["1' OR '1'='1"])
clf.predict(X)  # -> array([3])  (3 = boolean_blind)
```

## `branch1_comparison/` — Branch 1 architecture comparison (5-class)

Four candidates trained on the same 5-class data (`train/compare_branch1_architectures.py`).
The neural heads are sized from the data (true 5-class, no dead `stacked` neuron). TF-IDF+LogReg
was chosen for production on the latency/size trade-off — the F1 gap is negligible.

| Candidate | F1-macro | p50 latency | Size |
|---|---|---|---|
| `candidate_tfidf_logreg` | 0.9822 | 0.8 ms | 3.5 MB |
| `candidate_tfidf_lightgbm` | 0.9912 | 91.7 ms | 5.7 MB |
| `candidate_distilbert` | 0.9892 | 2.9 ms (GPU) | 256 MB |
| `candidate_cnn_sqltok` | 0.9838 | 0.3 ms | 0.11 MB (28.5K params) |

Only the two neural candidates' weights are hosted here (`candidate_distilbert/`:
`model.safetensors` + tokenizer/config; `candidate_cnn_sqltok/`: `model.pt` + `vocab.json`).

```python
# DistilBERT candidate
from transformers import AutoModelForSequenceClassification, AutoTokenizer
tok = AutoTokenizer.from_pretrained("branch1_comparison/candidate_distilbert")
model = AutoModelForSequenceClassification.from_pretrained("branch1_comparison/candidate_distilbert")
# model.config.num_labels == 5
```

The CNN candidate is a small TextCNN over a char-level SQL tokenizer; reconstruct it with the
helper in `train/compare_branch1_architectures.py` (`_build_textcnn_class`, `_encode_texts`) and the
saved `vocab.json` — see `train/notebooks/demo_detect.ipynb` for a runnable example.

## `branch1_no_*/` — leave-one-class-out ablations (zero-day coverage)

Four Branch-1 models, each trained with **one attack class excluded**
(`branch1_no_union_based`, `branch1_no_error_based`, `branch1_no_boolean_blind`,
`branch1_no_time_blind`). Used to measure whether the anomaly branch / combined system still
catches an attack family the supervised model was never trained on. Each scores F1-macro ~0.98 on
its remaining classes. Same file layout as `branch1_v1/`; `metadata.json` records `excluded_label`.

## `branch2_v1/` — Branch 2 (anomaly detection)

One-Class SVM trained on benign traffic only, using 4 structural features
(`length`, `special_char_ratio`, `sql_keyword_count`, `entropy`) — not TF-IDF, so it can
generalise to unseen attack syntax. Average precision ~0.982 (full PR curve / threshold sweep in
the project repo's `report/metrics/`).

```python
# From within the project repo (needs src.models.branch2_anomaly.AnomalyDetector):
from src.models.branch2_anomaly import AnomalyDetector
import numpy as np

detector = AnomalyDetector.load("branch2_v1")
X = np.array([[40, 0.05, 1, 3.6]])  # [length, special_char_ratio, sql_keyword_count, entropy]
detector.score(X)          # continuous anomaly score
detector.anomaly_flags(X)  # boolean flag
```

## `branch2_zeroday/` — Branch 2 (zero-day coverage variant)

One-Class SVM with the same config as `branch2_v1`, used in the zero-day coverage experiment
(pairs with the `branch1_no_*` ablations). Same loader as `branch2_v1`.

## `branch3_v1/` — Branch 3 (session-level sequence model)

A GRU over a session's stream of per-request feature vectors (`input_dim=7`, `hidden_dim=32`,
`max_len=64` requests/session). Classifies a **session** into 4 classes: `benign`,
`boolean_blind`, `time_blind`, `query_splitting`. Files: `model.pt` (state dict), `metadata.json`.
See the project repo for the model definition and the session-feature extractor needed to load it.

---

## Limitations

- Branch 1's `boolean_blind` class has ~13% measured label noise (catch-all bucket for unmatched
  attack rows) — see `data_contract.md` in the project repo.
- The uniformly high Branch-1 F1 (~0.98–0.99 across all 4 architectures) indicates the current
  dataset is easy to separate; it is **not** an adversarial/obfuscation benchmark.
- No adversarial/obfuscation robustness testing yet.
- License: mixed/unclear for the underlying training data (see the dataset repo's card) — treat as
  research / course-project artifacts, not cleared for unrestricted reuse.

## Project repo

Full source, training scripts, and documentation: the `VNU-Database2-Project` repo
(private / course project — ask the author for access).