File size: 12,289 Bytes
b8d911f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
---
license: apache-2.0
library_name: sklearn
tags:
  - malware
  - multi-class-classification
  - stacking
  - sban
pipeline_tag: text-classification
---

# MalwareDatasetClassification (SBAN)

**[Türkçe dokümantasyon](https://github.com/berkecanalan/MalwareDatasetClassification/blob/main/README.tr.md)** (Hub’da yalnızca İngilizce `README.md`; Türkçe metin GitHub’da)

Multiclass pipeline for **malware dataset origin classification** on [SBAN](https://github.com/ma-soreto/SBAN): four synchronized text views per sample → predict which sub-corpus it belongs to (`bodmas`, `dike`, `malwarebazaar`, `sorel20m`).

This repository contains **code, notebooks, and `sban_weighted_stacking_model.joblib`**. **No SBAN parquet or raw JSON** is distributed; obtain SBAN separately.

| Resource | Location |
|----------|----------|
| Source | [github.com/berkecanalan/MalwareDatasetClassification](https://github.com/berkecanalan/MalwareDatasetClassification) |
| Weights | Repo root + [huggingface.co/canalan/MalwareDatasetClassification](https://huggingface.co/canalan/MalwareDatasetClassification) |
| License | Apache-2.0 ([LICENSE](LICENSE)) |

---

## Task and labels

| | |
|---|---|
| **Input** | `assembly_code`, `binary_code`, `source_code`, `NLD` for one sample |
| **Output** | `dataset_name` ∈ {`bodmas`, `dike`, `malwarebazaar`, `sorel20m`} |
| **Scope** | Dataset **provenance** classification, not generic malware detection |

---

## Data preparation pipeline (scripts `01`–`11`)

End-to-end flow on local SBAN exports:

1. **`01_make_a_dataframe.py`** — Merge JSON shards under `data/M1/SBAN-MA-JUN25` into `SBAN.parquet` (four representations aligned by `ID`).
2. **`02_validate_data.py`** — Schema, missing values, duplicates, cross-dataset `ID` overlap, content fingerprints (see [Data quality](#data-quality-findings)).
3. **`03_make_clean_dataframe.py`** — Cleaning rules → `SBAN_clean.parquet`.
4. **`04_analyze_prompt_residue.py`** — Count LLM/prompt boilerplate phrases per representation ([Prompt residue](#prompt-residue-analysis)).
5. **`05_split_dataframe.py`** — Stratified train / validation / test parquet files.
6. **`06_make_features.py`** — Optional TF-IDF `.npz` features for alternate experiments.
7. **`07`–`11`** — Per-representation audits and source cleaning (`08_clean_source_code.py` uses `07_audit_source_code.py`).

Notebooks:

- **`baseline.ipynb`** — Early fusion / baseline stacking comparisons.
- **`svc_sban.ipynb`****Production model**: per-representation TF-IDF + numeric features, ID-based feature pruning, class-weight search, weighted `LinearSVC` bases, `HistGradientBoostingClassifier` meta learner, **joblib export**.
- **`inference.ipynb`** — Load exported bundle; validation/test metrics; synthetic **demo** row.

Canonical runtime entrypoint: **`inference.py`** (CLI + `StackingPredictor`).

---

## Data quality findings

Summaries below come from running the numbered scripts on the full merged SBAN table (before train/val/test split). Reproduce with your own copy of the data.

### Cross-dataset `ID` overlap (content match rate)

Shared `ID`s across corpus pairs; percentages = share of common IDs where that column’s text is **byte-identical** (`02_validate_data.py`, section 9).

| Pair | Common IDs | assembly | binary | source |
|------|------------|----------|--------|--------|
| bodmas × sorel20m | 806 | 72% | 74% | 91% |
| bodmas × malwarebazaar | 520 | 27% | 25% | 0% |
| bodmas × dike | 201 | 30% | 24% | 0% |
| dike × malwarebazaar | 82 | 23% | 22% | 4% |
| malwarebazaar × sorel20m | 99 | 41% | 47% | 0% |
| dike × sorel20m | 46 | 44% | 48% | 0% |

High overlap for **bodmas × sorel20m** (especially source) motivates careful splitting and explains why the classifier must use subtle cues, not only exact string identity across corpora.

### Rows with four aligned representations

After merge / alignment (`01_make_a_dataframe.py`):

| Dataset | Rows | Matched (4 repr.) |
|---------|------|-------------------|
| bodmas | 82,032 | 757 |
| dike | 5,342 | 669 |
| malwarebazaar | 6,048 | 905 |
| sorel20m | 71,319 | 726 |

“Matched” = samples where all four representation fields are present for labeling and training.

### Prompt residue analysis

`04_analyze_prompt_residue.py` scans fixed English phrases (e.g. “your code”, “here”, “additional”) across columns. Illustrative totals on cleaned data:

| Phrase | assembly | binary | source | NLD | Total |
|--------|----------|--------|--------|-----|-------|
| your code | 1 | 0 | 572 | 1 | 573 |
| add main function | 0 | 0 | 69 | 0 | 69 |
| code goes | 0 | 0 | 140 | 0 | 140 |
| implementation goes | 0 | 0 | 49 | 0 | 49 |
| corrected | 0 | 0 | 168 | 16 | 169 |
| here | 128 | 0 | 1,493 | 323 | 1,798 |
| no comments | 0 | 0 | 53 | 0 | 53 |
| additional | 37 | 0 | 76 | 1,103 | 1,185 |

Most residue sits in **source** and **NLD**; source cleaning scripts (`07`/`08`) target audit failures before modeling.

---

## Model architecture

Artifact: **`sban_weighted_stacking_model.joblib`** (`bundle_version: 1`, trained with **scikit-learn 1.6.1**).

```text
For each r ∈ {asm, binary, source, nld}:
  text → TF-IDF (binary: hex → byte tokens + instsep)
       + 6 numeric stats (length, tokens, entropy, …)
       → StandardScaler
       → sparse hstack → column subset (selected_indices from ID pruning)
       → LinearSVC (tuned class weights) → decision_function (4 scores)

Meta:
  hstack(all base decision scores + all scaled numeric blocks)
  → HistGradientBoostingClassifier
  → class probabilities
```

Bundle keys: `representation_order`, `representation_columns`, `numeric_feature_names`, `label_encoder`, `meta_model`, `representations` (vectorizer, scaler, indices, base model), `selected_class_weight_configs`, metadata.

Training details and ablations: **`svc_sban.ipynb`**.

### Feature pruning (TF-IDF columns)

Implemented in **`svc_sban.ipynb`** (cells after the first per-representation `LinearSVC` bases):

1. **Importance** — For each representation, mean `|coef_|` over classes from `final_base_models` (TF-IDF tokens + six numeric stats).
2. **Sort ascending** — Lowest-importance names are dropped first.
3. **Ratio sweep** — Validation macro-F1 was plotted for many removal ratios (roughly **5–60%** and **65–80%** in the analysis figures); the exported model uses a single setting.
4. **Production choice****`feature_pruning_ratio = 0.65`**: remove the lowest **65%** of the ranked feature list for TF-IDF vocabulary entries. The six numeric columns (`char_count`, `line_count`, `token_count`, `avg_line_length`, `unique_token_ratio`, `char_entropy`) are **always kept** and re-appended via fixed column indices after TF-IDF subsetting.

Validation macro-F1 at **65%** feature removal (same notebook run):

| Representation | Macro F1 (val) | Columns after prune |
|----------------|----------------|---------------------|
| asm | 0.7009 | 26,259 |
| binary | 0.6609 | 26,259 |
| nld | 0.4870 | 26,259 |
| source | 0.8802 | 26,257 |

These pruned column sets are stored in the joblib bundle as `representations[r]["selected_indices"]` (feature step only; ID pruning below may reuse the same index vector).

### ID pruning (training samples)

Overlapping **bodmas** vs **sorel20m** IDs motivate dropping ambiguous training rows before refitting bases:

1. Fix **feature pruning at 65%** and fit a temporary `LinearSVC` on pruned features.
2. **Score** each training row in **`bodmas`** and **`sorel20m`** only: sparse TF-IDF presence (binary) dotted with pruned-model TF-IDF coefficient magnitudes → `importance_score`.
3. **Grid** — For each representation, remove the lowest-scoring **`id_prune_ratios`** fraction **per class** (5%, 10%, …, 70%), refit on remaining train rows, measure validation macro-F1 → `id_pruning_summary` in the notebook.
4. **Production choice**`selected_id_prune_ratios`:

| Representation | ID remove ratio | Val macro F1 | Train rows kept | Removed bodmas / sorel20m |
|----------------|-----------------|--------------|-----------------|---------------------------|
| asm | **10%** | 0.6978 | 102,538 | 5,689 / 4,941 |
| binary | **5%** | 0.6625 | 107,854 | 2,844 / 2,470 |
| source | **40%** | 0.8712 | 70,646 | 22,756 / 19,766 |
| nld | **40%** | 0.4804 | 70,646 | 22,756 / 19,766 |

Final stacking retrains ID-pruned bases (5-fold OOF decision scores), then class-weight search and meta learner on top of that pipeline. **`dike`** and **`malwarebazaar`** rows are never removed by this step.

### Split sizes used in training notebook

| Split | Rows | bodmas | dike | malwarebazaar | sorel20m |
|-------|------|--------|------|---------------|----------|
| Train | 113,168 | 56,892 | 3,267 | 3,594 | 49,415 |
| Validation | 16,167 | 8,128 | 467 | 513 | 7,059 |
| Test | 32,334 | 16,255 | 933 | 1,027 | 14,119 |

(Test counts from `inference.ipynb` evaluation on exported bundle.)

### Base models on validation (`svc_sban.ipynb`)

Single-representation `LinearSVC` decision scores, validation set:

| Representation | Accuracy | Macro F1 | Weighted F1 |
|----------------|----------|----------|---------------|
| asm | 0.8983 | 0.7099 | 0.8902 |
| binary | 0.8426 | 0.6607 | 0.8345 |
| source | 0.9253 | 0.8811 | 0.9251 |
| nld | 0.6621 | 0.4997 | 0.6539 |

**Source** is the strongest single view; **nld** alone is weakest but adds complementary signal in the stack.

### Final exported model — validation & test

Metrics from **`inference.ipynb`** with `sban_weighted_stacking_model.joblib` (matches weighted meta validation in `svc_sban.ipynb` before export).

**Validation (n = 16,167)**

| | Accuracy | Macro F1 | Weighted F1 |
|---|----------|----------|-------------|
| Overall | **0.9413** | **0.9097** | **0.9412** |

| Class | Precision | Recall | F1 | Support |
|-------|-----------|--------|-----|---------|
| bodmas | 0.9551 | 0.9398 | 0.9474 | 8,128 |
| dike | 0.9125 | 0.8266 | 0.8674 | 467 |
| malwarebazaar | 0.8986 | 0.8635 | 0.8807 | 513 |
| sorel20m | 0.9306 | 0.9562 | 0.9433 | 7,059 |

**Test (n = 32,334)**

| | Accuracy | Macro F1 | Weighted F1 |
|---|----------|----------|-------------|
| Overall | **0.9379** | **0.9012** | **0.9378** |

| Class | Precision | Recall | F1 | Support |
|-------|-----------|--------|-----|---------|
| bodmas | 0.9532 | 0.9364 | 0.9447 | 16,255 |
| dike | 0.8909 | 0.8489 | 0.8694 | 933 |
| malwarebazaar | 0.8885 | 0.8150 | 0.8502 | 1,027 |
| sorel20m | 0.9272 | 0.9545 | 0.9407 | 14,119 |

Minority classes (`dike`, `malwarebazaar`) remain the hardest; weighted class tuning in `svc_sban.ipynb` targets that imbalance.

---

## Inference schema

| Column | Required for predict | Notes |
|--------|-------------------|--------|
| `assembly_code`, `binary_code`, `source_code`, `NLD` | Yes | |
| `ID` | No | Preserved in output |
| `dataset_name` | No | For `--evaluate` / notebook metrics |

---

## Installation

```bash
pip install -r requirements-inference.txt   # predict only
pip install -r requirements.txt             # full pipeline + notebooks
```

Use **scikit-learn 1.6.1** when loading the joblib bundle.

---

## Running inference

```bash
python inference.py \
  --model-path sban_weighted_stacking_model.joblib \
  --input /path/to/SBAN_test.parquet \
  --output predictions.parquet \
  --evaluate
```

```python
from inference import load_predictor
import pandas as pd

predictor = load_predictor("sban_weighted_stacking_model.joblib")
out = predictor.predict(pd.read_parquet("/path/to/samples.parquet"))
```

**Notebook:** `inference.ipynb` — Colab or local setup → demo row → validation/test cells (update parquet paths).

---

## Reproducing the production model

1. Obtain SBAN and build parquets via `01``05` (and cleaning/audit scripts as needed).
2. Open **`svc_sban.ipynb`** (Colab or local), point to `SBAN_train/val/test.parquet`.
3. Run training cells; export **`sban_weighted_stacking_model.joblib`** to the repo root.
4. Verify with **`inference.py`** or **`inference.ipynb`**.

---

## Citation and security

- Cite the **SBAN** dataset authors; this repo does not redistribute their files.
- **`joblib.load` uses pickle** — only load bundles from this project or your own exports.