rishitpant's picture
Update README
62a0361
|
Raw
History Blame Contribute Delete
17.7 kB
metadata
title: System Threat Forecaster
emoji: πŸ›‘οΈ
colorFrom: blue
colorTo: red
sdk: docker
pinned: false

πŸ›‘οΈ System Threat Forecaster

A production-grade end-to-end ML system that predicts Windows malware infection from hardware and OS telemetry. Upload a CSV of system records and receive a binary threat classification per row β€” downloadable as submission.csv.

Live Demo β†’ rishitpant/system-threat-forecaster on Hugging Face Spaces


Table of Contents


Overview

System Threat Forecaster is a binary classification system built on a modified version of a malware prediction dataset with Windows device telemetry covering engine versions, OS metadata, hardware specs, antivirus signature data, and more. The goal is to predict target (0 = clean, 1 = malware-infected) for each system record.

This was a graded academic ML project with an associated viva. The dataset was intentionally modified from the original source, making it a hard problem: the top leaderboard score in the class was 64.95% accuracy. The judging metric was accuracy.

What this project covers end-to-end:

  • Raw data cleaning (duplicate removal, column pruning, missing-target handling)
  • Rich feature engineering β€” version splitting, date decomposition, categorical grouping, derived hardware features
  • Sklearn ColumnTransformer preprocessing with type-aware pipelines
  • 7-model baseline evaluation with full metrics (accuracy, F1, precision, recall, ROC-AUC, average precision)
  • Bayesian hyperparameter tuning via Optuna (TPE sampler, 20 trials, AUC-optimized) on the top 3 models
  • Recursive Feature Elimination with Cross-Validation (RFECV) on the best tuned model
  • Per-stage visual eval reports: confusion matrix, ROC curve, PR curve for every model at every stage
  • SHAP analysis β€” global feature importance, summary plot, dependence plots, and single-prediction waterfall explanations
  • MLflow experiment tracking β€” all runs logged with metrics, params, and stage tags; final model registered in MLflow Model Registry
  • Serialized artifacts (model.pkl, preprocessor.pkl) for reproducible inference
  • Flask web app with drag-and-drop CSV upload, prediction summary, and CSV download
  • Dockerized deployment on Hugging Face Spaces (port 7860)
  • GitHub Actions CI/CD: pytest on every push, auto-deploy to HF Space on green main.

Project Structure

SystemThreatForecaster/
β”‚
β”œβ”€β”€ app.py                            # Flask entry point (routes: /, /predict, /download)
β”œβ”€β”€ Dockerfile                        # python:3.11-slim, exposes port 7860
β”œβ”€β”€ requirements.txt
β”œβ”€β”€ setup.py                          # Installable package: threatforecaster v0.0.1
β”‚
β”œβ”€β”€ src/
β”‚   β”œβ”€β”€ exception.py                  # CustomException with file + line traceback
β”‚   β”œβ”€β”€ logger.py                     # Timestamped file logger β†’ logs/<timestamp>.log
β”‚   β”œβ”€β”€ utils.py                      # save_object/load_object (dill), evaluate_models,
β”‚   β”‚                                 #   evaluate_final_model (3-panel plots), param I/O
β”‚   β”‚
β”‚   β”œβ”€β”€ components/
β”‚   β”‚   β”œβ”€β”€ data_cleaning.py          # Drops duplicates, unnecessary columns, null targets
β”‚   β”‚   β”œβ”€β”€ data_ingestion.py         # 80/20 train/val split β†’ artifacts/
β”‚   β”‚   β”œβ”€β”€ data_transformation.py    # Feature engineering + ColumnTransformer fit/save
β”‚   β”‚   β”œβ”€β”€ model_trainer.py          # Baseline β†’ Optuna tuning β†’ RFECV β†’ MLflow β†’ best model saved
β”‚   β”‚   └── model_pusher.py           # (reserved)
β”‚   β”‚
β”‚   └── pipeline/
β”‚       β”œβ”€β”€ train_pipeline.py         # Orchestrates full training run (clean β†’ ingest β†’ transform β†’ train)
β”‚       └── predict_pipeline.py       # Inference: FE β†’ preprocess β†’ RFECV mask β†’ predict
β”‚
β”œβ”€β”€ artifacts/
β”‚   β”œβ”€β”€ model.pkl                     # {model, selected (RFECV mask), name}
β”‚   β”œβ”€β”€ preprocessor.pkl              # Fitted ColumnTransformer
β”‚   β”œβ”€β”€ eval/
β”‚   β”‚   β”œβ”€β”€ baseline/                 # 7 Γ— 3-panel eval PNGs (confusion, ROC, PR)
β”‚   β”‚   β”œβ”€β”€ tuned/                    # 3 Γ— 3-panel eval PNGs (post-Optuna)
β”‚   β”‚   β”œβ”€β”€ final/                    # 1 Γ— 3-panel eval PNG (post-RFECV best model)
β”‚   β”‚   β”œβ”€β”€ shap_summary.png          # SHAP beeswarm β€” direction + magnitude per feature
β”‚   β”‚   β”œβ”€β”€ shap_bar.png              # Mean |SHAP| global importance
β”‚   β”‚   β”œβ”€β”€ shap_dependence_top3.png  # Feature value vs SHAP for top 3 features
β”‚   β”‚   β”œβ”€β”€ shap_waterfall_infected.png  # Single prediction explanation
β”‚   β”‚   └── feature_importance.png    # LightGBM gain + split count
β”‚   └── params/
β”‚       β”œβ”€β”€ lightgbm_best_params.json
β”‚       β”œβ”€β”€ xgboost_best_params.json
β”‚       └── random_forest_best_params.json
β”‚
β”œβ”€β”€ mlruns/                           # MLflow run data (auto-generated, not committed)
β”‚
β”œβ”€β”€ templates/
β”‚   β”œβ”€β”€ index.html                    # Drag-and-drop upload UI
β”‚   └── results.html                  # Prediction summary + 30-row preview + download
β”‚
β”œβ”€β”€ notebook/
β”‚   └── EDA.ipynb                     # 117-cell EDA + SHAP analysis notebook
β”‚
β”œβ”€β”€ tests/
β”‚   β”œβ”€β”€ test_app.py                   # Flask route tests
β”‚   β”œβ”€β”€ test_data_cleaning.py         # Unit tests for cleaning functions
β”‚   β”œβ”€β”€ test_data_transformation.py   # Unit tests for feature engineering
β”‚   └── test_predict_pipeline.py      # Integration tests (skipped if artifacts absent)
β”‚
β”œβ”€β”€ logs/                             # Timestamped runtime logs
β”‚
└── .github/
    └── workflows/
        └── cicd.yaml                 # CI: pytest on all pushes | CD: deploy to HF Space on main

ML Pipeline

Run the full pipeline end-to-end with a single command:

python src/pipeline/train_pipeline.py

This executes all four steps in sequence: clean β†’ ingest β†’ transform β†’ train.


Step 0 β€” Data Cleaning (data_cleaning.py)

Processes raw CSVs before ingestion:

  • Drops exact duplicate rows (165 found in training data)
  • Removes 15 irrelevant/redundant columns identified during EDA (MachineID, IsBetaUser, SMode, IsVirtualDevice, OSBuildLab, Processor, OSVersion, and others)
  • Drops rows with null target labels

Reads from notebook/data/train.csv and notebook/data/test.csv. Outputs train_eda_clean.csv and test_eda_clean.csv to notebook/data/.


Step 1 β€” Data Ingestion (data_ingestion.py)

Reads the cleaned CSVs and performs an 80/20 train/val split (random state 42). Writes artifacts/train.csv, artifacts/val.csv, and artifacts/test.csv.


Step 2 β€” Feature Engineering (data_transformation.py)

Applied identically at training and inference time via _apply_feature_engineering().

Version Splitting β€” EngineVersion, AppVersion, SignatureVersion, NumericOSVersion β†’ _Major, _Minor, _Build, _Revision numeric columns.

Date Decomposition β€” DateAS β†’ Malware_year/month/day/hour/minute; DateOS β†’ OS_year/month/day.

Categorical Grouping β€” high-cardinality string columns bucketed into clean groups:

Column Groups
MDC2FormFactor Desktop, Notebook, Tablet, Server
PrimaryDiskType HDD, SSD, Others
ChassisType Desktop, Notebook, Tablet, Others
PowerPlatformRole Desktop, Portable, Server, Others
OSEdition Core, Professional, Enterprise, Others
OSInstallType Upgrade, Clean, Others
AutoUpdateOptionsName Auto, Manual, Off, Unknown
LicenseActivationChannel Retail, Volume, OEM
FlightRing Retail, Insider, Disabled, Unknown

Derived Features:

Feature Formula
Days_since_OS_Installation Malware_day βˆ’ OS_day
Ram_per_core TotalPhysicalRAMMB / ProcessorCoreCount
Aspect_Ratio ResolutionH / ResolutionV
Pixel_Density (ResH Γ— ResV) / DiagonalInches
Primary_Disk_Allocated PrimaryDiskCapacityMB / SystemVolumeCapacityMB
Free_Disk_Space (SysVolCapacity βˆ’ PrimaryDiskCapacity) / PrimaryDiskCapacity

Preprocessing β€” A ColumnTransformer is fit on training data only and serialized to artifacts/preprocessor.pkl:

Column Type Detection Pipeline
Numerical float64/int64, non-binary, non-ID SimpleImputer(mean) β†’ MinMaxScaler
Binary exactly 2 unique values SimpleImputer(most_frequent) β†’ OrdinalEncoder
ID columns int/float with "ID"/"Identifier" in name SimpleImputer(most_frequent)
Categorical object dtype SimpleImputer(most_frequent) β†’ OneHotEncoder(handle_unknown='ignore')

Step 3 β€” Model Training (model_trainer.py)

Phase 1 β€” Baseline evaluation across 7 classifiers (default hyperparameters), scored on val accuracy, F1, precision, recall, ROC-AUC, and average precision. Eval reports (3-panel PNGs) saved for every model.

Context: This dataset was modified from its original source for an academic setting. The class leaderboard ceiling was 64.95% accuracy.

Model Val Accuracy
LightGBM 0.6204
XGBoost 0.6146
Random Forest 0.6090
AdaBoost 0.6035
Bagging 0.5724
Logistic Regression 0.5539
Decision Tree 0.5370

Phase 2 β€” Bayesian hyperparameter tuning (Optuna) on the top 3 models by AUC, using TPE sampler (20 trials, seed 42). Tuning maximizes ROC-AUC on the validation set. Best params are cached as JSON so re-runs skip re-tuning.

Model Key tuned params
LightGBM n_estimators, num_leaves, max_depth, learning_rate, min_child_samples, reg_lambda, colsample_bytree
XGBoost n_estimators, max_depth, learning_rate, subsample, colsample_bytree, reg_lambda, min_child_weight
Random Forest n_estimators, max_depth, min_samples_leaf, max_features

Phase 3 β€” RFECV feature selection on the best-tuned model (step=2, cv=5, scoring=roc_auc, n_jobs=-1). The boolean support mask is stored alongside the model for exact replay at inference time.

Phase 4 β€” Final model save. The winning model dict {model, selected, name} is written to artifacts/model.pkl. A floor of AUC β‰₯ 0.50 is enforced; training raises a CustomException if unmet.


Step 4 β€” Prediction Pipeline (predict_pipeline.py)

PredictPipeline.predict(df) mirrors the training path exactly:

  1. Load artifacts/preprocessor.pkl and artifacts/model.pkl
  2. Apply _apply_feature_engineering()
  3. Transform via the fitted ColumnTransformer
  4. Apply the RFECV boolean mask
  5. Return binary predictions + probability scores (0 = clean, 1 = infected)

Model Interpretability β€” SHAP Analysis

All SHAP analysis lives in notebook/EDA.ipynb (final section) and outputs are saved to artifacts/eval/.

SHAP (SHapley Additive exPlanations) assigns each feature a contribution value per prediction β€” more reliable than LightGBM's built-in importance because it's model-agnostic and accounts for feature interactions.

Key findings

The model is primarily driven by antivirus coverage and freshness. The top 3 features by mean |SHAP| are:

Feature Interpretation
NumAntivirusProductsInstalled Machines with 0 AV products get a SHAP push of ~+0.25 toward malware. Having even one product drops this sharply.
EngineVersion_Build Outdated AV engine build numbers correlate strongly with infection. Version splitting was key to surfacing this.
AntivirusConfigID Specific AV configurations have distinct risk profiles β€” certain configs push SHAP to +0.6 (very risky), others to -0.8 (very clean).

Other notable signals: low RAM (TotalPhysicalRAMMB) β†’ proxy for older/less-maintained hardware; IsGamer = 1 β†’ gaming machines often have AV disabled; non-standard OS install types (OSInstallType_Grouped_Others) β†’ higher risk than clean installs.

Plots generated

SHAP Summary (beeswarm) β€” each dot is one validation row. X-axis = SHAP value (positive = toward malware), colour = feature value (red = high, blue = low). Shows both direction and magnitude of each feature's influence.

SHAP Bar β€” mean |SHAP| across all rows. Global importance ranking comparable to but more reliable than gain importance.

Dependence plots (top 3) β€” feature value vs SHAP value, coloured by a second interacting feature. Shows whether relationships are linear, threshold-based, or non-monotonic.

Waterfall plot β€” single infected prediction explained step-by-step from the base rate (E[f(X)] = 0.017) to the final output (f(x) β‰ˆ 1.0). Example: a gaming machine with 0 AV products, non-standard OS install, and old engine build β€” every major risk factor aligned.


Experiment Tracking β€” MLflow

Every training run is tracked in MLflow under the Threat_Forecaster experiment. All three stages are logged with consistent metrics for side-by-side comparison.

Stage tag Run name pattern What's logged
baseline Baseline_{model} val_roc_auc, val_recall, val_precision, val_f1, val_ap
tuning Tuned_{model} same metrics + all hyperparameters via log_params
final Final_{model} same metrics + features_selected + hyperparameters + model artifact registered as Threat_Forecaster

Start the MLflow UI:

mlflow ui
# β†’ http://localhost:5000

Select any combination of runs across stages and click Compare to view metrics side by side.

The final model is registered in the MLflow Model Registry under Threat_Forecaster, enabling version tracking and stage promotion (Staging β†’ Production) if needed.

mlruns/ is auto-generated locally and excluded from version control via .gitignore.


Web Application

Flask app (app.py) with three routes:

Route Method Description
/ GET Drag-and-drop CSV upload form
/predict POST Accepts .csv, runs PredictPipeline, renders results with summary stats, 30-row preview, and confidence scores
/download POST Streams full predictions as submission.csv

Input: CSV matching the test data column layout. If no id column is present, a sequential index is used.

Output: id, target, confidence β€” one row per input record.


Getting Started

Prerequisites

  • Python 3.11+
  • Raw data at notebook/data/train.csv and notebook/data/test.csv

Install

pip install -r requirements.txt

# Or as an editable package
pip install -e .

Train (full pipeline)

python src/pipeline/train_pipeline.py

Runs cleaning β†’ ingestion β†’ transformation β†’ training and writes artifacts/model.pkl and artifacts/preprocessor.pkl.

View experiment results

mlflow ui
# β†’ http://localhost:5000

Run SHAP analysis

Open notebook/EDA.ipynb and run the SHAP Analysis section at the bottom. Requires trained artifacts in artifacts/.

Run the app

python app.py
# β†’ http://0.0.0.0:7860

Docker

docker build -t system-threat-forecaster .
docker run -p 7860:7860 system-threat-forecaster

The Dockerfile uses python:3.11-slim, installs libgomp1 (required by LightGBM), and serves on port 7860 to match Hugging Face Spaces.


CI/CD

GitHub Actions (.github/workflows/cicd.yaml) runs on every push:

test job (all branches):

  1. Checkout repo
  2. Setup Python 3.11
  3. Install requirements.txt
  4. Set PYTHONPATH to workspace root
  5. Run pytest -v

deploy job (main branch only, after tests pass):

  1. Checkout with full history and LFS
  2. Authenticate to Hugging Face via HF_TOKEN secret
  3. Force-push to rishitpant/system-threat-forecaster on HF Spaces

Testing

pytest -v
File Tests What they check
test_data_cleaning.py 6 Duplicate removal, column dropping (with and without missing cols), null target handling
test_data_transformation.py 5 Version column splitting, original columns dropped, date decomposition, RAM/core ratio, graceful handling of minimal input
test_predict_pipeline.py 4 Two outputs returned, output length matches input, predictions are binary {0,1}, probabilities in [0,1] β€” skipped if artifacts absent
test_app.py 3 Homepage 200 OK, no-file POST doesn't crash, wrong file type POST doesn't crash

Predict pipeline tests are skipped automatically in CI since artifacts/model.pkl and the cleaned CSVs are not committed to the repo. All other 14 tests run on every push.


Dependencies

pandas, numpy
scikit-learn==1.8.0
xgboost, lightgbm
optuna
mlflow
shap
flask
dill
seaborn, matplotlib
pytest

Author

Rishit Pant β€” rishitpant100@gmail.com