---
license: mit
tags:
- phishing-detection
- url-classification
---
# Overview
The objective of the project is to classify if an URL is phishing or not. This model repo contains the required encoders (for url,dom and tld), scaler (for digit_cnt and is_https) and the trained model (RandomForest Classifier).
# Dataset & Attribution
This project uses the URL-Phish dataset. The dataset was obtained from Kaggle, where it is available as [Phishing URL Detection (111K URLs, 22 Features)](https://www.kaggle.com/datasets/sahandnamvar/phishing-url-detection-111k-urls-22-features).
The dataset is licensed under **[Creative Commons Attribution 4.0 International (CC BY 4.0)](https://creativecommons.org/licenses/by/4.0/)**,
which permits sharing, redistribution, and adaptation with appropriate credit.
**Dataset citation**
>Dam Minh, Linh; Tran Cong, Hung (2025).
>URL-Phish: A Feature-Engineered Dataset for Phishing Detection.
>Mendeley Data, V1.
>DOI: https://doi.org/10.17632/65z9twcx3r.1
**Original data sources referenced by the dataset authors**
>PhishTank – Community-driven phishing URL repository
>Research Organization Registry (ROR) dataset – Source of trusted benign domain URLs
**Paper citation**
>Dam Minh Linh, Tran Cong Hung,
>A feature-engineered dataset of benign and phishing URLs for machine learning and large language models evaluation,
>Data in Brief,
>Volume 63,
>2025,
>112162,
>ISSN 2352-3409,
>https://doi.org/10.1016/j.dib.2025.112162.
**Modifications:**
The following preprocessing was applied to the original dataset:
- Duplicate rows and null/missing values were checked for and removed, if present
- Feature scaling applied to selected numeric features
- TF-IDF encoding applied to selected URL/text-derived feature(s)
- Data split into train / validation / test sets
**Feature usage:**
The final model was trained using a selected subset
of the features; the remaining
features were excluded at training time via feature selection, not by removing them from
the stored datasets.
# License
- **Code**: MIT License — see `LICENSE`
- **Data**: Raw and processed datasets are redistributed under **Creative Commons Attribution 4.0 International (CC BY 4.0) license**, consistent with the original dataset's license (see Dataset & Attribution above).
- **Model & preprocessors**: MIT License — trained artifacts are provided under the same
terms as the codebase. Model trained on **Creative Commons Attribution 4.0 International (CC BY 4.0) license** data; see Dataset & Attribution section for details.