NhanChanDeo
Upload Vietnamese Toxic Comment Classification project
8058cb3
|
Raw
History Blame Contribute Delete
2.48 kB
# DLNLP - Vietnamese Hate Speech Detection
This repository contains a local notebook workflow for Vietnamese hate speech detection.
## Project Structure
- `data/raw`: raw CSV files
- `data/processed`: cleaned CSV files generated by preprocessing
- `notebooks`: notebook pipeline
- `outputs/figures`: saved charts and confusion matrices
- `outputs/models`: trained model artifacts
- `outputs/results`: metrics, reports, and prediction files
## Dữ liệu đầu vào
The preprocessing notebook expects these files in `data/raw`:
- `train_raw.csv`
- `val_raw.csv`
- `test_raw.csv`
Each file should contain at least:
- `free_text`
- `label_id`
## Setup
Create and activate a virtual environment, then install dependencies:
```bash
python -m venv .venv
.venv\Scripts\activate
pip install -r requirements.txt
```
If you already have a working Python environment, installing from `requirements.txt` is enough.
## How to run locally
Run the notebooks in this order:
1. `notebooks/01_data_preprocessing.ipynb`
2. `notebooks/02_baseline_tfidf_svm.ipynb`
3. `notebooks/03_bilstm_train_evaluate.ipynb`
4. `notebooks/04_phobert_train_evaluate.ipynb`
5. `notebooks/05_results_error_analysis.ipynb`
## Demo giao diện
Sau khi đã có các file model trong `outputs/models`, chạy giao diện demo bằng:
```bash
streamlit run app.py
```
## Notes
- To make the repository clone-and-run for the demo app, keep the trained artifacts in `outputs/models`.
- Do not commit local environment folders such as `.venv`, notebook caches, or temporary training checkpoints.
- The notebooks resolve the project root automatically from the local workspace, so Colab Drive mounting is not required.
- Notebook 1 creates `data/processed` from the raw CSV files.
- Notebooks 2, 3, and 4 write outputs into `outputs/models`, `outputs/results`, and `outputs/figures`.
- Notebook 5 reads the artifacts from those local output folders and performs comparison/error analysis.
- For notebook 3 and notebook 4, the package list includes `pyvi`, `transformers`, `torch`, and `accelerate`.
## Expected outputs
After a full run, you should see files such as:
- `data/processed/train_processed.csv`
- `outputs/models/svm_model.joblib`
- `outputs/models/bilstm_best.pt`
- `outputs/models/phobert_base/`
- `outputs/results/svm_metrics.csv`
- `outputs/results/bilstm_metrics.csv`
- `outputs/results/phobert_metrics.csv`
- `outputs/results/phobert_test_predictions.csv`
- `outputs/figures/*.png`