Delete README.md
Browse files
README.md
DELETED
|
@@ -1,129 +0,0 @@
|
|
| 1 |
-
|
| 2 |
-
|
| 3 |
-
# Multilingual Sentiment Analysis
|
| 4 |
-
|
| 5 |
-
An end-to-end NLP project for three-class sentiment classification: **negative**, **neutral**, and **positive**. The repository contains two XLM-RoBERTa experiments:
|
| 6 |
-
|
| 7 |
-
1. an archived English airline-tweet model, and
|
| 8 |
-
2. the current general multilingual Twitter-sentiment model.
|
| 9 |
-
|
| 10 |
-
The current model is exposed through a Gradio app and is evaluated both overall and separately for every supported language.
|
| 11 |
-
|
| 12 |
-
## 🎮 Demo: Gradio App
|
| 13 |
-
|
| 14 |
-
Run the interactive demo locally:
|
| 15 |
-
```bash
|
| 16 |
-
python app.py
|
| 17 |
-
```
|
| 18 |
-
|
| 19 |
-

|
| 20 |
-
|
| 21 |
-
## Results at a glance
|
| 22 |
-
|
| 23 |
-
| Experiment | Data | Evaluation | Result |
|
| 24 |
-
| --- | --- | --- | --- |
|
| 25 |
-
| Airline tweets (legacy) | English airline tweets | Validation set | 84.24% accuracy, 0.842 weighted F1 |
|
| 26 |
-
| General multilingual (current) | 8-language Twitter sentiment dataset | Held-out test set (6,960 tweets) | 66.52% accuracy, 0.663 macro F1 |
|
| 27 |
-
|
| 28 |
-
The two scores are **not directly comparable**: the airline model was evaluated on a narrow, English-only domain, whereas the multilingual model is evaluated across eight languages and broader Twitter content.
|
| 29 |
-
|
| 30 |
-
### Current multilingual model: confusion matrix
|
| 31 |
-
|
| 32 |
-
<img src="results/multilingual/eval/confusion_matrix.png" width="500">
|
| 33 |
-
|
| 34 |
-
### Airline-tweet model: evaluation plots
|
| 35 |
-
|
| 36 |
-
<img src="results/airline_tweets/eval/confusion_matrix.png" width="400">
|
| 37 |
-
|
| 38 |
-
<img src="results/airline_tweets/eval/loss_curve.png" width="400">
|
| 39 |
-
|
| 40 |
-
|
| 41 |
-
### Current multilingual model: test macro F1 by language
|
| 42 |
-
|
| 43 |
-
| Language | Macro F1 |
|
| 44 |
-
| --- | ---: |
|
| 45 |
-
| German | 0.731 |
|
| 46 |
-
| French | 0.721 |
|
| 47 |
-
| Portuguese | 0.699 |
|
| 48 |
-
| English | 0.692 |
|
| 49 |
-
| Spanish | 0.667 |
|
| 50 |
-
| Arabic | 0.647 |
|
| 51 |
-
| Italian | 0.625 |
|
| 52 |
-
| Hindi | 0.516 |
|
| 53 |
-
|
| 54 |
-
|
| 55 |
-
The neutral class is the most difficult overall (0.588 F1). Hindi is the weakest supported language, so this project does not claim equal quality across all languages.
|
| 56 |
-
|
| 57 |
-
## How it works
|
| 58 |
-
|
| 59 |
-
1. `multilingual_sentiment_analysis.preprocess` downloads and prepares the multilingual dataset.
|
| 60 |
-
2. Text is cleaned by removing URLs and mentions, normalizing whitespace, and removing only the `#` marker. Unicode text, accents, emojis, and non-Latin scripts are preserved.
|
| 61 |
-
3. `multilingual_sentiment_analysis.train` fine-tunes `xlm-roberta-base` for three-class sequence classification.
|
| 62 |
-
4. `multilingual_sentiment_analysis.evaluate` reports overall metrics, per-language metrics, and a confusion matrix.
|
| 63 |
-
5. `app.py` provides single-text and batch inference with the saved model.
|
| 64 |
-
|
| 65 |
-
## Dataset
|
| 66 |
-
|
| 67 |
-
The current experiment uses [Cardiff NLP's Tweet Sentiment Multilingual dataset](https://huggingface.co/datasets/cardiffnlp/tweet_sentiment_multilingual). It provides the same label mapping used by this project:
|
| 68 |
-
|
| 69 |
-
```text
|
| 70 |
-
0 = negative
|
| 71 |
-
1 = neutral
|
| 72 |
-
2 = positive
|
| 73 |
-
```
|
| 74 |
-
|
| 75 |
-
It includes Arabic, English, French, German, Hindi, Italian, Portuguese, and Spanish. The preparation script loads each language configuration separately so it can preserve the official train/validation/test split and retain the language for evaluation.
|
| 76 |
-
|
| 77 |
-
## Repository layout
|
| 78 |
-
|
| 79 |
-
```text
|
| 80 |
-
src/
|
| 81 |
-
multilingual_sentiment_analysis/
|
| 82 |
-
config.py # shared paths, labels, and dataset settings
|
| 83 |
-
preprocess.py # dataset download, cleaning, tokenization, and saving
|
| 84 |
-
train.py # fine-tuning entry point
|
| 85 |
-
evaluate.py # test and per-language evaluation
|
| 86 |
-
infer.py # lazy model loading and inference helpers
|
| 87 |
-
tests/ # regression tests for UI and preprocessing behavior
|
| 88 |
-
data/
|
| 89 |
-
airline_tweets/ # archived English-only source data and prepared splits
|
| 90 |
-
multilingual/ # prepared multilingual splits
|
| 91 |
-
results/
|
| 92 |
-
airline_tweets/ # archived English-only checkpoints
|
| 93 |
-
multilingual/ # current checkpoints and evaluation outputs
|
| 94 |
-
sentiment_model/
|
| 95 |
-
airline-tweets-sentiment-model/ # archived English-only model
|
| 96 |
-
general-multilingual-sentiment-model/ # current app model
|
| 97 |
-
```
|
| 98 |
-
|
| 99 |
-
Data, checkpoints, and model weights are intentionally ignored by Git.
|
| 100 |
-
|
| 101 |
-
|
| 102 |
-
## Training configuration
|
| 103 |
-
|
| 104 |
-
- Base checkpoint: `xlm-roberta-base`
|
| 105 |
-
- Maximum token length: 128
|
| 106 |
-
- Epochs: 3
|
| 107 |
-
- Learning rate: 2e-5 with the Trainer's linear schedule
|
| 108 |
-
- Train/evaluation batch sizes: 16 / 32
|
| 109 |
-
- Best checkpoint criterion: weighted F1 on the validation set
|
| 110 |
-
- Random seed: 42
|
| 111 |
-
|
| 112 |
-
The multilingual training run completed in approximately 4 minutes 40 seconds on an NVIDIA RTX 5060 Laptop GPU.
|
| 113 |
-
|
| 114 |
-
## Limitations and next steps
|
| 115 |
-
|
| 116 |
-
- The model was evaluated on Twitter-style text only; it is not validated for reviews, support tickets, or other domains.
|
| 117 |
-
- It is evaluated for eight languages, not all languages supported by the XLM-R tokenizer.
|
| 118 |
-
- Hindi and neutral-sentiment performance require further improvement.
|
| 119 |
-
- A useful next experiment is to initialize from a Twitter-adapted multilingual encoder such as XLM-T and compare per-language macro F1 against this baseline.
|
| 120 |
-
|
| 121 |
-
## Testing
|
| 122 |
-
|
| 123 |
-
```powershell
|
| 124 |
-
python -m pytest -q
|
| 125 |
-
```
|
| 126 |
-
|
| 127 |
-
## Attribution
|
| 128 |
-
|
| 129 |
-
The multilingual dataset is provided by Cardiff NLP. See its [dataset card](https://huggingface.co/datasets/cardiffnlp/tweet_sentiment_multilingual) and the associated [XLM-T paper](https://aclanthology.org/2022.lrec-1.27/).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|