| --- |
| language: |
| - ar |
| - fr |
| - en |
| license: mit |
| base_model: xlm-roberta-base |
| tags: |
| - text-classification |
| - comment-moderation |
| - darija |
| - code-switching |
| - nlp |
| metrics: |
| - accuracy |
| --- |
| |
| # Relevant / Irrelevant Comment Classifier |
|
|
| A fine-tuned `xlm-roberta-base` model that classifies social media comments as **relevant** or **irrelevant** to the topic they were posted under. |
|
|
| Trained on real-world telecom customer comments (French, Darija, and Arabic), where "irrelevant" covers off-topic chatter, spam, or unrelated remarks mixed in with genuine customer feedback. |
|
|
| ## Usage |
|
|
| ```python |
| from transformers import pipeline |
| |
| classifier = pipeline("text-classification", model="Chaima-KHENAFIF/relevant-comment-detector") |
| classifier("la connexion est très lente") |
| ``` |
|
|
| ## Training |
|
|
| Fine-tuned on ~2,000 labeled comments scraped from telecom operator social media pages, covering French, Darija (Latin and Arabic script), and Arabic. |
|
|
| ## Training results |
|
|
| | Epoch | Train Loss | Val Loss | Val Accuracy | |
| |---|---|---|---| |
| | 1 | 0.5005 | 0.2117 | 93.56% | |
| | 2 | 0.2171 | 0.1577 | **95.05%** | |
| | 3 | 0.1743 | 0.1530 | 94.06% | |
| | 4 | 0.1322 | 0.1546 | 93.56% | |
|
|
| Best validation accuracy of **95.05%** reached at epoch 2. |
|
|
| ## Links |
|
|
| - Full code and training pipeline: [GitHub repo](https://github.com/chaima-Khenafif03/relevant-comment-detector) |