Update README.md
Browse files---
language:
- ar
- fr
- en
license: mit
base_model: xlm-roberta-base
tags:
- text-classification
- comment-moderation
- darija
- code-switching
- nlp
metrics:
- accuracy
---
# Relevant / Irrelevant Comment Classifier
A fine-tuned `xlm-roberta-base` model that classifies social media comments as **relevant** or **irrelevant** to the topic they were posted under.
Trained on real-world telecom customer comments (French, Darija, and Arabic), where "irrelevant" covers off-topic chatter, spam, or unrelated remarks mixed in with genuine customer feedback.
## Usage
```python
from transformers import pipeline
classifier = pipeline("text-classification", model="Chaima-KHENAFIF/relevant-comment-detector")
classifier("la connexion est très lente")
```
## Training
Fine-tuned on ~2,000 labeled comments scraped from telecom operator social media pages, covering French, Darija (Latin and Arabic script), and Arabic.
## Training results
| Epoch | Train Loss | Val Loss | Val Accuracy |
|---|---|---|---|
| 1 | 0.5005 | 0.2117 | 93.56% |
| 2 | 0.2171 | 0.1577 | **95.05%** |
| 3 | 0.1743 | 0.1530 | 94.06% |
| 4 | 0.1322 | 0.1546 | 93.56% |
Best validation accuracy of **95.05%** reached at epoch 2.
## Links
- Full code and training pipeline: [GitHub repo](https://github.com/chaima-Khenafif03/relevant-comment-detector)
|
@@ -1,3 +1,17 @@
|
|
| 1 |
---
|
| 2 |
license: mit
|
| 3 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
| 2 |
license: mit
|
| 3 |
+
language:
|
| 4 |
+
- ar
|
| 5 |
+
- en
|
| 6 |
+
- fr
|
| 7 |
+
metrics:
|
| 8 |
+
- accuracy
|
| 9 |
+
base_model:
|
| 10 |
+
- FacebookAI/xlm-roberta-base
|
| 11 |
+
tags:
|
| 12 |
+
- text-classification
|
| 13 |
+
- comment-moderation
|
| 14 |
+
- darija
|
| 15 |
+
- code-switching
|
| 16 |
+
- nlp
|
| 17 |
+
---
|