File size: 9,423 Bytes
7618f62
d62f97b
 
7618f62
d62f97b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3924eee
d62f97b
 
3924eee
d62f97b
 
3924eee
d62f97b
7618f62
d62f97b
 
 
 
 
 
 
 
 
 
 
 
 
 
3924eee
d62f97b
3924eee
d62f97b
3924eee
d62f97b
 
 
 
 
3924eee
d62f97b
 
 
3924eee
 
 
d62f97b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3924eee
 
 
d62f97b
 
 
 
 
 
 
3924eee
d62f97b
3924eee
d62f97b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3924eee
d62f97b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
---
language:
  - ar
license: apache-2.0
library_name: transformers
pipeline_tag: text-classification
tags:
  - arabic
  - toxicity
  - hate-speech
  - cyberbullying
  - content-moderation
  - offensive-language
  - profanity
  - multi-dialect
  - arabic-nlp
  - MARBERTv2
  - bert
datasets:
  - youssefreda9/HAYAA
metrics:
  - accuracy
  - f1
base_model: UBC-NLP/MARBERTv2
model-index:
  - name: Hayā (HAYAA)
    results:
      - task:
          type: text-classification
          name: Text Classification
        dataset:
          type: youssefreda9/HAYAA
          name: HAYAA
          split: test
        metrics:
          - type: accuracy
            value: 0.9784
            name: Accuracy
          - type: f1
            value: 0.9412
            name: F1 (Toxic)
          - type: f1
            value: 0.9845
            name: F1 (Safe)
---

# Hayā (حياء) — Arabic Toxic Content Classifier 🛡️

[![Dataset on HF](https://img.shields.io/badge/🤗_Dataset-Hayā-green)](https://huggingface.co/datasets/youssefreda9/HAYAA)
[![GitHub](https://img.shields.io/badge/GitHub-HAYAA-black?logo=github)](https://github.com/youssefreda10/HAYAA)

## Model Description

**Hayā** is a fine-tuned [UBC-NLP/MARBERTv2](https://huggingface.co/UBC-NLP/MARBERTv2) model for **binary Arabic toxicity classification** (Safe / Toxic). It is designed to detect offensive language, hate speech, cyberbullying, profanity, and other forms of toxic content across **all major Arabic dialects**.

Hayā is the core classifier in a larger **defense-in-depth content moderation system** that includes rule-based layers for explicit profanity and a Chrome extension that blurs toxic content in real time.

### Key Highlights

- 🎯 **97.84% accuracy** on a held-out test set of ~100K sentences
- 🌍 **Every major Arabic dialect**: Egyptian, Levantine, Gulf, Maghrebi, Iraqi, Sudanese, MSA
- 📊 Trained on **nearly 1 million examples (997K)** from 51 curated public Arabic datasets
- ⚡ Optimized for real-time inference via serverless API (Modal + FastAPI)
- 🧠 **V2 Update:** Enhanced with continued fine-tuning on 34K hard edge-cases to conquer implicit hate and subtle toxicity.

---

## Performance

Evaluated on a **held-out test set of 99,759 sentences** (stratified, zero data leakage):

| Metric | Score |
|--------|-------|
| **Accuracy** | **97.84%** |
| **F1 (Toxic)** | **94.12%** |
| **F1 (Safe)** | **98.45%** |

> **Note on the numbers:** Manual error analysis showed the model frequently *outperformed the original human annotations* — many counted "errors" were actually mislabels in the source datasets. Real-world performance on correctly-labeled data is therefore higher than the raw scores suggest.

---

## Intended Use

### Primary Use Cases
- ✅ Classifying Arabic text as **Safe** or **Toxic**
- ✅ Content moderation for Arabic social media, forums, and comment sections
- ✅ Child safety and parental control applications
- ✅ Research on Arabic hate speech and toxicity detection

### Out-of-Scope Use
- ❌ Generating toxic or hateful content
- ❌ Surveillance or profiling of individuals
- ❌ Final automated decisions without human review
- ❌ Non-Arabic text classification

---

## Usage

### Quick Start with Transformers Pipeline

```python
from transformers import pipeline

classifier = pipeline("text-classification", model="youssefreda9/HAYAA", top_k=None)

results = classifier("أنت إنسان رائع ومحترم")
print(results)
# [[{'label': 'Safe', 'score': 0.99}, {'label': 'Toxic', 'score': 0.01}]]

results = classifier("يا حمار أنت")
print(results)
# [[{'label': 'Toxic', 'score': 0.98}, {'label': 'Safe', 'score': 0.02}]]
```

### Direct Model Usage

```python
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch

tokenizer = AutoTokenizer.from_pretrained("youssefreda9/HAYAA")
model = AutoModelForSequenceClassification.from_pretrained("youssefreda9/HAYAA")

text = "كلامك جميل جدا"
inputs = tokenizer(text, return_tensors="pt", max_length=128, truncation=True, padding=True)

with torch.no_grad():
    outputs = model(**inputs)
    probs = torch.softmax(outputs.logits, dim=-1)
    pred = torch.argmax(probs, dim=-1).item()

labels = {0: "Safe", 1: "Toxic"}
print(f"Prediction: {labels[pred]} ({probs[0][pred]:.2%})")
```

### Batch Inference

```python
from transformers import pipeline

classifier = pipeline("text-classification", model="youssefreda9/HAYAA")

texts = [
    "صباح الخير يا أصدقاء",
    "أنت واطي ومحترمش حد",
    "الجو جميل النهارده",
]

results = classifier(texts)
for text, result in zip(texts, results):
    print(f"{result['label']} ({result['score']:.2%}): {text}")
```

---

## Training Details

### Base Model
- **Model**: [UBC-NLP/MARBERTv2](https://huggingface.co/UBC-NLP/MARBERTv2)
- **Architecture**: BERT-base (Arabic-specific pre-training)
- **Task**: Binary sequence classification (Safe / Toxic)

### Training Data
- **Dataset**: [youssefreda9/HAYAA](https://huggingface.co/datasets/youssefreda9/HAYAA)
- **Training examples**: ~798K
- **Validation examples**: ~100K
- **Sources**: 51 curated public Arabic hate-speech and abuse datasets
- **Dialects**: All major Arabic dialects + MSA

### Hyperparameters

| Parameter | Value |
|-----------|-------|
| Max sequence length | 128 |
| Training Strategy | 2-Stage (4 epochs base corpus + 3 epochs hard edge-cases) |
| Batch size | 16 (effective: 32 with gradient accumulation) |
| Learning rate | 2e-5 (Stage 1), 5e-6 (Stage 2) |
| Warmup ratio | 0.1 |
| Weight decay | 0.01 |
| FP16 | ✅ |
| Loss | Weighted cross-entropy (class imbalance) |
| Best model metric | F1 (weighted) |
| Optimizer | AdamW |
| Seed | 42 |

### Class Weighting
The training set is imbalanced (more Safe than Toxic). A **weighted cross-entropy loss** was used, with the Toxic class weight dynamically computed as `safe_count / toxic_count` to ensure the model doesn't under-predict toxicity.

---

## What It Detects

The **Toxic** label covers:

| Category | Examples |
|----------|----------|
| **Profanity** | Explicit swear words across all dialects |
| **Hate speech** | Incitement against groups based on religion, ethnicity, nationality |
| **Cyberbullying** | Personal insults, harassment, directed attacks |
| **Racism** | Racial slurs and discriminatory language |
| **Sexism** | Misogynistic and sexually explicit language |
| **Religious hate** | Blasphemy, sectarian attacks |
| **Obfuscated toxicity** | Intentional typos, spaced letters, homoglyphs |

### Dialects

| Dialect | Coverage |
|---------|----------|
| Egyptian | ✅ |
| Levantine (Syrian, Lebanese, Jordanian, Palestinian) | ✅ |
| Gulf (Saudi, Emirati, Kuwaiti, Bahraini, Omani, Qatari) | ✅ |
| Maghrebi (Moroccan, Algerian, Tunisian, Libyan) | ✅ |
| Iraqi | ✅ |
| Sudanese | ✅ |
| Modern Standard Arabic (MSA) | ✅ |

---

## Part of a Larger System

This model is one layer in the **Hayā defense-in-depth pipeline**:

| Layer | Function |
|-------|----------|
| **L0 — Sanitize** | Unicode normalization, homoglyph folding, emoji analysis |
| **L1 — Dictionary** | Context-aware instant matching (100% precision) |
| **L1.5 — De-obfuscation** | Resolves masked/spaced evasion attempts |
| **L2 — This Model** | Deep-learning classification for implicit toxicity |

The full pipeline ships as a **Chrome Extension (Manifest V3)** with PIN-protected parental controls. See the [GitHub repository](https://github.com/youssefreda10/HAYAA) for the complete system.

---

## Limitations & Biases

- **Platform bias**: Training data is predominantly from Twitter; performance may vary on other platforms (forums, messaging apps, gaming chat).
- **Annotation noise**: Despite extensive cleaning, some label noise from the 51 original source datasets may persist.
- **Dialect imbalance**: Egyptian and MSA are better represented than Sudanese or Yemeni dialects.
- **Max length**: Sequences longer than 128 tokens are truncated, which may affect classification of very long texts.
- **Evolving language**: Arabic internet slang and evasion tactics evolve constantly; periodic retraining is recommended.
- **Context limitations**: The model classifies individual texts in isolation; conversational context is not considered.

---

## Ethical Considerations

- This model is designed to **protect people** — especially children — from online abuse.
- It should be used as a **tool to assist human moderators**, not as a sole decision-maker.
- False positives (Safe text flagged as Toxic) can suppress legitimate speech; false negatives (Toxic text missed) can leave users exposed to harm.
- The model was trained on publicly available datasets with appropriate licenses.

---

## Citation

```bibtex
@misc{hayaa2026,
  title={Hayā: Fine-tuned MARBERTv2 for Multi-Dialect Arabic Toxicity Classification},
  author={Youssef Reda},
  year={2026},
  url={https://huggingface.co/youssefreda9/HAYAA},
}
```

---

## License

This model is released under the [Apache 2.0 License](https://www.apache.org/licenses/LICENSE-2.0).

---

## Links

- 🤗 **Dataset**: [youssefreda9/HAYAA](https://huggingface.co/datasets/youssefreda9/HAYAA)
- 💻 **GitHub**: [youssefreda10/HAYAA](https://github.com/youssefreda10/HAYAA)
- 🧠 **Base Model**: [UBC-NLP/MARBERTv2](https://huggingface.co/UBC-NLP/MARBERTv2)