Instructions to use oddadmix/dialect-router-v0.2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use oddadmix/dialect-router-v0.2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="oddadmix/dialect-router-v0.2")# Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("oddadmix/dialect-router-v0.2") model = AutoModelForSequenceClassification.from_pretrained("oddadmix/dialect-router-v0.2", device_map="auto") - Notebooks
- Google Colab
- Kaggle
# Load model directly
from transformers import AutoTokenizer, AutoModelForSequenceClassification
tokenizer = AutoTokenizer.from_pretrained("oddadmix/dialect-router-v0.2")
model = AutoModelForSequenceClassification.from_pretrained("oddadmix/dialect-router-v0.2", device_map="auto")dialect-router-v0.2
A lightweight Arabic dialect identification model that classifies input text into one of 15 language codes: 13 Arabic dialects, Modern Standard Arabic, and English. It is the routing backbone in the Lahgtna pipeline, automatically selecting the correct voice reference and Chatterbox language token for speech synthesis.
v0.2 is a fine-tune of asafaya/bert-mini-arabic and expands coverage from 10 to 13 Arabic dialects, and adds an English label.
Model Details
| Property | Value |
|---|---|
| Base model | asafaya/bert-mini-arabic |
| Architecture | BERT-mini encoder + sequence classification head |
| Task | Multi-class text classification (15 classes) |
| Input | Raw text (up to 512 tokens) |
| Output | One of 15 dialect / language codes |
| Languages | Arabic (ar), English (en) |
| License | MIT |
Evaluation Results
| Metric | Score |
|---|---|
| Accuracy | 0.9359 |
| F1 Macro | 0.9052 |
| Eval Loss | 0.4537 |
Dialect Labels
| ID | Label | Dialect / Language | Region |
|---|---|---|---|
| 0 | ar | Modern Standard Arabic (MSA) | — |
| 1 | bh | Bahraini | Bahrain |
| 2 | dz | Algerian | Algeria |
| 3 | eg | Egyptian | Egypt |
| 4 | en | English | — |
| 5 | iq | Iraqi | Iraq |
| 6 | lb | Lebanese | Lebanon |
| 7 | ly | Libyan | Libya |
| 8 | ma | Moroccan (Darija) | Morocco |
| 9 | ps | Palestinian | Palestine |
| 10 | sa | Saudi | Saudi Arabia |
| 11 | sd | Sudanese | Sudan |
| 12 | sy | Syrian | Syria |
| 13 | tn | Tunisian | Tunisia |
| 14 | ye | Yemeni | Yemen |
What's New in v0.2
- 13 Arabic dialects (up from 10): adds Bahraini (
bh), Algerian (dz), and Yemeni (ye) - English label (
en) — English input is now routed explicitly instead of being out-of-scope - Moroccan label renamed
mo→ma(ISO 3166 country code) - New base model:
asafaya/bert-mini-arabic— smaller and faster for routing workloads - Retrained on an expanded multi-dialect corpus
Intended Use
Primary use Dialect-aware TTS routing — given an Arabic utterance, predict the dialect so the correct speaker reference audio and Chatterbox language code can be selected automatically.
Secondary use Standalone Arabic dialect identification for NLP pipelines, content filtering, dataset analysis, or any application that needs to distinguish Arabic dialects programmatically.
Out-of-scope use
- Languages other than Arabic and English
- Code-switched text (Arabic + English mixed)
- Dialect intensity scoring or fine-grained subdialect classification
- High-stakes decisions without human review
How to Use
Direct inference
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
model_id = "oddadmix/dialect-router-v0.2"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)
model.eval()
text = "اه ياراسي الواحد دماغه وجعاه"
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512)
with torch.no_grad():
logits = model(**inputs).logits
pred_id = torch.argmax(logits, dim=-1).item()
dialect = model.config.id2label[pred_id]
print(dialect) # e.g. "eg"
With the Transformers pipeline
from transformers import pipeline
classifier = pipeline(
"text-classification",
model="oddadmix/dialect-router-v0.2",
)
result = classifier("اه ياراسي الواحد دماغه وجعاه")
print(result)
# [{'label': 'eg', 'score': 0.94}]
Inside Lahgtna TTS
from inference import run_pipeline
# Dialect is detected automatically
run_pipeline(
text="اه ياراسي الواحد دماغه وجعاه",
output_path="output.wav",
)
Training Procedure
Hyperparameters
- learning_rate: 3e-05
- train_batch_size: 64
- eval_batch_size: 64
- seed: 42
- optimizer: AdamW (betas=(0.9, 0.999), epsilon=1e-08)
- lr_scheduler_type: cosine
- num_epochs: 20
- mixed_precision_training: Native AMP
Limitations & Biases
- Short texts (< 5 tokens) may produce unreliable predictions — the model benefits from sentence-length input.
- Code-switched text (e.g. Arabic + French in Maghrebi dialects, or Arabic + English) may confuse the classifier; heavily mixed input may be routed to
en. - Dialect continuum — dialects from geographically adjacent regions (e.g. sy / lb / ps, ma / dz / tn, sa / bh) may be confused by the model.
- Corpus bias — label distribution in training data may not reflect real-world dialect prevalence; some dialects (e.g. sd, ly, bh, ye) may have lower recall.
- This model should not be used for identity classification of individuals.
Citation
@misc{lahgtna-dialect-router-2026,
title = {dialect-router-v0.2: Arabic Dialect Identification for TTS Routing},
author = {Oddadmix},
year = {2026},
url = {https://huggingface.co/oddadmix/dialect-router-v0.2}
}
- Downloads last month
- 18
Model tree for oddadmix/dialect-router-v0.2
Base model
asafaya/bert-mini-arabic
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="oddadmix/dialect-router-v0.2")