How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("text-classification", model="oddadmix/dialect-router-v0.2")
# Load model directly
from transformers import AutoTokenizer, AutoModelForSequenceClassification

tokenizer = AutoTokenizer.from_pretrained("oddadmix/dialect-router-v0.2")
model = AutoModelForSequenceClassification.from_pretrained("oddadmix/dialect-router-v0.2", device_map="auto")
Quick Links

dialect-router-v0.2

A lightweight Arabic dialect identification model that classifies input text into one of 15 language codes: 13 Arabic dialects, Modern Standard Arabic, and English. It is the routing backbone in the Lahgtna pipeline, automatically selecting the correct voice reference and Chatterbox language token for speech synthesis.

v0.2 is a fine-tune of asafaya/bert-mini-arabic and expands coverage from 10 to 13 Arabic dialects, and adds an English label.

Model Details

Property Value
Base model asafaya/bert-mini-arabic
Architecture BERT-mini encoder + sequence classification head
Task Multi-class text classification (15 classes)
Input Raw text (up to 512 tokens)
Output One of 15 dialect / language codes
Languages Arabic (ar), English (en)
License MIT

Evaluation Results

Metric Score
Accuracy 0.9359
F1 Macro 0.9052
Eval Loss 0.4537

Dialect Labels

ID Label Dialect / Language Region
0 ar Modern Standard Arabic (MSA)
1 bh Bahraini Bahrain
2 dz Algerian Algeria
3 eg Egyptian Egypt
4 en English
5 iq Iraqi Iraq
6 lb Lebanese Lebanon
7 ly Libyan Libya
8 ma Moroccan (Darija) Morocco
9 ps Palestinian Palestine
10 sa Saudi Saudi Arabia
11 sd Sudanese Sudan
12 sy Syrian Syria
13 tn Tunisian Tunisia
14 ye Yemeni Yemen

What's New in v0.2

  • 13 Arabic dialects (up from 10): adds Bahraini (bh), Algerian (dz), and Yemeni (ye)
  • English label (en) — English input is now routed explicitly instead of being out-of-scope
  • Moroccan label renamed moma (ISO 3166 country code)
  • New base model: asafaya/bert-mini-arabic — smaller and faster for routing workloads
  • Retrained on an expanded multi-dialect corpus

Intended Use

Primary use Dialect-aware TTS routing — given an Arabic utterance, predict the dialect so the correct speaker reference audio and Chatterbox language code can be selected automatically.

Secondary use Standalone Arabic dialect identification for NLP pipelines, content filtering, dataset analysis, or any application that needs to distinguish Arabic dialects programmatically.

Out-of-scope use

  • Languages other than Arabic and English
  • Code-switched text (Arabic + English mixed)
  • Dialect intensity scoring or fine-grained subdialect classification
  • High-stakes decisions without human review

How to Use

Direct inference

from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch

model_id = "oddadmix/dialect-router-v0.2"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)
model.eval()

text = "اه ياراسي الواحد دماغه وجعاه"
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512)
with torch.no_grad():
    logits = model(**inputs).logits

pred_id = torch.argmax(logits, dim=-1).item()
dialect = model.config.id2label[pred_id]
print(dialect)  # e.g. "eg"

With the Transformers pipeline

from transformers import pipeline

classifier = pipeline(
    "text-classification",
    model="oddadmix/dialect-router-v0.2",
)
result = classifier("اه ياراسي الواحد دماغه وجعاه")
print(result)
# [{'label': 'eg', 'score': 0.94}]

Inside Lahgtna TTS

from inference import run_pipeline

# Dialect is detected automatically
run_pipeline(
    text="اه ياراسي الواحد دماغه وجعاه",
    output_path="output.wav",
)

Training Procedure

Hyperparameters

  • learning_rate: 3e-05
  • train_batch_size: 64
  • eval_batch_size: 64
  • seed: 42
  • optimizer: AdamW (betas=(0.9, 0.999), epsilon=1e-08)
  • lr_scheduler_type: cosine
  • num_epochs: 20
  • mixed_precision_training: Native AMP

Limitations & Biases

  • Short texts (< 5 tokens) may produce unreliable predictions — the model benefits from sentence-length input.
  • Code-switched text (e.g. Arabic + French in Maghrebi dialects, or Arabic + English) may confuse the classifier; heavily mixed input may be routed to en.
  • Dialect continuum — dialects from geographically adjacent regions (e.g. sy / lb / ps, ma / dz / tn, sa / bh) may be confused by the model.
  • Corpus bias — label distribution in training data may not reflect real-world dialect prevalence; some dialects (e.g. sd, ly, bh, ye) may have lower recall.
  • This model should not be used for identity classification of individuals.

Citation

@misc{lahgtna-dialect-router-2026,
  title  = {dialect-router-v0.2: Arabic Dialect Identification for TTS Routing},
  author = {Oddadmix},
  year   = {2026},
  url    = {https://huggingface.co/oddadmix/dialect-router-v0.2}
}
Downloads last month
18
Safetensors
Model size
11.6M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for oddadmix/dialect-router-v0.2

Finetuned
(5)
this model