Laya Darija 25-domain classifier

A fine-tune of Laya multilingual, whose encoder is mmBERT-base, that assigns a Moroccan Darija message to one of 25 topic domains. It reads Arabic script, Arabizi and French-mixed text.

This is a closed-set classifier. Every input receives one of the 25 labels, and there is no "unknown" option. It does not generate replies or take actions.

Same reserved test (9,147 messages, 25 labels) Accuracy Macro F1
Original Laya multilingual (zero-shot) 27.54% 0.256
Earlier fine-tune (last 4 encoder blocks, 1 epoch) 39.79% 0.390
This model 78.52% 0.784

Usage

import json, laya
from huggingface_hub import hf_hub_download

repo = "mballouch/jev-darija"
agent = laya.load(repo, device="cuda")
questions = json.load(open(hf_hub_download(repo, "questions.json")))

result = agent.system_one("الويفي ما خدامش عندي من البارح", questions, max_len=2048, head_max_len=768)
print(result)  # -> intent: Internet_and_Telecom

Use max_len=2048 and head_max_len=768, so the full 25-option question fits without truncation. questions.json contains the exact English instructions and domain descriptions used in training. Changing that wording changes the model's inputs.

Labels: Arts_and_Entertainment, Autos_and_Vehicles, Beauty_and_Fitness, Books_and_Literature, Business_and_Industrial, Computers_and_Electronics, Finance, Food_and_Drink, Games, Health, Hobbies_and_Leisure, Home_and_Garden, Internet_and_Telecom, Jobs_and_Education, Law_and_Government, News, Online_Communities, People_and_Society, Pets_and_Animals, Real_Estate, Science, Sensitive_Subjects, Shopping, Sports, Travel_and_Transportation.

Training

  • Corpus: atlasia/moroccan_darija_domain_classifier_dataset, revision 76cfb771. Of its 188,848 rows, 1,616 empty, duplicate or contradictory rows were removed. The rest were split by near-duplicate group into 159,140 train / 9,363 validation / 9,363 calibration / 9,366 test rows, with no group crossing a split.
  • Extra data: 4,146 Darija messages from an earlier version of mballouch/jev-darija-synthetic, filtered against every corpus split.
  • Recipe: continued from an earlier fine-tune. All 22 encoder blocks plus the decision layers are trained (125.1M parameters); embeddings, final norm and act head are frozen. Two epochs, encoder LR 2e-5, head LR 5e-5, cosine decay with 3% warm-up, label smoothing 0.1, effective batch 32, BF16 with gradient checkpointing, and shuffled option order.
  • Selection and calibration: the checkpoint was chosen by validation macro F1. Temperature 0.886 was fitted on the calibration split. The test split was used once, after training.
  • Hardware: one RTX 5060 Laptop GPU (8 GB). The run took 4.3 hours, with 7.34 GiB peak reserved memory.

Evaluation

  • Corpus test: 78.52% accuracy and 0.784 macro F1 over 9,147 rows. The weakest domains are People_and_Society (F1 0.64) and Hobbies_and_Leisure (0.66); the strongest are Science (0.94), Books_and_Literature (0.91) and Pets_and_Animals (0.91).

Limitations

  • The source labels are noisy. The corpus is synthetic and unvalidated; for example, a taxi ride labelled Autos_and_Vehicles. Test accuracy therefore measures agreement with noisy labels. Many confident "errors" are cases where the model is right and the label is wrong.
  • Arabizi is weaker than Arabic script.
  • Neighbouring domains are often confused: Computers ↔ Internet/Telecom, Travel ↔ Autos, and People_and_Society ↔ Sensitive_Subjects.

Licence note

The base model is Apache-2.0 and its mmBERT encoder is MIT.

Note: this is an experimental fan project, not a research or commercial project.

Downloads last month
-
Safetensors
Model size
0.3B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mballouch/jev-darija

Finetuned
(152)
this model

Datasets used to train mballouch/jev-darija