Instructions to use mballouch/jev-darija with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Laya
How to use mballouch/jev-darija with Laya:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Laya Darija 25-domain classifier
A fine-tune of Laya multilingual, whose encoder is mmBERT-base, that assigns a Moroccan Darija message to one of 25 topic domains. It reads Arabic script, Arabizi and French-mixed text.
This is a closed-set classifier. Every input receives one of the 25 labels, and there is no "unknown" option. It does not generate replies or take actions.
| Same reserved test (9,147 messages, 25 labels) | Accuracy | Macro F1 |
|---|---|---|
| Original Laya multilingual (zero-shot) | 27.54% | 0.256 |
| Earlier fine-tune (last 4 encoder blocks, 1 epoch) | 39.79% | 0.390 |
| This model | 78.52% | 0.784 |
Usage
import json, laya
from huggingface_hub import hf_hub_download
repo = "mballouch/jev-darija"
agent = laya.load(repo, device="cuda")
questions = json.load(open(hf_hub_download(repo, "questions.json")))
result = agent.system_one("الويفي ما خدامش عندي من البارح", questions, max_len=2048, head_max_len=768)
print(result) # -> intent: Internet_and_Telecom
Use max_len=2048 and head_max_len=768, so the full 25-option question fits without truncation. questions.json contains the exact English instructions and domain descriptions used in training. Changing that wording changes the model's inputs.
Labels: Arts_and_Entertainment, Autos_and_Vehicles, Beauty_and_Fitness, Books_and_Literature, Business_and_Industrial, Computers_and_Electronics, Finance, Food_and_Drink, Games, Health, Hobbies_and_Leisure, Home_and_Garden, Internet_and_Telecom, Jobs_and_Education, Law_and_Government, News, Online_Communities, People_and_Society, Pets_and_Animals, Real_Estate, Science, Sensitive_Subjects, Shopping, Sports, Travel_and_Transportation.
Training
- Corpus:
atlasia/moroccan_darija_domain_classifier_dataset, revision76cfb771. Of its 188,848 rows, 1,616 empty, duplicate or contradictory rows were removed. The rest were split by near-duplicate group into 159,140 train / 9,363 validation / 9,363 calibration / 9,366 test rows, with no group crossing a split. - Extra data: 4,146 Darija messages from an earlier version of
mballouch/jev-darija-synthetic, filtered against every corpus split. - Recipe: continued from an earlier fine-tune. All 22 encoder blocks plus the decision layers are trained (125.1M parameters); embeddings, final norm and act head are frozen. Two epochs, encoder LR 2e-5, head LR 5e-5, cosine decay with 3% warm-up, label smoothing 0.1, effective batch 32, BF16 with gradient checkpointing, and shuffled option order.
- Selection and calibration: the checkpoint was chosen by validation macro F1. Temperature 0.886 was fitted on the calibration split. The test split was used once, after training.
- Hardware: one RTX 5060 Laptop GPU (8 GB). The run took 4.3 hours, with 7.34 GiB peak reserved memory.
Evaluation
- Corpus test: 78.52% accuracy and 0.784 macro F1 over 9,147 rows. The weakest domains are
People_and_Society(F1 0.64) andHobbies_and_Leisure(0.66); the strongest areScience(0.94),Books_and_Literature(0.91) andPets_and_Animals(0.91).
Limitations
- The source labels are noisy. The corpus is synthetic and unvalidated; for example, a taxi ride labelled
Autos_and_Vehicles. Test accuracy therefore measures agreement with noisy labels. Many confident "errors" are cases where the model is right and the label is wrong. - Arabizi is weaker than Arabic script.
- Neighbouring domains are often confused: Computers ↔ Internet/Telecom, Travel ↔ Autos, and People_and_Society ↔ Sensitive_Subjects.
Licence note
The base model is Apache-2.0 and its mmBERT encoder is MIT.
Note: this is an experimental fan project, not a research or commercial project.
- Downloads last month
- -
Model tree for mballouch/jev-darija
Base model
convaiinnovations/laya