Instructions to use mballouch/jev-darija-telecom-model with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Laya
How to use mballouch/jev-darija-telecom-model with Laya:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Jev-Darija Telecom Model
A classifier for Moroccan Darija telecom customer messages, in Arabic script, Arabizi, or Darija mixed with French. For each message it returns:
- category: one of 12 (billing, prepaid credit, mobile data, home internet, home line, network and calls, SIM and line, plans, devices, customer care, mobile money, out of scope)
- intent: one of 75 inside that category (e.g.
recharge_failed,home_outage,lost_stolen,mm_send_receive) - product: prepaid or postpaid mobile, mobile broadband, ADSL, fibre, 4G/5G box, landline, TV, mobile money, digital services, or unknown
- urgency: 0 to 3
- needs_human: whether the message should go to a human agent
It is a fine-tune of Laya multilingual (322M parameters), trained on mballouch/jev-darija-telecom. The taxonomy, with every definition, is in telecom_taxonomy.json.
Usage
pip install laya==0.3.28 transformers==4.57.6 huggingface_hub
from huggingface_hub import hf_hub_download
import importlib.util
spec = importlib.util.spec_from_file_location("predict", hf_hub_download("mballouch/jev-darija-telecom-model", "predict.py"))
predict = importlib.util.module_from_spec(spec)
spec.loader.exec_module(predict)
classifier = predict.TelecomClassifier("mballouch/jev-darija-telecom-model")
print(classifier.predict("chargit 20dh w ma wslatni walo, 3afak chofo lia"))
# {'category': 'prepaid_credit', 'intent': 'recharge_failed', 'product': 'prepaid_mobile', 'urgency': 2, 'needs_human': False, ...}
Or from the command line: python predict.py "الكونيكسيون مقطوعة من البارح".
predict.py asks the questions in two passes, as in training: category, product, urgency and needs_human first, then the intent question of the predicted category. The question wording must match telecom_taxonomy.json exactly, which predict.py takes care of. Each prediction takes about 30 ms on a laptop GPU.
Examples
| Message | Intent | Category | Urgency | Human |
|---|---|---|---|---|
| الكونيكسيون مقطوعة من البارح ف الدار والضو ديال الروتور حمر | home_outage |
home_internet | 2 | no |
| chargit 20dh w ma wslatni walo, 3afak chofo lia | recharge_failed |
prepaid_credit | 2 | no |
| tsre9 lia tel f tram, bghit nbloki la puce daba daba | lost_stolen |
sim_line | 3 | no |
| وصلني ميساج كيقول ليا ربحتي 5000 درهم صيفط الكود، واش هادا منكم؟ | scam_spam_report |
sim_line | 2 | yes |
| bghit nsift 300dh l mama f Fès b l wallet, kifach? | mm_send_receive |
mobile_money | 0 | no |
| غادي نسافر لإسبانيا، واش الخط غيخدم تما؟ | roaming |
plans_subscription | 0 | no |
| الضو مقطوع ف الحومة كاملة من الصباح | non_telecom |
out_of_scope | 0 | no |
Evaluation
On the 1,107-message test split of mballouch/jev-darija-telecom (synthetic messages written separately from the training data, with different writer personas and cities):
| Metric | Score |
|---|---|
| Intent, end to end (category and intent both right, 75 classes) | 84.6% |
| Category (12 classes) | 89.4% |
| Intent, given the right category | 93.3% |
| Product | 89.9% |
| Urgency (exact level) | 80.8% |
| Needs human | 95.3% |
Intent accuracy by script: Arabic 84.8%, Arabizi 82.9%, mixed 86.5%.
Training
- Starting point:
mballouch/jev-darija, a Laya multilingual model already fine-tuned on Moroccan Darija topic classification. - Data: 18,350 training messages; 10% held out, by intent, for epoch selection and confidence calibration.
- What was trained: all 22 encoder blocks and the decision layers (125M parameters); the token embeddings stayed frozen.
- Recipe: soft cross-entropy with 0.05 label smoothing, shuffled option order, AdamW (encoder learning rate 3e-5, head 1e-4), 6% warm-up then cosine decay, effective batch 32, fp16 on a Kaggle Tesla T4. The best of 8 epochs on validation (epoch 7) was kept.
- Calibration: confidence temperatures were fitted on the validation split.
- Downloads last month
- -
Model tree for mballouch/jev-darija-telecom-model
Base model
convaiinnovations/laya