Instructions to use TypicaAI/MagBERT-NER-AR with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use TypicaAI/MagBERT-NER-AR with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="TypicaAI/MagBERT-NER-AR")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("TypicaAI/MagBERT-NER-AR") model = AutoModelForTokenClassification.from_pretrained("TypicaAI/MagBERT-NER-AR", device_map="auto") - Notebooks
- Google Colab
- Kaggle
MagBERT-NER-AR — Arabic Named Entity Recognition, Moroccan-Context Oriented
MagBERT-NER-AR is a BERT-based named entity recognition (NER) model for Modern Standard Arabic (MSA), fine-tuned from asafaya/bert-base-arabic. It recognizes 12 entity types (OntoNotes-style) using the BILUO tagging scheme.
What sets it apart is its Moroccan-context orientation: the training corpus was manually curated from Moroccan newspapers. The model therefore works on standard MSA text from any Arab country, while being particularly well adapted to the entities, naming conventions and editorial nuances of the Moroccan press — Moroccan personal and family names (including Amazigh-origin names such as Aït …), cities, regions and provinces, national institutions and public offices, the territorial administration (wali, ‘amil, provinces, communes), Moroccan laws and codes, and Maghrebi date conventions (e.g. يناير، شتنبر، دجنبر rather than Levantine month names).
The model was developed by Typica.ai, kept private since 2021, and is now open-sourced for educational and research purposes.
Model Details
| Developed by | Hicham Assoudi — Typica.ai |
| Model type | BERT-base token classification (BertForTokenClassification) |
| Language | Arabic — Modern Standard Arabic (ar), Arabic script |
| Domain | News / journalistic text, Moroccan context |
| Base model | asafaya/bert-base-arabic |
| Tagging scheme | BILUO (Begin, Inside, Last, Unit, Outside) |
| Version | v0.1.1 |
| License | cc-by-nc-4.0 |
| Contact | assoudi@typica.ai |
Version History
| Date | Version | Notes |
|---|---|---|
| April 2021 | v0.0.x | First training on the Moroccan news corpus — private, internal use at Typica.ai |
| August 2023 | v0.1.1 | Retrained — this release |
| 2026 | v0.1.1 | Open-sourced on the Hugging Face Hub |
Entity Types
| Label | Description | Illustrative examples (Moroccan context) |
|---|---|---|
CARDINAL |
Other numerals | آلاف، 120 |
DATE |
Absolute or relative dates and periods | الثلاثاء المقبل، شهر شتنبر |
EVENT |
Named events, festivals, operations | عملية مرحبا، مهرجان كناوة |
GPE |
Countries, cities, regions, provinces | الرباط، جهة سوس ماسة، إقليم ورزازات |
LAW |
Named laws, codes, legal documents | مدونة الأسرة |
LOC |
Non-GPE locations: mountains, rivers, seas | جبال الأطلس، نهر أم الربيع |
NORP |
Nationalities, religious or political groups | المغاربة، مغاربة العالم |
ORDINAL |
Ordinals | الدورة الثالثة |
ORG |
Companies, institutions, parties, clubs | بنك المغرب، المكتب الوطني للسكك الحديدية |
PERCENT |
Percentages | 12 في المائة |
PERSON |
People, including fictional | سلمى العلوي، يوسف آيت الحاج |
PRODUCT |
Objects, vehicles, products | — |
Tags use the prefixes B, I, L, U, plus O, PAD: 42 labels in total (PAD is a technical label from training — treat it as O).
Intended Uses
Direct intended uses:
- Research and education in Arabic NLP and information extraction.
- Extracting people, organizations, places, dates and figures from Arabic news, especially Moroccan press content.
- Building blocks for search, knowledge graphs, media monitoring or document indexing prototypes in Arabic.
- Benchmarking Arabic NER, including comparisons between Moroccan and other Arab-country news text.
Out-of-scope uses:
- ❌ Moroccan Darija (dialect) or social-media text — the model was trained on formal MSA news.
- ❌ Arabizi or Latin-script text (French, English).
- ❌ Fully automated high-stakes decisions (legal, identity, surveillance) without human review.
- ❌ Commercial use, if the license is non-commercial, without a separate agreement with Typica.ai.
How to Use
Quick start (pipeline)
from transformers import pipeline
ner = pipeline(
"token-classification",
model="TypicaAI/MagBERT-NER-AR",
aggregation_strategy="first",
)
text = "يعقد بنك المغرب اجتماعه الفصلي بالرباط يوم الثلاثاء المقبل لتدارس تطور نسبة التضخم."
print(ner(text))
The generic pipeline aggregation only understands
B-/I-prefixes, so multi-word entities tagged withL-can come back split into several pieces. For clean entity spans, use the BILUO-aware helper below.
Recommended: BILUO-aware entity extraction
import torch
from transformers import AutoTokenizer, AutoModelForTokenClassification
repo = "TypicaAI/MagBERT-NER-AR"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForTokenClassification.from_pretrained(repo).eval()
def biluo_spans(text, words):
entities, cur = [], None
def close():
nonlocal cur
if cur is not None:
entities.append({
"entity_group": cur["type"],
"word": text[cur["start"]:cur["end"]],
"score": round(sum(cur["scores"]) / len(cur["scores"]), 4),
"start": cur["start"],
"end": cur["end"],
})
cur = None
for start, end, tag, score in words:
if tag in ("O", "PAD"):
close()
continue
prefix, _, etype = tag.partition("-")
if prefix in ("B", "U") or cur is None or cur["type"] != etype:
close()
cur = {"type": etype, "start": start, "end": end, "scores": [score]}
else:
cur["end"] = end
cur["scores"].append(score)
if prefix in ("L", "U"):
close()
close()
return entities
def extract_entities(text, max_length=512):
enc = tokenizer(text, return_offsets_mapping=True, truncation=True,
max_length=max_length, return_tensors="pt")
offsets = enc.pop("offset_mapping")[0].tolist()
word_ids = enc.word_ids(0)
with torch.no_grad():
probs = model(**enc).logits[0].softmax(-1)
words, prev = [], None
for i, wid in enumerate(word_ids):
if wid is None:
continue
start, end = offsets[i]
if wid != prev: # label of the first sub-token of each word
score, idx = probs[i].max(-1)
words.append([start, end, model.config.id2label[int(idx)], float(score)])
prev = wid
else:
words[-1][1] = end
return biluo_spans(text, words)
text = "أكدت الباحثة سلمى العلوي، أستاذة بجامعة القاضي عياض بمراكش، أن مراجعة مدونة الأسرة تتطلب نقاشا مجتمعيا واسعا."
for e in extract_entities(text):
print(f"{e['entity_group']:9} | {e['word']:30} | {e['score']:.3f}")
Each returned entity has the form {"entity_group", "word", "score", "start", "end"}, with character offsets into the original text. For long documents, split the text into sentences first: the model was trained at sentence level.
Training Data
The model was trained on a proprietary Arabic NER corpus manually curated from Moroccan newspapers by Typica.ai:
- Source: Moroccan news articles written in Modern Standard Arabic (politics, economy, society, culture, sport, regional news).
- Annotation: 12 OntoNotes-style entity types, BILUO scheme, word-level labels.
- Why it matters: the Moroccan editorial context exposes the model to entities and patterns under-represented in pan-Arab or Middle-East-centric corpora: Moroccan and Amazigh-origin names, provinces and communes, national offices and agencies, Moroccan legal texts and Maghrebi month names.
The training corpus is not released.
Training Procedure
| Hyperparameter | Value |
|---|---|
| Base checkpoint | asafaya/bert-base-arabic |
| Objective | Token classification (cross-entropy), full fine-tuning |
| Optimizer | AdamW, lr 3e-5, eps 1e-8 |
| LR schedule | Linear decay, no warm-up |
| Batch size | 8 |
| Epochs | 3 |
| Max sequence length | 175 sub-word tokens |
| Gradient clipping | 1.0 |
| Train / validation split | 75 % / 25 % (random, seed 2021) |
| Sub-word labelling | Each sub-token inherits its word's label |
Limitations & Bias
- Moroccan-press bias: entity coverage reflects the Moroccan news agenda; entities specific to other countries are recognized through general MSA patterns but may be less precise.
- Rare entity types: classes with little training support (e.g.
LAW,ORDINAL,PRODUCT,LOC) are less reliable thanPERSON,GPE,ORGorDATE. - Uncovered types: amounts of money, measurements, times of day and facilities are not separate classes in this version.
- Attached clitics: Arabic proclitics are not split (e.g. بالرباط is returned as one span including بال). Post-process if you need bare forms.
- Register: formal MSA only; Darija, Arabizi and noisy social-media text are out of scope.
- Sequence length: trained on sentences (≤175 sub-tokens); segment long documents.
- Temporal drift: news vocabulary, institutions and public figures change over time; data predates 2023.
- Annotation subjectivity: boundaries between
GPE/LOC,ORG/FACandEVENT/DATEfollow annotator judgement.
Ethical Considerations
NER can be used to extract personal names and link them to places and organizations. Use this model in line with applicable data-protection rules (including Moroccan Law 09-08 on personal data), avoid profiling individuals, and keep humans in the loop when outputs inform decisions.
Citation
@misc{assoudi2023magbertnerar,
title = {MagBERT-NER-AR: Arabic Named Entity Recognition, Moroccan-Context Oriented},
author = {Assoudi, Hicham},
year = {2023},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/TypicaAI/MagBERT-NER-AR}},
note = {Typica.ai. First trained April 2021, retrained August 2023 (v0.1.1)}
}
Contact
Hicham Assoudi — Founder & Applied AI Researcher, Typica.ai · PhD (AI/NLP) Typica.ai — Independent applied research initiative 📧 assoudi@typica.ai · LinkedIn · 🌐 typica.ai · 🤗 TypicaAI on Hugging Face
- Downloads last month
- -
Model tree for TypicaAI/MagBERT-NER-AR
Base model
asafaya/bert-base-arabic