MagBERT-NER-AR — Arabic Named Entity Recognition, Moroccan-Context Oriented

MagBERT-NER-AR is a BERT-based named entity recognition (NER) model for Modern Standard Arabic (MSA), fine-tuned from asafaya/bert-base-arabic. It recognizes 12 entity types (OntoNotes-style) using the BILUO tagging scheme.

What sets it apart is its Moroccan-context orientation: the training corpus was manually curated from Moroccan newspapers. The model therefore works on standard MSA text from any Arab country, while being particularly well adapted to the entities, naming conventions and editorial nuances of the Moroccan press — Moroccan personal and family names (including Amazigh-origin names such as Aït …), cities, regions and provinces, national institutions and public offices, the territorial administration (wali, ‘amil, provinces, communes), Moroccan laws and codes, and Maghrebi date conventions (e.g. يناير، شتنبر، دجنبر rather than Levantine month names).

The model was developed by Typica.ai, kept private since 2021, and is now open-sourced for educational and research purposes.

Model Details

Developed by Hicham Assoudi — Typica.ai
Model type BERT-base token classification (BertForTokenClassification)
Language Arabic — Modern Standard Arabic (ar), Arabic script
Domain News / journalistic text, Moroccan context
Base model asafaya/bert-base-arabic
Tagging scheme BILUO (Begin, Inside, Last, Unit, Outside)
Version v0.1.1
License cc-by-nc-4.0
Contact assoudi@typica.ai

Version History

Date Version Notes
April 2021 v0.0.x First training on the Moroccan news corpus — private, internal use at Typica.ai
August 2023 v0.1.1 Retrained — this release
2026 v0.1.1 Open-sourced on the Hugging Face Hub

Entity Types

Label Description Illustrative examples (Moroccan context)
CARDINAL Other numerals آلاف، 120
DATE Absolute or relative dates and periods الثلاثاء المقبل، شهر شتنبر
EVENT Named events, festivals, operations عملية مرحبا، مهرجان كناوة
GPE Countries, cities, regions, provinces الرباط، جهة سوس ماسة، إقليم ورزازات
LAW Named laws, codes, legal documents مدونة الأسرة
LOC Non-GPE locations: mountains, rivers, seas جبال الأطلس، نهر أم الربيع
NORP Nationalities, religious or political groups المغاربة، مغاربة العالم
ORDINAL Ordinals الدورة الثالثة
ORG Companies, institutions, parties, clubs بنك المغرب، المكتب الوطني للسكك الحديدية
PERCENT Percentages 12 في المائة
PERSON People, including fictional سلمى العلوي، يوسف آيت الحاج
PRODUCT Objects, vehicles, products —

Tags use the prefixes B, I, L, U, plus O, PAD: 42 labels in total (PAD is a technical label from training — treat it as O).

Intended Uses

Direct intended uses:

  • Research and education in Arabic NLP and information extraction.
  • Extracting people, organizations, places, dates and figures from Arabic news, especially Moroccan press content.
  • Building blocks for search, knowledge graphs, media monitoring or document indexing prototypes in Arabic.
  • Benchmarking Arabic NER, including comparisons between Moroccan and other Arab-country news text.

Out-of-scope uses:

  • ❌ Moroccan Darija (dialect) or social-media text — the model was trained on formal MSA news.
  • ❌ Arabizi or Latin-script text (French, English).
  • ❌ Fully automated high-stakes decisions (legal, identity, surveillance) without human review.
  • ❌ Commercial use, if the license is non-commercial, without a separate agreement with Typica.ai.

How to Use

Quick start (pipeline)

from transformers import pipeline

ner = pipeline(
    "token-classification",
    model="TypicaAI/MagBERT-NER-AR",
    aggregation_strategy="first",
)

text = "يعقد بنك المغرب اجتماعه الفصلي بالرباط يوم الثلاثاء المقبل لتدارس تطور نسبة التضخم."
print(ner(text))

The generic pipeline aggregation only understands B-/I- prefixes, so multi-word entities tagged with L- can come back split into several pieces. For clean entity spans, use the BILUO-aware helper below.

Recommended: BILUO-aware entity extraction

import torch
from transformers import AutoTokenizer, AutoModelForTokenClassification

repo = "TypicaAI/MagBERT-NER-AR"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForTokenClassification.from_pretrained(repo).eval()


def biluo_spans(text, words):
    entities, cur = [], None

    def close():
        nonlocal cur
        if cur is not None:
            entities.append({
                "entity_group": cur["type"],
                "word": text[cur["start"]:cur["end"]],
                "score": round(sum(cur["scores"]) / len(cur["scores"]), 4),
                "start": cur["start"],
                "end": cur["end"],
            })
            cur = None

    for start, end, tag, score in words:
        if tag in ("O", "PAD"):
            close()
            continue
        prefix, _, etype = tag.partition("-")
        if prefix in ("B", "U") or cur is None or cur["type"] != etype:
            close()
            cur = {"type": etype, "start": start, "end": end, "scores": [score]}
        else:
            cur["end"] = end
            cur["scores"].append(score)
        if prefix in ("L", "U"):
            close()
    close()
    return entities


def extract_entities(text, max_length=512):
    enc = tokenizer(text, return_offsets_mapping=True, truncation=True,
                    max_length=max_length, return_tensors="pt")
    offsets = enc.pop("offset_mapping")[0].tolist()
    word_ids = enc.word_ids(0)
    with torch.no_grad():
        probs = model(**enc).logits[0].softmax(-1)

    words, prev = [], None
    for i, wid in enumerate(word_ids):
        if wid is None:
            continue
        start, end = offsets[i]
        if wid != prev:  # label of the first sub-token of each word
            score, idx = probs[i].max(-1)
            words.append([start, end, model.config.id2label[int(idx)], float(score)])
            prev = wid
        else:
            words[-1][1] = end
    return biluo_spans(text, words)


text = "أكدت الباحثة سلمى العلوي، أستاذة بجامعة القاضي عياض بمراكش، أن مراجعة مدونة الأسرة تتطلب نقاشا مجتمعيا واسعا."
for e in extract_entities(text):
    print(f"{e['entity_group']:9} | {e['word']:30} | {e['score']:.3f}")

Each returned entity has the form {"entity_group", "word", "score", "start", "end"}, with character offsets into the original text. For long documents, split the text into sentences first: the model was trained at sentence level.

Training Data

The model was trained on a proprietary Arabic NER corpus manually curated from Moroccan newspapers by Typica.ai:

  • Source: Moroccan news articles written in Modern Standard Arabic (politics, economy, society, culture, sport, regional news).
  • Annotation: 12 OntoNotes-style entity types, BILUO scheme, word-level labels.
  • Why it matters: the Moroccan editorial context exposes the model to entities and patterns under-represented in pan-Arab or Middle-East-centric corpora: Moroccan and Amazigh-origin names, provinces and communes, national offices and agencies, Moroccan legal texts and Maghrebi month names.

The training corpus is not released.

Training Procedure

Hyperparameter Value
Base checkpoint asafaya/bert-base-arabic
Objective Token classification (cross-entropy), full fine-tuning
Optimizer AdamW, lr 3e-5, eps 1e-8
LR schedule Linear decay, no warm-up
Batch size 8
Epochs 3
Max sequence length 175 sub-word tokens
Gradient clipping 1.0
Train / validation split 75 % / 25 % (random, seed 2021)
Sub-word labelling Each sub-token inherits its word's label

Limitations & Bias

  • Moroccan-press bias: entity coverage reflects the Moroccan news agenda; entities specific to other countries are recognized through general MSA patterns but may be less precise.
  • Rare entity types: classes with little training support (e.g. LAW, ORDINAL, PRODUCT, LOC) are less reliable than PERSON, GPE, ORG or DATE.
  • Uncovered types: amounts of money, measurements, times of day and facilities are not separate classes in this version.
  • Attached clitics: Arabic proclitics are not split (e.g. بالرباط is returned as one span including بال). Post-process if you need bare forms.
  • Register: formal MSA only; Darija, Arabizi and noisy social-media text are out of scope.
  • Sequence length: trained on sentences (≤175 sub-tokens); segment long documents.
  • Temporal drift: news vocabulary, institutions and public figures change over time; data predates 2023.
  • Annotation subjectivity: boundaries between GPE/LOC, ORG/FAC and EVENT/DATE follow annotator judgement.

Ethical Considerations

NER can be used to extract personal names and link them to places and organizations. Use this model in line with applicable data-protection rules (including Moroccan Law 09-08 on personal data), avoid profiling individuals, and keep humans in the loop when outputs inform decisions.

Citation

@misc{assoudi2023magbertnerar,
  title        = {MagBERT-NER-AR: Arabic Named Entity Recognition, Moroccan-Context Oriented},
  author       = {Assoudi, Hicham},
  year         = {2023},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/TypicaAI/MagBERT-NER-AR}},
  note         = {Typica.ai. First trained April 2021, retrained August 2023 (v0.1.1)}
}

Contact

Hicham Assoudi — Founder & Applied AI Researcher, Typica.ai · PhD (AI/NLP) Typica.ai — Independent applied research initiative 📧 assoudi@typica.ai · LinkedIn · 🌐 typica.ai · 🤗 TypicaAI on Hugging Face

Downloads last month
-
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for TypicaAI/MagBERT-NER-AR

Finetuned
(21)
this model