| --- |
| language: nso |
| tags: |
| - xlm-roberta |
| - masked-language-model |
| - sepedi |
| - african-language |
| - sediba |
| - nlu |
| - nlp |
| license: mit |
| --- |
| |
| # sediba-XLM-R |
|
|
| A Sepedi masked-language model built on XLM-RoBERTa for Sepedi NLU research and token-level tasks. |
|
|
| ## Model Details |
|
|
| ### Model Description |
|
|
| - **Developed by:** Sediba AI NPC |
| - **Funded by:** Sediba AI NPC / TST (Pty) Ltd — Mankweng, Limpopo |
| - **Shared by:** Sediba AI NPC |
| - **Model type:** Masked language model (encoder-only Transformer) |
| - **Language(s) (NLP):** Sepedi (engpossible code-switched Sepedi/English contexts) |
| - **License:** MIT |
| - **Finetuned from model:** XLM-RoBERTa-base (Facebook/Google XLM-R) |
|
|
| ### Model Sources |
|
|
| - **Repository:** huggingface.co/Sediba-AI/sediba-XLM-R |
| - **Demo:** `huggingface.co/spaces/Sediba-AI/leotsha` (7-mode CDI datafication game surface) |
| - **Paper:** [forthcoming] |
|
|
| ## Uses |
|
|
| ### Direct Use |
|
|
| - Masked-token prediction and probing for Sepedi |
| - Token-level NLU research (intent, NER probes) built on Sepedi representations |
| - Evaluation baseline for Sepedi-language NLP tasks |
|
|
| ### Downstream Use |
|
|
| - Fine-tuning on Sepedi downstream tasks (classification, NER, etc.) |
| - Serving as the Sepedi representation anchor for the broader Sediba model family |
|
|
| ### Out-of-Scope Use |
|
|
| - Generative/conversational Sepedi (this is an MLM, not a generative LM) |
| - High-stakes automated decisions about Sepedi speakers without human review |
|
|
| ## Sepedi Families Context |
|
|
| This repo is the **Sediba Sepedi MLM anchor**. It is one model in a family with distinct roles: |
|
|
| || Model | Role | Status | Repo / location | |
| ||---|---|---|--|---| |
| || **sediba-XLM-R** (this repo) | Masked-language Sepedi NLU anchor | **Shipped** | `Sediba-AI/sediba-XLM-R` on HF | |
| || **Leotša la Sediba** | Conversational Sepedi (post-training / RLHF tier) | **Pending — publication holds until 1B+ base tier** | held in Sediba vaults; HF publication pending 1B+ base | |
| || **SedibaLM V7** | QLoRA experiment: Qwen2.5-1.5B-Instruct + custom Sepedi vocab extension | **In training (Kaggle trail)** | Kaggle trail (`kaggle.com/sedibaai/sedibalm-v7-sepedi-qlora`) + local `v7_launch_bundle/` | |
|
|
| **Leotša la Sediba** and SedibaLM V7 are **not** published under this repo ID. They are tracked separately in Sediba's vaults and training pipeline, and differ from this model in base architecture, parameter tier, and task. |
|
|
| ## Evaluation |
|
|
| - Self-language sanity probe via lglang on a held-out Sepedi sample |
| - Sepedi-width fertility check comparing model output diversity against a Sepedi baseline sample |
| - Exact results are tracked in the Sediba vaults (model evaluation log) and are available on request. |
|
|
| ## Limitations |
|
|
| - Single-language Sepedi focus; performance on heavily code-switched or multi-register Sepedi is not fully characterised |
| - Masked-LM only — not a generative model; do not use for Sepedi text generation |
| - Training data scope and provenance are tracked in Sediba's data-governance stack (Esethu / NOODL / Itsebe) and the Sepedi corpus register |
|
|
| ## Citation |
|
|
| If you use this model, cite Sediba AI NPC and link the repo: |
|
|
| ``` |
| @misc{sediba-xlmr, |
| author = {Sediba AI NPC}, |
| title = {sediba-XLM-R: A Sepedi masked-language model built on XLM-RoBERTa}, |
| year = {2026}, |
| url = {https://huggingface.co/Sediba-AI/sediba-XLM-R} |
| } |
| ``` |
|
|
| Build by Sediba AI | Mankweng, Limpopo. |
|
|