sediba-XLM-R
A Sepedi masked-language model built on XLM-RoBERTa for Sepedi NLU research and token-level tasks.
Model Details
Model Description
- Developed by: Sediba AI NPC
- Funded by: Sediba AI NPC / TST (Pty) Ltd — Mankweng, Limpopo
- Shared by: Sediba AI NPC
- Model type: Masked language model (encoder-only Transformer)
- Language(s) (NLP): Sepedi (engpossible code-switched Sepedi/English contexts)
- License: MIT
- Finetuned from model: XLM-RoBERTa-base (Facebook/Google XLM-R)
Model Sources
- Repository: huggingface.co/Sediba-AI/sediba-XLM-R
- Demo:
huggingface.co/spaces/Sediba-AI/leotsha(7-mode CDI datafication game surface) - Paper: [forthcoming]
Uses
Direct Use
- Masked-token prediction and probing for Sepedi
- Token-level NLU research (intent, NER probes) built on Sepedi representations
- Evaluation baseline for Sepedi-language NLP tasks
Downstream Use
- Fine-tuning on Sepedi downstream tasks (classification, NER, etc.)
- Serving as the Sepedi representation anchor for the broader Sediba model family
Out-of-Scope Use
- Generative/conversational Sepedi (this is an MLM, not a generative LM)
- High-stakes automated decisions about Sepedi speakers without human review
Sepedi Families Context
This repo is the Sediba Sepedi MLM anchor. It is one model in a family with distinct roles:
|| Model | Role | Status | Repo / location |
||---|---|---|--|---|
|| sediba-XLM-R (this repo) | Masked-language Sepedi NLU anchor | Shipped | Sediba-AI/sediba-XLM-R on HF |
|| Leotša la Sediba | Conversational Sepedi (post-training / RLHF tier) | Pending — publication holds until 1B+ base tier | held in Sediba vaults; HF publication pending 1B+ base |
|| SedibaLM V7 | QLoRA experiment: Qwen2.5-1.5B-Instruct + custom Sepedi vocab extension | In training (Kaggle trail) | Kaggle trail (kaggle.com/sedibaai/sedibalm-v7-sepedi-qlora) + local v7_launch_bundle/ |
Leotša la Sediba and SedibaLM V7 are not published under this repo ID. They are tracked separately in Sediba's vaults and training pipeline, and differ from this model in base architecture, parameter tier, and task.
Evaluation
- Self-language sanity probe via lglang on a held-out Sepedi sample
- Sepedi-width fertility check comparing model output diversity against a Sepedi baseline sample
- Exact results are tracked in the Sediba vaults (model evaluation log) and are available on request.
Limitations
- Single-language Sepedi focus; performance on heavily code-switched or multi-register Sepedi is not fully characterised
- Masked-LM only — not a generative model; do not use for Sepedi text generation
- Training data scope and provenance are tracked in Sediba's data-governance stack (Esethu / NOODL / Itsebe) and the Sepedi corpus register
Citation
If you use this model, cite Sediba AI NPC and link the repo:
@misc{sediba-xlmr,
author = {Sediba AI NPC},
title = {sediba-XLM-R: A Sepedi masked-language model built on XLM-RoBERTa},
year = {2026},
url = {https://huggingface.co/Sediba-AI/sediba-XLM-R}
}
Build by Sediba AI | Mankweng, Limpopo.
- Downloads last month
- 20