AI & ML interests
Large language models for low-resource African languages. Research initiative under JHUB Africa, JKUAT, funded by NVIDIA. Focus languages: Swahili, Kikuyu, Kamba, Luo.
Recent Activity
LughaGen Organization
Bridging the gap from global foundational AI to local African authenticity.
About LughaGen
LughaGen is a student-led research initiative under JHUB Africa at Jomo Kenyatta University of Agriculture and Technology (JKUAT), funded by the NVIDIA Academic Grant Program.
Our mission is to advance data-efficient, culturally appropriate, and trustworthy foundational Large Language Models for low-resource African languages, with a primary focus on Kenyan and East African languages.
We combine rigorous corpus engineering, compute-efficient training, safety & cultural alignment, and scalable deployment to create models that truly serve African contexts.
Core Objectives
- Data Efficiency — High-quality, ethically sourced corpora for low-resource languages
- Compute Efficiency — Optimized training under constrained GPU resources (NVIDIA A100)
- Cultural Relevance — Models that respect and reflect local linguistic and cultural norms
- Trust & Safety — Robust evaluation frameworks for safety, bias, and explainability
- Real-World Impact — Scalable and deployable models for African use cases
Research Pillars & Team
| Role | Lead | Focus Area |
|---|---|---|
| Data Leads | Nicolette Nkirote & Baraka Innocent | Corpus curation, cleaning & governance |
| Model Lead | Derek Mayabi | Efficient pretraining & customization |
| Trust & Evaluation | Kevin Musembi | Safety, cultural alignment & explainability |
| Deployment Lead | Godfrey Koros | Scalable inference & real-world deployment |
Principal Investigator: Dr. Lawrence Nderu Host Institution: JHUB Africa Center of Excellence, JKUAT
Current Work
- LughaGen Corpus — Multilingual Kenyan language dataset (Swahili, Kikuyu, Kamba, Dholuo) combining secondary, synthetic, and native speaker-informed data
- LughaGen-F — Foundational models trained on African corpora
- Synthetic data generation pipelines for low-resource languages
- Code-switching (Sheng) research
- Cultural alignment and safety benchmarks
Resources
- Datasets: LughaGen Multilingual African Corpus
- GitHub: LughaGen-Organization
- Models: Coming soon
License & Ethics
All datasets and models are released with clear provenance, licensing (primarily CC-BY-SA 4.0), and ethical documentation. We prioritize transparency, community benefit, and respect for the linguistic heritage of Kenyan and African communities.
Acknowledgments
Powered by NVIDIA A100 GPUs through the NVIDIA Academic Grant Program.
We deeply thank the creators of all open African language resources and the Kenyan language communities whose linguistic data makes this work possible.