AI & ML interests

Large language models for low-resource African languages. Research initiative under JHUB Africa, JKUAT, funded by NVIDIA. Focus languages: Swahili, Kikuyu, Kamba, Luo.

Recent Activity

Organization Card

LughaGen Organization

Bridging the gap from global foundational AI to local African authenticity.


About LughaGen

LughaGen is a student-led research initiative under JHUB Africa at Jomo Kenyatta University of Agriculture and Technology (JKUAT), funded by the NVIDIA Academic Grant Program.

Our mission is to advance data-efficient, culturally appropriate, and trustworthy foundational Large Language Models for low-resource African languages, with a primary focus on Kenyan and East African languages.

We combine rigorous corpus engineering, compute-efficient training, safety & cultural alignment, and scalable deployment to create models that truly serve African contexts.


Core Objectives

  • Data Efficiency — High-quality, ethically sourced corpora for low-resource languages
  • Compute Efficiency — Optimized training under constrained GPU resources (NVIDIA A100)
  • Cultural Relevance — Models that respect and reflect local linguistic and cultural norms
  • Trust & Safety — Robust evaluation frameworks for safety, bias, and explainability
  • Real-World Impact — Scalable and deployable models for African use cases

Research Pillars & Team

Role Lead Focus Area
Data Leads Nicolette Nkirote & Baraka Innocent Corpus curation, cleaning & governance
Model Lead Derek Mayabi Efficient pretraining & customization
Trust & Evaluation Kevin Musembi Safety, cultural alignment & explainability
Deployment Lead Godfrey Koros Scalable inference & real-world deployment

Principal Investigator: Dr. Lawrence Nderu Host Institution: JHUB Africa Center of Excellence, JKUAT


Current Work

  • LughaGen Corpus — Multilingual Kenyan language dataset (Swahili, Kikuyu, Kamba, Dholuo) combining secondary, synthetic, and native speaker-informed data
  • LughaGen-F — Foundational models trained on African corpora
  • Synthetic data generation pipelines for low-resource languages
  • Code-switching (Sheng) research
  • Cultural alignment and safety benchmarks

Resources


License & Ethics

All datasets and models are released with clear provenance, licensing (primarily CC-BY-SA 4.0), and ethical documentation. We prioritize transparency, community benefit, and respect for the linguistic heritage of Kenyan and African communities.


Acknowledgments

Powered by NVIDIA A100 GPUs through the NVIDIA Academic Grant Program.

We deeply thank the creators of all open African language resources and the Kenyan language communities whose linguistic data makes this work possible.


Contact & Collaboration

models 0

None public yet