--- title: README emoji: 💻 colorFrom: purple colorTo: yellow sdk: static pinned: false license: cc-by-sa-4.0 short_description: LughaGen - Building LLMs for low-resource African languages --- # LughaGen Organization > *Bridging the gap from global foundational AI to local African authenticity.* --- ## About LughaGen **LughaGen** is a student-led research initiative under **JHUB Africa** at **Jomo Kenyatta University of Agriculture and Technology (JKUAT)**, funded by the **NVIDIA Academic Grant Program**. Our mission is to advance **data-efficient, culturally appropriate, and trustworthy** foundational Large Language Models for low-resource African languages, with a primary focus on **Kenyan and East African languages**. We combine rigorous corpus engineering, compute-efficient training, safety & cultural alignment, and scalable deployment to create models that truly serve African contexts. --- ## Core Objectives - **Data Efficiency** — High-quality, ethically sourced corpora for low-resource languages - **Compute Efficiency** — Optimized training under constrained GPU resources (NVIDIA A100) - **Cultural Relevance** — Models that respect and reflect local linguistic and cultural norms - **Trust & Safety** — Robust evaluation frameworks for safety, bias, and explainability - **Real-World Impact** — Scalable and deployable models for African use cases --- ## Research Pillars & Team | Role | Lead | Focus Area | |---|---|---| | **Data Leads** | Nicolette Nkirote & Baraka Innocent | Corpus curation, cleaning & governance | | **Model Lead** | Derek Mayabi | Efficient pretraining & customization | | **Trust & Evaluation** | Kevin Musembi | Safety, cultural alignment & explainability | | **Deployment Lead** | Godfrey Koros | Scalable inference & real-world deployment | **Principal Investigator:** Dr. Lawrence Nderu **Host Institution:** JHUB Africa Center of Excellence, JKUAT --- ## Current Work - **LughaGen Corpus** — Multilingual Kenyan language dataset (Swahili, Kikuyu, Kamba, Dholuo) combining secondary, synthetic, and native speaker-informed data - **LughaGen-F** — Foundational models trained on African corpora - Synthetic data generation pipelines for low-resource languages - Code-switching (Sheng) research - Cultural alignment and safety benchmarks --- ## Resources - **Datasets:** [LughaGen Multilingual African Corpus](https://huggingface.co/datasets/LughaGen-Organization/LughaGen-Multilingual-African-Corpus) - **GitHub:** [LughaGen-Organization](https://github.com/LughaGen-Organization/LughaGen) - **Models:** Coming soon --- ## License & Ethics All datasets and models are released with clear provenance, licensing (primarily **CC-BY-SA 4.0**), and ethical documentation. We prioritize transparency, community benefit, and respect for the linguistic heritage of Kenyan and African communities. --- ## Acknowledgments Powered by **NVIDIA A100 GPUs** through the NVIDIA Academic Grant Program. We deeply thank the creators of all open African language resources and the Kenyan language communities whose linguistic data makes this work possible. --- ## Contact & Collaboration - **Email:** nicolette.nkirote@students.jkuat.ac.ke | papai.baraka@students.jkuat.ac.ke - **Organization:** [LughaGen-Organization on HuggingFace](https://huggingface.co/LughaGen-Organization)