Spaces:
Running
Running
| title: README | |
| emoji: π» | |
| colorFrom: purple | |
| colorTo: yellow | |
| sdk: static | |
| pinned: false | |
| license: cc-by-sa-4.0 | |
| short_description: LughaGen - Building LLMs for low-resource African languages | |
| # LughaGen Organization | |
| > *Bridging the gap from global foundational AI to local African authenticity.* | |
| --- | |
| ## About LughaGen | |
| **LughaGen** is a student-led research initiative under **JHUB Africa** at **Jomo Kenyatta University of Agriculture and Technology (JKUAT)**, funded by the **NVIDIA Academic Grant Program**. | |
| Our mission is to advance **data-efficient, culturally appropriate, and trustworthy** foundational Large Language Models for low-resource African languages, with a primary focus on **Kenyan and East African languages**. | |
| We combine rigorous corpus engineering, compute-efficient training, safety & cultural alignment, and scalable deployment to create models that truly serve African contexts. | |
| --- | |
| ## Core Objectives | |
| - **Data Efficiency** β High-quality, ethically sourced corpora for low-resource languages | |
| - **Compute Efficiency** β Optimized training under constrained GPU resources (NVIDIA A100) | |
| - **Cultural Relevance** β Models that respect and reflect local linguistic and cultural norms | |
| - **Trust & Safety** β Robust evaluation frameworks for safety, bias, and explainability | |
| - **Real-World Impact** β Scalable and deployable models for African use cases | |
| --- | |
| ## Research Pillars & Team | |
| | Role | Lead | Focus Area | | |
| |---|---|---| | |
| | **Data Leads** | Nicolette Nkirote & Baraka Innocent | Corpus curation, cleaning & governance | | |
| | **Model Lead** | Derek Mayabi | Efficient pretraining & customization | | |
| | **Trust & Evaluation** | Kevin Musembi | Safety, cultural alignment & explainability | | |
| | **Deployment Lead** | Godfrey Koros | Scalable inference & real-world deployment | | |
| **Principal Investigator:** Dr. Lawrence Nderu | |
| **Host Institution:** JHUB Africa Center of Excellence, JKUAT | |
| --- | |
| ## Current Work | |
| - **LughaGen Corpus** β Multilingual Kenyan language dataset (Swahili, Kikuyu, Kamba, Dholuo) combining secondary, synthetic, and native speaker-informed data | |
| - **LughaGen-F** β Foundational models trained on African corpora | |
| - Synthetic data generation pipelines for low-resource languages | |
| - Code-switching (Sheng) research | |
| - Cultural alignment and safety benchmarks | |
| --- | |
| ## Resources | |
| - **Datasets:** [LughaGen Multilingual African Corpus](https://huggingface.co/datasets/LughaGen-Organization/LughaGen-Multilingual-African-Corpus) | |
| - **GitHub:** [LughaGen-Organization](https://github.com/LughaGen-Organization/LughaGen) | |
| - **Models:** Coming soon | |
| --- | |
| ## License & Ethics | |
| All datasets and models are released with clear provenance, licensing (primarily **CC-BY-SA 4.0**), and ethical documentation. We prioritize transparency, community benefit, and respect for the linguistic heritage of Kenyan and African communities. | |
| --- | |
| ## Acknowledgments | |
| Powered by **NVIDIA A100 GPUs** through the NVIDIA Academic Grant Program. | |
| We deeply thank the creators of all open African language resources and the Kenyan language communities whose linguistic data makes this work possible. | |
| --- | |
| ## Contact & Collaboration | |
| - **Email:** nicolette.nkirote@students.jkuat.ac.ke | papai.baraka@students.jkuat.ac.ke | |
| - **Organization:** [LughaGen-Organization on HuggingFace](https://huggingface.co/LughaGen-Organization) | |