README / README.md
niqqyniqqy's picture
Update README.md
117aa38 verified
|
Raw
History Blame Contribute Delete
3.36 kB
---
title: README
emoji: πŸ’»
colorFrom: purple
colorTo: yellow
sdk: static
pinned: false
license: cc-by-sa-4.0
short_description: LughaGen - Building LLMs for low-resource African languages
---
# LughaGen Organization
> *Bridging the gap from global foundational AI to local African authenticity.*
---
## About LughaGen
**LughaGen** is a student-led research initiative under **JHUB Africa** at **Jomo Kenyatta University of Agriculture and Technology (JKUAT)**, funded by the **NVIDIA Academic Grant Program**.
Our mission is to advance **data-efficient, culturally appropriate, and trustworthy** foundational Large Language Models for low-resource African languages, with a primary focus on **Kenyan and East African languages**.
We combine rigorous corpus engineering, compute-efficient training, safety & cultural alignment, and scalable deployment to create models that truly serve African contexts.
---
## Core Objectives
- **Data Efficiency** β€” High-quality, ethically sourced corpora for low-resource languages
- **Compute Efficiency** β€” Optimized training under constrained GPU resources (NVIDIA A100)
- **Cultural Relevance** β€” Models that respect and reflect local linguistic and cultural norms
- **Trust & Safety** β€” Robust evaluation frameworks for safety, bias, and explainability
- **Real-World Impact** β€” Scalable and deployable models for African use cases
---
## Research Pillars & Team
| Role | Lead | Focus Area |
|---|---|---|
| **Data Leads** | Nicolette Nkirote & Baraka Innocent | Corpus curation, cleaning & governance |
| **Model Lead** | Derek Mayabi | Efficient pretraining & customization |
| **Trust & Evaluation** | Kevin Musembi | Safety, cultural alignment & explainability |
| **Deployment Lead** | Godfrey Koros | Scalable inference & real-world deployment |
**Principal Investigator:** Dr. Lawrence Nderu
**Host Institution:** JHUB Africa Center of Excellence, JKUAT
---
## Current Work
- **LughaGen Corpus** β€” Multilingual Kenyan language dataset (Swahili, Kikuyu, Kamba, Dholuo) combining secondary, synthetic, and native speaker-informed data
- **LughaGen-F** β€” Foundational models trained on African corpora
- Synthetic data generation pipelines for low-resource languages
- Code-switching (Sheng) research
- Cultural alignment and safety benchmarks
---
## Resources
- **Datasets:** [LughaGen Multilingual African Corpus](https://huggingface.co/datasets/LughaGen-Organization/LughaGen-Multilingual-African-Corpus)
- **GitHub:** [LughaGen-Organization](https://github.com/LughaGen-Organization/LughaGen)
- **Models:** Coming soon
---
## License & Ethics
All datasets and models are released with clear provenance, licensing (primarily **CC-BY-SA 4.0**), and ethical documentation. We prioritize transparency, community benefit, and respect for the linguistic heritage of Kenyan and African communities.
---
## Acknowledgments
Powered by **NVIDIA A100 GPUs** through the NVIDIA Academic Grant Program.
We deeply thank the creators of all open African language resources and the Kenyan language communities whose linguistic data makes this work possible.
---
## Contact & Collaboration
- **Email:** nicolette.nkirote@students.jkuat.ac.ke | papai.baraka@students.jkuat.ac.ke
- **Organization:** [LughaGen-Organization on HuggingFace](https://huggingface.co/LughaGen-Organization)