Spaces:
Running
Running
File size: 3,361 Bytes
d53cfd2 ff69380 d53cfd2 117aa38 ff69380 117aa38 ff69380 117aa38 ff69380 117aa38 ff69380 117aa38 ff69380 117aa38 ff69380 117aa38 ff69380 117aa38 ff69380 117aa38 ff69380 117aa38 ff69380 117aa38 ff69380 117aa38 ff69380 117aa38 ff69380 117aa38 ff69380 117aa38 ff69380 117aa38 ff69380 117aa38 ff69380 117aa38 ff69380 117aa38 ff69380 117aa38 ff69380 117aa38 ff69380 117aa38 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 | ---
title: README
emoji: π»
colorFrom: purple
colorTo: yellow
sdk: static
pinned: false
license: cc-by-sa-4.0
short_description: LughaGen - Building LLMs for low-resource African languages
---
# LughaGen Organization
> *Bridging the gap from global foundational AI to local African authenticity.*
---
## About LughaGen
**LughaGen** is a student-led research initiative under **JHUB Africa** at **Jomo Kenyatta University of Agriculture and Technology (JKUAT)**, funded by the **NVIDIA Academic Grant Program**.
Our mission is to advance **data-efficient, culturally appropriate, and trustworthy** foundational Large Language Models for low-resource African languages, with a primary focus on **Kenyan and East African languages**.
We combine rigorous corpus engineering, compute-efficient training, safety & cultural alignment, and scalable deployment to create models that truly serve African contexts.
---
## Core Objectives
- **Data Efficiency** β High-quality, ethically sourced corpora for low-resource languages
- **Compute Efficiency** β Optimized training under constrained GPU resources (NVIDIA A100)
- **Cultural Relevance** β Models that respect and reflect local linguistic and cultural norms
- **Trust & Safety** β Robust evaluation frameworks for safety, bias, and explainability
- **Real-World Impact** β Scalable and deployable models for African use cases
---
## Research Pillars & Team
| Role | Lead | Focus Area |
|---|---|---|
| **Data Leads** | Nicolette Nkirote & Baraka Innocent | Corpus curation, cleaning & governance |
| **Model Lead** | Derek Mayabi | Efficient pretraining & customization |
| **Trust & Evaluation** | Kevin Musembi | Safety, cultural alignment & explainability |
| **Deployment Lead** | Godfrey Koros | Scalable inference & real-world deployment |
**Principal Investigator:** Dr. Lawrence Nderu
**Host Institution:** JHUB Africa Center of Excellence, JKUAT
---
## Current Work
- **LughaGen Corpus** β Multilingual Kenyan language dataset (Swahili, Kikuyu, Kamba, Dholuo) combining secondary, synthetic, and native speaker-informed data
- **LughaGen-F** β Foundational models trained on African corpora
- Synthetic data generation pipelines for low-resource languages
- Code-switching (Sheng) research
- Cultural alignment and safety benchmarks
---
## Resources
- **Datasets:** [LughaGen Multilingual African Corpus](https://huggingface.co/datasets/LughaGen-Organization/LughaGen-Multilingual-African-Corpus)
- **GitHub:** [LughaGen-Organization](https://github.com/LughaGen-Organization/LughaGen)
- **Models:** Coming soon
---
## License & Ethics
All datasets and models are released with clear provenance, licensing (primarily **CC-BY-SA 4.0**), and ethical documentation. We prioritize transparency, community benefit, and respect for the linguistic heritage of Kenyan and African communities.
---
## Acknowledgments
Powered by **NVIDIA A100 GPUs** through the NVIDIA Academic Grant Program.
We deeply thank the creators of all open African language resources and the Kenyan language communities whose linguistic data makes this work possible.
---
## Contact & Collaboration
- **Email:** nicolette.nkirote@students.jkuat.ac.ke | papai.baraka@students.jkuat.ac.ke
- **Organization:** [LughaGen-Organization on HuggingFace](https://huggingface.co/LughaGen-Organization)
|