File size: 1,754 Bytes
eb77fe9
 
 
 
 
 
 
 
 
a6072fb
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
---
title: README
emoji: 🐠
colorFrom: green
colorTo: yellow
sdk: static
pinned: false
---

# NaathNLP

Open NLP infrastructure for **Thok Naath** (the Nuer language) — built by volunteers working toward language preservation and revitalization.

Nuer is spoken by millions of people across South Sudan, Ethiopia, and diaspora communities, but remains a low-resource language with almost no native NLP tooling. NaathNLP exists to change that — building translation models, parallel corpora, and language tools with quality reviewed by native speakers at every step.

## What we're building

-  **Translation** — a fine-tuned NLLB model for English–Nuer translation
-  **Corpus** — a large-scale English–Nuer parallel corpus (1M+ pairs)
-  **Chatbot** — a pivot-pipeline conversational prototype
-  **ASR** — automatic speech recognition for spoken Nuer
-  **TTS** — text-to-speech for Nuer, to support learners and low-literacy speakers
-  **Reasoning** — working toward a native Nuer reasoning model that thinks in Nuer directly, rather than pivoting through English

## Why this matters

Most language models have never seen meaningful amounts of Nuer text or speech. Without deliberate investment, low-resource languages like Naath risk falling further behind as AI tools become part of everyday life — for education, communication, and access to information. NaathNLP is a volunteer effort to make sure Nuer speakers aren't left out of that shift, and to help preserve the language for future generations.

## Get involved

This is volunteer-driven work, and contributions are welcome — whether that's data, native-speaker review, compute, or code. Reach out if you'd like to help.

 ngunartaban@gmail.com
 github.com/bielng