AI & ML interests

NaathNLP is a volunteer-led initiative building open NLP infrastructure for Thok Naath (the Nuer language), including translation models, parallel corpora, and tools aimed at language preservation and revitalization.

Recent Activity

dayomtechnologies  updated a Space 3 days ago
NaathNLP/README
dayomtechnologies  published a Space 3 days ago
NaathNLP/README
View all activity

Organization Card

NaathNLP

Open NLP infrastructure for Thok Naath (the Nuer language) — built by volunteers working toward language preservation and revitalization.

Nuer is spoken by millions of people across South Sudan, Ethiopia, and diaspora communities, but remains a low-resource language with almost no native NLP tooling. NaathNLP exists to change that — building translation models, parallel corpora, and language tools with quality reviewed by native speakers at every step.

What we're building

  • Translation — a fine-tuned NLLB model for English–Nuer translation
  • Corpus — a large-scale English–Nuer parallel corpus (1M+ pairs)
  • Chatbot — a pivot-pipeline conversational prototype
  • ASR — automatic speech recognition for spoken Nuer
  • TTS — text-to-speech for Nuer, to support learners and low-literacy speakers
  • Reasoning — working toward a native Nuer reasoning model that thinks in Nuer directly, rather than pivoting through English

Why this matters

Most language models have never seen meaningful amounts of Nuer text or speech. Without deliberate investment, low-resource languages like Naath risk falling further behind as AI tools become part of everyday life — for education, communication, and access to information. NaathNLP is a volunteer effort to make sure Nuer speakers aren't left out of that shift, and to help preserve the language for future generations.

Get involved

This is volunteer-driven work, and contributions are welcome — whether that's data, native-speaker review, compute, or code. Reach out if you'd like to help.

ngunartaban@gmail.com github.com/bielng

models 0

None public yet

datasets 0

None public yet