UNICEF WCARO NLP Landscape

West and Central Africa Language Technology
Last updated: 2026-04-23
Work in progress. Some data may contain inaccuracies. Contribute on GitHub

Umbaji

Startup early

Organization Information

Startup
small
early
open
Lomé, Togo
Umbaji

Coverage

Togo

Description

Togo-based organization building NLP resources for Togolese languages. Developed the Eyaa-Tom multilingual speech corpus (30.9 hours across 10 Togolese languages) and the Yodi Data Hub, a community platform for crowdsourcing speech and text with fair compensation (415 contributors from 6 countries, 60,000+ clips collected). Fine-tuned Whisper and NLLB models for Togolese languages and introduced the Lom Bench, a community-based TTS evaluation benchmark. Uses the Nwulite Obodo License (NOODL) balancing open access with community ownership. Also sells robotics and hardware products.

🦄 UNICEF Relevance

Covers Ewe, a UNICEF priority language. Community-driven data collection with fair compensation aligns with UNICEF values. Togo is a WCA country with no other dedicated language technology actor. Ethical data practices and community ownership model.

Key People

Projects

Publications

Partnerships

Notes

Primary business appears to be robotics/hardware distribution alongside NLP research. NLP work is legitimate and published at peer-reviewed venues. Founder well-connected in African AI community.

Last updated: 2026-03-31