Non-profit research lab providing research opportunities to independent and underrepresented researchers. Runs multiple African language NLP projects including AfriCaption (image captioning in 20 African languages), data augmentation for African MT, Nigerian language embeddings, and Yoruba QA robustness evaluation. Connected to the NaijaVoices speech dataset initiative. Provides mentorship, compute resources, and publication support to researchers working on African language technology.
AfriCaption covers 8 UNICEF focus languages including Yoruba, Igbo, Hausa, Ewe, Lingala, Nigerian Fulfulde, Dyula, and Bambara β the broadest WCA language coverage of any single project. Multimodal AI (image captioning) could support visual content accessibility. Provides research mentorship pipeline for African NLP talent.