European Commission Smart Development Hack prize (1M+ euros). Mozilla Foundation support for Common Voice work. Gates Foundation (via African Next Voices project).
Rwandan AI and open data company building voice infrastructure for African languages. "Umuganda" is the Rwandan tradition of community service -- Digital Umuganda applies this model to crowdsourced data collection for language technology. Africa's largest contributor to Mozilla Common Voice with 2,200+ hours of Kinyarwanda speech data.
While based in Rwanda, their work extends to WCA languages: the AfriVoice dataset includes Lingala (517hrs), Fulani/Fulfulde (527hrs), and Wolof (531hrs) alongside Shona, Malagasy, and Somali. Their Open Data For All (OD4A) initiative has digitalized 17 languages across 14 countries with 4,500 audio hours recorded. Machine translation project for Rwanda done in collaboration with CLEAR Global.
🦄 UNICEF Relevance
Past contact of CLEAR Global (MT Rwanda collaboration). AfriVoice dataset directly covers WCA languages: Lingala (517hrs), Fulani/Fulfulde (527hrs), and Wolof (531hrs) -- all priority languages for UNICEF WCA. Their community-driven "digital umuganda" data collection methodology is replicable for other WCA languages that lack speech data. OD4A initiative spans 14 countries and 17 languages. ANV project partners include WCA actors RobotsMali and Data Science Nigeria.
Key People
Audace Niyonkuru - Founder & CEO
Samuel Rutunda - Chief Technology Officer
Projects
AfriVoice Dataset: ~3,200 hours of audio across 6 African languages. Each datapoint contains images, corresponding audio descriptions (WAV), and transcriptions. CC-BY-4.0. WCA-relevant languages: Lingala (517hrs total, 101hrs transcribed), Fulani/Fulfulde (527hrs total, 102hrs transcribed), Wolof (531hrs total, 103hrs transcribed). Also covers Shona (574hrs), Malagasy (516hrs), Somali (536hrs).
Open Data For All (OD4A): Large-scale data collection initiative building open datasets for African language technology. Collects voice recordings, text samples, and translations to support ASR, TTS, and MT. 17 languages digitalized, 552.8 million language speakers represented, 14 countries involved, 4,500 audio hours recorded. Datasets released open-source on HuggingFace.
Mozilla Common Voice - Kinyarwanda: Africa's largest Common Voice contribution. 2,200+ hours of validated Kinyarwanda speech data crowdsourced using the "digital umuganda" community model.
kin
Mbaza Innovation Hub: Innovation hub focused on real-world problem-solving. Co-creation with governments, private sector, and NGOs to develop tailored AI solutions.
Machine Translation Rwanda: Machine translation project for Kinyarwanda done in collaboration with CLEAR Global (Translators without Borders).
kin
WAXAL Dataset (with Google Research): Large-scale multilingual African speech corpus led by Google Research (Jan 2021 - Mar 2024). Digital Umuganda collected ASR data for 4 languages (Fulani 124.2hrs, Lingala 101.5hrs, Malagasy 182.5hrs, Shona 99.2hrs). Total dataset: ~1,250 hours ASR (14 languages) + ~186 hours TTS (10 languages), 21 languages total representing 100M+ speakers. WCA-relevant languages across the full dataset include Akan, Dagaare, Dagbani, Ewe, Fante, Fulani, Hausa, Igbo, Ikposo, Lingala, Twi, and Yoruba. CC-BY-4.0. Partners: Google Research (lead), University of Ghana, Makerere University, Media Trust Limited. Paper: arxiv 2602.02734.
AfriVoice / African Next Voices (ANV): Pan-African speech dataset initiative creating 9,000+ hours of speech across 18 languages. Partners: University of Pretoria, Maseno University, Data Science Nigeria, RobotsMali, Masakhane, Lelapa AI, Lanfrica. Funded by $2.2M Gates Foundation grant + Meta. Featured in Nature.
Rwanda-based (not WCA), but directly relevant due to: (1) AfriVoice dataset covering Lingala, Fulfulde, and Wolof, (2) CLEAR Global collaboration on MT Rwanda, (3) ANV project partnerships with WCA actors (RobotsMali, DSN), (4) replicable community data collection model. Their 2,200hrs Kinyarwanda Common Voice contribution demonstrates what's achievable with community mobilization. OD4A's scale (4,500hrs across 17 languages, 14 countries) makes them one of the larger open data producers for African languages.