Community-driven initiative building Africa's largest speech datasets for Nigerian languages. The NaijaVoices dataset contains 1,800+ hours of speech from 5,000+ speakers across Igbo, Hausa, and Yoruba -- making it the largest multi-speaker African speech dataset to date. Uses a "data farming" methodology where communities cultivate, own, and benefit from their language data (as opposed to extractive "data mining"). Fine-tuning ASR models on the dataset showed dramatic improvements: 75.86% WER reduction (Whisper), 52.06% (MMS), 42.33% (XLSR). Runs the Language Heritage Micro-Grants program, reinvesting dataset revenue into community-driven documentation of endangered Nigerian languages (2025 cohort covers Ehugbo, Nupe, Tyap, Gbagyi, Ekpeye, Ibono, Obolo).
Large open speech dataset for Nigeria's three major languages (Igbo, Hausa, Yoruba) directly relevant to UNICEF Nigeria -- a priority country. Community-driven "data farming" methodology aligns with UNICEF values around participation and ownership. Fine-tuned ASR models show dramatic WER improvements. Micro-grants program expands coverage to endangered Nigerian languages. Chris Emezue is a central node connecting NaijaVoices, Lanfrica, and Masakhane.
Dataset available at naijavoices.com (requires membership registration) and on HuggingFace. Strong cross-connections with Intron Health (Chris Emezue co-authored AfriSpeech papers), Lanfrica (same founder), and Masakhane. Intron Health lists NaijaVoices as a partner.