Research lab within the Department of Computer Science at University of Ghana, led by Prof. Isaac Wiafe. Focused on Human-Computer Interaction, persuasive technology, and AI/ML for African languages. The lab has produced the largest open speech corpus for Ghanaian languages: UGSpeechData with 5,000 hours across 5 languages (Akan, Ewe, Dagbani, Dagaare, Ikposo - 1,000 hours each, 100 hours transcribed per language). Current flagship project Tɛkyerɛma Pa ("Good Tongue") develops ASR for people with speech impairments in collaboration with Google Research Africa and UCL GDI Hub. Additional work includes maternal health translation systems and TTS development.
🦄 UNICEF Relevance
Extremely relevant for UNICEF Ghana and regional language technology needs. The 5,000-hour UGSpeechData corpus is the largest open speech dataset for Ghanaian languages - a foundational resource for any voice-based application. The Tɛkyerɛma Pa project directly serves people with disabilities, aligning with UNICEF's inclusive technology mandate. Maternal health translation work supports health communication programs. The lab has proven ability to collaborate with major tech companies (Google) and international organizations (UCL). Could support AGOO platform, Asempa AI companion, and health messaging in Akan, Ewe, Dagbani, Dagaare. Academic setting enables research partnerships and student involvement.
Key People
Prof. Isaac Wiafe - Lab Lead, Associate Professor PhD Informatics from University of Reading. 15+ years researching HCI and persuasive technologies. Principal investigator on UGSpeechData and Tɛkyerɛma Pa. Research areas: Persuasive Technologies, HCI, VR, NLP, ASR. 947 citations. HuggingFace: IsaacWiafe
Prof. Jamal-Deen Abdulai - Head of Department, Computer Science PhD Computer Science from University of Glasgow (2009). Co-author of UGSpeechData. Research: AI/ML, wireless networks, sensor networks, embedded systems. Top 10 Computer Science academics in Ghana. 1,696 citations.
Kelvin Nketia-Achiampong - Research Assistant, MSc Student Also known as "Fiifi". Over 5 years experience in scalable systems. Develops ASR and TTS for Akan/Twi using VITS and YourTTS. Focus on maternal health speech tech. Personal site: fiifinketia.github.io. HuggingFace: fiifinketia
Mark Atta Mensah - Researcher Developed Akan Whisper model for ASR. Focus on ASR, NLP, and persuasive technologies. HuggingFace: GiftMark
Fiifi Baffoe Payin Winful - Front-end Developer, Researcher Develops Ewe language ASR models. Co-author on UGSpeechData. Also worked on Kinyarwanda speech recognition. Farmerline collaboration. HuggingFace: Paywinful
Projects
UGSpeechData (5000 hours): MASSIVE multilingual speech corpus: 5,000 hours across 5 Ghanaian languages (Akan, Ewe, Dagbani, Dagaare, Ikposo). 1,000 hours per language from native speakers, 100 hours transcribed per language. 970,148 audio files total. Licensed CC BY-NC-ND 4.0. DOI: 10.57760/sciencedb.22298
Tɛkyerɛma Pa (Good Tongue): AI-powered ASR for people with non-standard speech (cerebral palsy, stroke, Down syndrome, ALS, Parkinson's) in 5 Ghanaian languages. $40k Google grant. Collaboration with Google Research Africa and UCL GDI Hub. Creating first open-source dataset of impaired speech in Ghanaian languages. Mobile app runs locally without internet.
UGSpeech Akan 100hrs (HuggingFace): Clean subset of UGSpeechData: 100 hours Akan speech on HuggingFace. 12.7k downloads. Major resource for Akan ASR development.
BibleTTS Asante Twi Segmented: 9 hours of segmented Asante Twi speech from BibleTTS for TTS training (max 29 second segments, 22050Hz). 3.25k downloads on HuggingFace.
WAXAL Dataset (with Google Research): Large-scale multilingual African speech corpus led by Google Research (Jan 2021 - Mar 2024). University of Ghana collected ASR data for 6 languages (Akan 101.9hrs, Dagaare 104.7hrs, Dagbani 98.5hrs, Ewe 99.8hrs, Ikposo 103.8hrs, Fante) and TTS data for 2 languages (Akan 15.8hrs, Fante 24.0hrs). Total dataset: ~1,250 hours ASR (14 languages) + ~186 hours TTS (10 languages), 21 languages total representing 100M+ speakers. WCA-relevant languages across the full dataset include Akan, Dagaare, Dagbani, Ewe, Fante, Fulani, Hausa, Igbo, Ikposo, Lingala, Twi, and Yoruba. CC-BY-4.0. Partners: Google Research (lead), Digital Umuganda, Makerere University, Media Trust Limited. Paper: arxiv 2602.02734.
University of Ghana, Department of Computer Science
Notes
Major resource hub for Ghanaian language technology. The 5,000-hour UGSpeechData corpus is a game-changer for the region. Strong leadership with Prof. Isaac Wiafe (947 citations) and Prof. Jamal-Deen Abdulai (1,696 citations, current HOD). Google Research collaboration through Katrin Tomanek brings world-class expertise. The Tɛkyerɛma Pa hackathon engaged students across Ghana (winner: Kasa Noma team from UESD, $2,500 prize). Lab members also collaborate with Farmerline on agricultural speech tech. Individual researchers contribute via personal HuggingFace accounts (fiifinketia, GiftMark, Paywinful, IsaacWiafe).