UNICEF WCARO NLP Landscape

West and Central Africa Language Technology
Last updated: 2026-04-23
Work in progress. Some data may contain inaccuracies. Contribute on GitHub

UBC Deep Learning & NLP Lab

University established

Organization Information

University
~2018
medium
established
open
Vancouver, Canada (University of British Columbia)
University of British Columbia. Canadian research funding agencies.
UBC-NLP
β€”

Coverage

Description

Research lab at the University of British Columbia led by Muhammad Abdul-Mageed, focused on deep representation learning and natural language socio-pragmatics. While based in Canada, the lab has become one of the most prolific producers of African language NLP resources globally, with work spanning language modeling, machine translation, language identification, and benchmarking for hundreds of African languages. Their SERENGETI and Cheetah language models cover 517 African languages, Toucan provides machine translation for 156 African language pairs, and AfroLID identifies 517 African languages. They also maintain the SimbaBench and Sahara leaderboards for benchmarking African speech and NLP models.

πŸ¦„ UNICEF Relevance

Major producer of open African language technology resources covering hundreds of WCA languages. Their SERENGETI/Cheetah models (517 African languages), Toucan MT system (156 language pairs), and AfroLID language identification tool are directly usable in UNICEF programmes. The SimbaBench and Sahara benchmarks provide standardized evaluation frameworks for assessing African language technology - valuable for UNICEF's model selection decisions. Their "Towards Afrocentric NLP" position paper provides a useful framework for ethical language technology development in the region. Not Africa-based, but models and tools are open and freely available.

Key People

Projects

Publications

View all publications β†’

Partnerships

Notes

Canada-based lab, not African-led, but with significant African language output. Their scale of coverage (517 languages) is among the broadest of any research group. Strong publication record at top NLP venues (ACL, EMNLP, COLING). Also has extensive Arabic NLP work (ARBERT, MARBERT, AraT5, JASMINE, NileChat) which may be relevant for Shuwa Arabic and Mauritania. 45 models and 28 datasets on HuggingFace as of January 2026.

Last updated: 2026-01-26