For years, Sheriff Issaka kept running into the same problem while building language technology for Africa: there was not enough data. The data needed to train systems for African languages was scarce, and much of what was available was not good enough.
That problem changed what African Languages Lab, the AI research and deployment company Issaka founded in 2020, spent its time doing. Instead of building systems only with the available data, the company began collecting the data it could not find.
The company says it has since built the largest collection of African-language data in existence, covering more than 70 languages, including Amharic, Hausa, Zulu, Twi, Igbo, and Yoruba.
On Tuesday, African Languages Lab launched Mansa, a multilingual and multimodal AI platform that puts about 30 of those languages into production. It is available on the web and mobile, with APIs for developers and businesses.
The company says its datasets now contain more than 100 billion curated tokens, or pieces of text used to train AI models, alongside more than 19,000 hours of speech recordings that have been reviewed and validated by language experts.
The problem is bigger than poor translations
Africa is home to more than 2,000 languages, but fewer than 5% have the resources needed for natural language processing. Most African languages remain poorly represented in digital datasets, limiting how well AI systems can understand them.
“The models are just very bad at understanding our languages,” Issaka said. “If you try to have a full-blown conversation with an LLM in most African languages, they just cannot pull it off.”
There is also a cost problem. African languages can require significantly more tokens than English for the same input. Chioma Agwuedo,…
Source link
Read Full Article by John Adoyi at techcabal.com
Source link
