Saturday, August 29, 2026

Top 5 This Week

spot_img

Related Posts

The African-built tool helping the world’s under-resourced languages enter the AI age

By HER Staff Reporter

A software tool developed by South African researchers to address one of the greatest barriers facing African languages in the digital age is now being used by researchers around the world.

TextAugment was designed to help developers create language technologies when they do not have the vast quantities of digital text normally required to train artificial intelligence systems. Created by Professor Vukosi Marivate, director of the African Institute for Data Science and AI (AfriDSAI) and holder of the Absa UP Chair of Data Science at the University of Pretoria, and researcher Tshephisho Sefara, the open-source software has recorded more than 286,000 downloads and is being used in work involving languages ranging from Swahili and Arabic to Uzbek.

The software’s growing reach was recognised on 16 July 2026, when Marivate and the TextAugment team received the inaugural NSTF-SADiLaR Research Software Award for Human Language Technologies. The award recognises their work in developing open-source software that helps researchers create synthetic training data for languages with limited digital resources.

The recognition adds to a series of recent honours for Marivate. On 19 May 2026, the Ga-Rankuwa-born computer scientist was awarded the Order of Mapungubwe in Silver for his contributions to data science, artificial intelligence and natural language processing.

The Presidency commended him for “his excellent contributions to data science, artificial intelligence (AI), and natural language processing (NLP) that have significantly advanced both national and continental technological capabilities”.

But his work is about more than teaching machines to understand language. It is about asking a bigger question: what happens when the technologies shaping the future cannot understand the languages spoken by millions of people?

“AI just existing is not enough for it to actually impact people,” he said. “What you have to do is to have the correct conditions, environments to amplify the good, and then you must also reduce the chances of the negatives that come with any type of technology coming into our world.”

African languages are described in artificial intelligence research as “low-resource” languages. Marivate says the term does not mean that these languages have few speakers. Rather, they lack the digital data and tools that allow machines to process them effectively.

“One of the biggest hurdles in training robust machine learning models for low-resource languages (like isiZulu, Sesotho, or Yoruba) is data scarcity. Traditional AI models require mountains of clean, labelled data to perform well. If that data doesn’t exist, these languages are effectively locked out of the modern AI revolution,” Marivate explained.

That gap becomes particularly visible when people interact with online systems. Sentiment analysis systems, for example, can determine whether an online review is positive, negative, or neutral. But systems developed primarily for English cannot automatically provide the same service for African languages.

“In South Africa we have 12 official languages, and we’re not actually catering for most of our people. To interact with all these online systems, you’re being forced to write in English because that’s how the processing happens.”

A language can have millions of speakers and still be poorly represented online. This is where one of Marivate’s most widely used research tools enters the picture.

TextAugment was developed to help researchers work around the shortage of high-quality language data.

Marivate explained that TextAugment was developed to help address the limited amount of training data available for some languages. The tool does this by synthetically expanding existing datasets, including by replacing words in sentences with synonyms while attempting to preserve their original meaning.

The technology can support translation, sentiment analysis, summarisation and other language-processing applications. It can also help researchers build tools such as spell checkers and systems capable of processing written language.

TextAugment was eventually packaged as open-source software so that researchers beyond Marivate’s own team could use it. It has since been used by researchers working with languages including Swahili, Uzbek and Arabic.

But increasing the amount of language data available to AI models is only part of the challenge. African languages carry cultural meanings, idioms and proverbs that cannot always be translated word-for-word. Marivate acknowledges that some augmentation methods have limitations because replacing a word with a synonym does not necessarily account for the context in which that word is being used.

“Writing in English and then translating in isiXhosa is not the same alternative as a person who writes in isiXhosa, because they write in a way that touches culture, geography, and lots of things that are inside there.”

For Marivate, African language technology should not simply mean taking English-language systems and making them produce African words. The technology must be capable of engaging with the cultural context in which those languages are spoken.

His research group is therefore also examining how AI models understand idioms and proverbs in African languages, including isiZulu and Tshivenda.

The challenge extends beyond how AI processes individual words. The broader question of whose languages and knowledge are represented in AI was also explored by University of Ghana Vice-Chancellor Professor Nana Aba Appiah Amfo at the University of Warwick’s Distinguished Africa Lecture 2026. In a lecture titled “Whose Language Counts? African Voices, Knowledge Systems, and the Future of AI,” Amfo examined the implications of language exclusion in emerging technologies and challenged prevailing assumptions about whose knowledge is represented in AI systems.

According to Amfo, the absence of African languages from digital ecosystems is not merely a technical limitation but a matter of representation and inclusion.

“When a language is absent from the digital corpus, it is not merely a translation problem. It is a visibility problem. It is a knowledge problem. And ultimately, it becomes a question of justice,” she said.

It would be easy to frame Africa’s AI development as something that began when ChatGPT arrived. Marivate rejects that narrative. He points to an African AI ecosystem that has been developing for more than a decade.

The Deep Learning Indaba began in 2016 as a gathering for people working in African AI research, development, and innovation. At its 2026 gathering in Lagos, Marivate said about 1,000 people participated.

The Masakhane Research Foundation, which emerged from that wider community in 2019, has grown to more than 3,000 researchers and innovators working on African languages, according to Marivate.

“The African AI ecosystem had not been waiting,” Marivate said. “They had not been waiting; these organisations started a decade ago.”

Investment is needed in the languages themselves, including the development of digital tools, research capacity and education. Governments and private companies also need to support locally developed technology rather than simply importing or reselling systems developed elsewhere.

“We do need to put investments in developing our languages, not for AI, but developing our languages, making sure there’s tooling that’s available to get people to be able to engage.”

He argues that without procurement and investment that support African technology companies and researchers, a sustainable ecosystem cannot emerge. That is particularly important because poorly performing AI systems can have consequences beyond inconvenience.

If an AI service works well in English but produces unreliable results in an African language, users may assume the technology is equally capable across languages when it is not. Marivate warns that this can lead to genuinely harmful experiences.

“If we’re not having an active ecosystem that is doing this work, there is a higher likelihood that people are going to have really bad experiences.”

The next phase of his work is therefore moving towards evaluating how well existing large language models actually perform in African languages.

Africa, he acknowledged, remains heavily dependent on technology developed elsewhere. But he argues that building local research and development capacity means the continent can anticipate technological change rather than simply react to it.

“We need to be building ourselves,” he said, “so that we can then learn how to do it better.”

TextAugment began as a response to the limited digital representation of African languages. Its adoption shows the wider opportunity: when African researchers build for the continent’s realities, they can produce technology with global value.

LEAVE A REPLY

Please enter your comment!
Please enter your name here

Popular Articles