News
Dartmouth Researchers Build First Propaganda Detection Tool for Kinyarwanda

A Dartmouth College team has built KinyaProp, the first dataset for detecting propaganda in Kinyarwanda, the language of 15 million people in Rwanda that models like Claude and ChatGPT have so far struggled to handle for manipulation detection.
Contents
Language models that can write flawless sentences in Kinyarwanda turn out to be helpless when asked to spot propaganda in that same language. A team of researchers at Dartmouth College has built the first public dataset for detecting manipulation techniques in Rwanda's national language, hoping it will become a template for hundreds of other African languages largely overlooked in AI development.
A blind spot in large models
The team tested how widely used language models, including Claude and ChatGPT, handle propaganda detection in Kinyarwanda. The result was unambiguous: even though these systems can generate fluent text in the language, they proved, in the researchers' words, essentially non-functional when it came to recognizing manipulative techniques.
If you ask them to complete a task that requires understanding, such as detecting propaganda, they fail significantly - Fabrice Niyigaba, Dartmouth College
That distinction matters, because the ability to generate fluent text in a given language is often mistaken for actually understanding it. A model can correctly conjugate Kinyarwanda verbs while completely missing that a passage is stoking fear, playing on prejudice, or deliberately sowing confusion among readers, which are exactly the classic propaganda techniques KinyaProp is designed to teach models to recognize.
How the dataset was built
Building KinyaProp did not involve simply machine-translating existing English-language propaganda datasets. Three native Kinyarwanda speakers spent three months manually analyzing more than 600 news articles, assigning each to categories drawn from sixteen recognized manipulation techniques. Researcher Ivory Yang says it was only through this ground-up work that the team saw how different local disinformation patterns are from those known in English.
We learned that misinformation is expressed much differently in Kinyarwanda - Ivory Yang, Dartmouth College
After integrating KinyaProp with existing commercial models, their accuracy in detecting propaganda in Kinyarwanda rose markedly, approaching the level those same systems achieve in languages well represented in training data, such as English or Arabic.
The shadow of the 1994 genocide
This work's context is no accident. During the 1994 genocide in Rwanda, more than 800,000 people, mostly Tutsi civilians, were killed by Hutu extremists over the course of a hundred days. News media, including radio, was one of the key tools of the mass manipulation that preceded and fueled the wave of violence.
The news media was one of the platforms used for large-scale manipulation - Fabrice Niyigaba, Dartmouth College
That history gives the ability to quickly and systematically detect propaganda in Kinyarwanda significance for Rwanda that goes well beyond an academic exercise. Project advisor Soroush Vosoughi notes that AI systems which appear competent in Western languages can hide serious gaps in places no one has systematically tested them before.
Systems that appear capable in English may leave serious blind spots - Soroush Vosoughi, Dartmouth College
Plans to expand to other Bantu languages
The authors see KinyaProp as a starting point, not a finished product. The team says it plans to build similar datasets for Swahili and Zulu, two Bantu languages with tens of millions of speakers each. The methodology developed for Kinyarwanda, built on manual labeling by native speakers rather than machine-translating ready-made English datasets, is meant to serve as a template for the entire Bantu language family.
To a Polish reader the story may seem geographically distant, but it illustrates a pattern well known in Central Europe too: AI models trained mainly on English-language data systematically perform worse on languages with fewer digital text resources. Polish is nowhere near as underrepresented as Kinyarwanda, but research into detecting disinformation in less common languages has a direct bearing on how effectively global AI platforms handle content moderation outside the biggest markets.
The work by Niyigaba, Yang and Vosoughi also feeds into a broader body of research on the limits of large language models in tasks that require deep understanding of cultural context, not just grammatical correctness. The next test for KinyaProp will be whether commercial language model developers choose to build similar datasets permanently into their safety systems, rather than treating it as a one-off research exercise.

