BERTopic-Virality Prioritisation: A Scalable Framework for Thematic and Comparative Analysis of COVID-19 and Monkeypox Misinformation on Twitter
This paper introduces BERTopic-VP, a scalable framework that integrates contextual embedding clustering with a virality-prioritisation layer and a hybrid misinformation detector to enable the early identification and comparative analysis of high-impact health misinformation narratives across different pandemics.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
When a new disease emerges, the world faces two simultaneous crises: the virus itself and the flood of false information that spreads alongside it. This second crisis, often called an infodemic, can be just as dangerous as the biological one. On social media platforms like Twitter, where messages are short and move at lightning speed, false stories about cures, origins, and safety measures can gain traction faster than official health guidance. Traditional methods for tracking these stories often fail because they rely on counting how often specific words appear. In the chaotic, rapidly changing language of social media, where people use slang, hashtags, and emojis, simple word counts break down. They miss the deeper meaning behind a post, grouping unrelated ideas together or failing to spot a new, dangerous rumor until it has already spread too far.
To solve this, researchers at Manchester Metropolitan University developed a new way to listen to the digital noise. They created a system that does not just count words but understands the meaning behind them, while also paying attention to how fast a story is spreading. Their goal was to build a tool that could help public health officials spot the most dangerous lies early, even when they are just starting to circulate. By combining a deep understanding of language with a measure of how much people are sharing a post, they created a framework that can sift through batches of tweets to find the specific narratives that pose the greatest risk to public health.
The researchers applied this new system, which they named BERTopic-VP, to real data from two major health crises: the global pandemic caused by the novel coronavirus and the 2022 outbreak of monkeypox. They fed the system batches of tweets related to these events, processing them in groups ranging from 10,000 to 100,000. Instead of looking for specific keywords, the system first grouped the tweets into clusters based on their actual meaning. It used advanced language models to understand that two posts saying different things could still be talking about the same core idea, such as distrust in government or fear of vaccines. Once the stories were grouped, the system added a second layer of analysis: it looked at how much engagement each story received. It counted how many people liked, shared, or replied to the posts within each group. This allowed the system to rank the stories, highlighting the ones that were not just popular, but were actively spreading.
The results revealed distinct differences in how misinformation behaved during these two outbreaks. During the coronavirus pandemic, the false stories were highly varied and complex. They covered a wide range of topics, from conspiracy theories about the virus's origin to debates over mask efficacy and lockdown policies. The language used in these posts was often dense and difficult to read, mimicking the tone of an expert to sound more credible. In contrast, the misinformation surrounding the monkeypox outbreak was much more focused. The false narratives centered heavily on a few specific themes, such as failures in vaccine distribution and deep distrust in health authorities. These stories were linguistically simpler, easier to read, and carried a much stronger emotional charge, particularly fear and anger.
The study found that the system was effective at its job. It successfully identified clusters of misinformation with high viral potential, flagging the top one percent of spreading stories for human review. In the monkeypox data, where the researchers had access to actual sharing numbers, the system produced semantically interpretable topics and identified the most active groups based on observed engagement. For the coronavirus data, where sharing numbers were not available, the system used a clever workaround: it learned from the monkeypox data what linguistic and psychological features usually lead to high sharing, and then applied that knowledge to predict which coronavirus stories were likely to spread. This allowed the system to prioritize risky narratives even without direct access to sharing counts.
Crucially, the researchers showed that a story's ability to spread is not always linked to whether it is true or false. The system found many highly viral groups of posts that were actually factual, such as news updates or official health advice. This is why the system is designed to work in two steps: first, it finds the stories that are spreading the fastest; second, it uses a separate check to determine if those stories are true. This combination allows health officials to focus their limited time on the specific lies that are gaining the most momentum, rather than trying to debunk every false claim or ignoring the fast-moving ones that happen to be true.
The study also highlighted that the emotional tone of a story is a powerful driver of its spread. The most viral misinformation was often written in simple language that was easy to understand quickly, paired with strong emotions like fear and anger. This suggests that lies that are easy to digest and make people feel intense emotions are the most likely to be shared. By identifying these patterns, the new framework offers a practical way for public health teams to monitor the digital landscape. It acts as an early warning system, surfacing the specific, high-risk narratives that need immediate attention before they can cause widespread harm. The work demonstrates that by understanding both the meaning of the words and the mechanics of how they travel, we can better protect communities from the dangers of misinformation during health emergencies.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.