← Latest papers
🤖 machine learning

Predicting, Evaluating, and Explaining Top Misinformation Spreaders via Archetypal User Behavior

This paper proposes a framework for predicting, evaluating, and explaining top misinformation spreaders by modeling three distinct user archetypes (amplifiers, super-spreaders, and coordinated accounts), demonstrating that super-spreader traits dominate top rankings while temporal dynamics and explainable AI enhance predictive accuracy and interpretability for proactive content moderation.

Original authors: Enrico Verdolotti, Luca Luceri, Silvia Giordano

Published 2026-08-18
📖 5 min read🧠 Deep dive

Original authors: Enrico Verdolotti, Luca Luceri, Silvia Giordano

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

On social media, information travels at the speed of a thought, but not everyone who shares it plays the same role. Some people create the original stories, while others simply pass them along, and a few work together to make sure a message reaches as many eyes as possible. When false information spreads, it is rarely the fault of a single person; instead, it is the result of a complex web of behaviors where different types of users amplify, coordinate, and launch narratives. Understanding who these people are and how they operate is crucial because false stories can distort public opinion, erode trust in science, and cause real-world harm. The challenge for platforms is that by the time a false story is flagged and removed, it has often already traveled too far to be stopped. To stop the spread before it gains momentum, researchers need a way to identify the specific accounts most likely to cause damage, not by guessing their intentions, but by watching how they actually behave.

A team of researchers set out to solve this problem by looking at the digital footprints left behind by users who spread low-credibility content. They focused on three distinct patterns of behavior, or "archetypes," that these users tend to follow. The first group, which they call amplifiers, rarely creates their own posts. Instead, they act as a megaphone, constantly resharing content made by others to boost its reach. The second group, known as super-spreaders, are the original creators of viral content. They write posts that consistently get shared thousands of times, acting as the spark that ignites a fire. The third group consists of coordinated accounts, which are groups of users working in sync to push the same message at the same time, creating an illusion of widespread public agreement. The researchers realized that these roles are not always separate; a single user might act as a super-spreader for one story and an amplifier for another, or a group might coordinate their actions to mimic a single influential voice.

To test how well they could predict who would spread misinformation next, the researchers analyzed two massive collections of data from Twitter during the pandemic. One dataset contained conversations from Italy, while the other was a global collection of posts in many languages. They first labeled the content based on the trustworthiness of the websites linked in the posts, using a system that rates news domains from zero to one hundred. They then built different mathematical models to rank users based on the three behavioral archetypes. Some models counted how many times a user reshared false stories, while others looked at how early a user shared a story after it was posted, or how closely a user's activity matched that of others in a coordinated group. They also created a new, time-sensitive version of a classic influence score that accounts for how a user's activity changes over days and weeks, rather than just looking at a single snapshot in time.

The results showed that the most effective way to find the biggest troublemakers was to look for the super-spreaders. When the researchers ranked users based on their ability to generate original content that others would reshare, these accounts consistently appeared at the very top of the list. These are the individuals responsible for the initial burst of virality that allows false narratives to take hold. However, as the researchers looked further down the list of influential users, the picture became more complex. In the middle and lower ranks, the behavior of amplifiers and coordinated groups became much more prominent. This suggests that while a few super-spreaders start the fire, it is the amplifiers who keep it burning and the coordinated groups who fan the flames across the network.

To get a clearer picture of how these different behaviors work together, the team trained a machine learning model that combined all three archetypes into a single system. They then used a technique that explains how the model makes its decisions, allowing them to see which behaviors mattered most. They found that the model relied heavily on the number of times a user reshared low-credibility content, a trait associated with amplifiers. This was surprising because one might expect the creators of viral content to be the most important signal. Instead, the data showed that the act of resharing is a powerful indicator of risk. The model also confirmed that the super-spreaders, identified by their ability to generate consistent engagement, were critical, especially among the very top-ranked users. As the ranking moved down, the influence of coordinated behavior and the frequency of resharing grew stronger, revealing a layered ecosystem where different types of actors play different roles in the same problem.

The study also tested how well these methods work when there is very little data to go on, simulating a scenario where a platform needs to make a decision quickly after a new story appears. They found that the time-sensitive influence score remained stable and accurate even with limited information, outperforming older methods that did not account for the timing of user activity. The machine learning model, however, struggled when trained on very small amounts of data, suggesting that while complex models are powerful, they need enough history to learn the patterns. The researchers concluded that the best approach for platforms is likely a combination of these tools: using simple, time-aware scores to quickly identify the most dangerous accounts, and using more detailed models to understand the full range of behaviors that sustain the spread of misinformation. By focusing on these observable patterns rather than trying to guess a user's intent, platforms can intervene earlier, targeting the specific accounts that drive the most harm before false information becomes impossible to contain.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →