Wisdom of the LLM Crowd: A Large Scale Benchmark of Multi-Label U.S. Election-Related Harmful Social Media Content
This paper introduces USE24-XD, a large-scale, multi-label dataset of nearly 100k U.S. election-related social media posts annotated by a "wisdom-of-the-crowd" of six LLMs and validated against human raters to enable scalable detection of harmful content while revealing systematic subjectivity in labeling behaviors.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the internet during an election is like a massive, chaotic town square. Everyone is shouting, posting signs, and spreading stories. Some stories are true, some are jokes, some are scary rumors, and some are outright lies designed to make people angry or afraid.
The researchers in this paper wanted to build a super-smart "filter" to sort through this noise and find the dangerous stuff (like lies, hate speech, or scary rumors) before it spreads too far. But sorting through 100,000 posts by hand would take a team of people years to do. So, they asked a new question: Can we use Artificial Intelligence (AI) to do the sorting for us?
Here is the story of how they tried to answer that, explained simply:
1. The Problem: The "Needle in a Haystack"
The team collected nearly 100,000 posts from X (formerly Twitter) during the 2024 U.S. election. They wanted to tag them into five specific "buckets" of bad content:
- Conspiracy: "The election was stolen by aliens!" (Secret plots).
- Sensationalism: "BREAKING: The world is ending tomorrow!" (Scary, over-the-top fear).
- Speculation: "I bet the candidate will drop out next week." (Wild guesses without proof).
- Hate Speech: Attacking people based on who they are.
- Satire: Funny jokes that look like news but aren't (like The Onion).
2. The Experiment: The AI "Book Club"
Instead of hiring one person to read every post, the researchers hired six different AI models (think of them as six different super-smart book club members with different personalities).
- They gave each AI the same 1,000 posts and asked: "Does this post fit into any of our five bad-content buckets?"
- They also hired 34 real humans (via a crowdsourcing website) to do the same job, just to see how the AI compared to real people.
3. The Big Surprise: The AI "Crowd" Was More Consistent
Usually, we think humans are the gold standard for judging things like "is this a joke?" or "is this hate speech?" because humans understand nuance. But the researchers found something interesting:
- The Humans were Messy: When different humans looked at the same post, they often disagreed. One person might think a post was a "joke" (Satire), while another thought it was "hate speech." This is because humans have their own political biases and backgrounds.
- The AI was Steady: The six AIs agreed with each other much more often than the humans did. They were like a well-oiled machine.
- The "Wisdom of the AI Crowd": The researchers realized that if they took the majority vote of the three best AIs, they got a result that was even better than any single AI or any single human. It was like asking three experts for their opinion and going with the answer two of them agreed on.
4. The Human Factor: Politics Changes How We See Things
The study also looked at why the humans disagreed. They found that who you are changes what you see.
- If a human annotator was very conservative, they were more likely to label a post as "Hate Speech."
- If they were a "Centrist" or independent, they were more likely to label it as "Speculation" (just a wild guess).
- The Metaphor: Imagine looking at a painting through different colored glasses. A person wearing red glasses sees the whole painting as red; a person wearing blue glasses sees it as blue. The AI didn't have colored glasses, so it saw the "shape" of the content more consistently, even if it missed some of the subtle human context.
5. The Result: A New Tool for the Future
The team created a massive, clean dataset called USE24-XD.
- What they found: About 60% of the election posts they looked at had at least one "bad" label. This means harmful content isn't a rare glitch; it's a huge part of the conversation.
- The Win: They proved that using a "team" of AIs is a cheap, fast, and surprisingly accurate way to clean up social media. It's not perfect (it sometimes misses subtle jokes), but it's way faster and cheaper than hiring armies of humans.
The Takeaway
Think of this study as building a new kind of sieve.
In the past, we tried to sift through the internet's noise with tiny, hand-made sieves (humans), which was slow and inconsistent. This paper shows that we can build a giant, automated sieve made of AI. While the AI sieve isn't perfect, if you use a few different sieves together and take the majority vote, you catch almost all the "bad apples" without needing to hire a thousand people to pick through the fruit basket.
This gives researchers and platforms a powerful new tool to understand and fight misinformation before it causes real-world damage.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.