← Latest papers
💻 computer science

CommunityFact: A Dynamic, Multilingual, Multi-domain Benchmark for Misinformation Detection in the Wild

The paper introduces CommunityFact, a dynamic, multilingual, and multi-domain benchmark for misinformation detection that evaluates LLMs in real-world settings, revealing that while web access significantly improves performance, current models exhibit misaligned source-selection policies compared to human raters and highlighting Community Notes as a potential training signal for future verification systems.

Original authors: Sahajpreet Singh, Insyirah Mujtahid, Min-Yen Kan, Kokil Jaidka

Published 2026-05-29
📖 5 min read🧠 Deep dive

Original authors: Sahajpreet Singh, Insyirah Mujtahid, Min-Yen Kan, Kokil Jaidka

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Picture: A "Living" Test for AI Fact-Checkers

Imagine you are trying to teach a robot how to spot fake news. In the past, researchers gave the robot a static textbook full of old news stories and asked, "Is this true or false?" The problem is that the real world doesn't work like a textbook. News changes every minute, spreads in different languages, and moves across the internet like a wildfire. A robot that memorized an old textbook might fail miserably when faced with a brand-new rumor.

This paper introduces COMMUNITYFACT, a new, "living" test designed to see how well AI models can fact-check rumors in the wild, right now.

The Source: The "Community Watch"

Instead of using a textbook, the researchers built their test using X's Community Notes.

  • The Analogy: Think of Community Notes as a giant, global neighborhood watch. When someone posts a suspicious claim on X (formerly Twitter), volunteers step in to write notes correcting the misinformation. These notes are voted on by the community; if enough people from different viewpoints agree the note is helpful, it gets published.
  • The Innovation: The researchers didn't just copy the tweets and notes. They acted like translators and editors, turning those messy social media interactions into clean, standalone "quiz questions."
    • Example: Instead of a messy thread saying, "Did you see this video? It's fake!" they created a clean claim: "The video shows a real event that happened in 2024."
    • They labeled these claims TRUE or FALSE based on what the helpful Community Note said.

The Test: 16,000 Questions in 5 Languages

The final dataset is a massive quiz with 15,992 questions covering:

  • 5 Languages: English, Spanish, French, Japanese, and Portuguese.
  • 2 Topics: Politics and Finance.
  • The "Freshness" Factor: Unlike old tests, this one is built from recent data (2025). It's designed to be "refreshable," meaning the researchers can run the same process again next year to get a brand-new test without starting from scratch.

The Experiment: How Smart Are the AI Models?

The researchers put 10 different AI models (LLMs) through this test under four different conditions, like giving a student different tools during an exam:

  1. Closed-Book (No Internet): The AI has to rely only on what it memorized during training.
    • Result: The AI struggled. It was like trying to solve a math problem from 2025 using a textbook from 2020. It often got the answers wrong because the information was too new or too specific.
  2. Thinking Mode: The AI was told to "think step-by-step" before answering.
    • Result: This was a mixed bag. For some models, thinking helped; for others, it actually made them slower and more prone to errors. It wasn't a magic fix.
  3. Open-Book (Web Search): The AI was allowed to search the internet for answers.
    • Result: Huge improvement. Giving the AI access to the internet was the single biggest help. It proved that for fact-checking, having access to current information is far more important than just having a "big brain" (large model size).
  4. The "Hint" Mode (Evidence-Guided Search): The AI was allowed to search the web, but the researchers also gave it a list of specific links (URLs) that the human Community Notes volunteers had already used as proof.
    • Result: This was the most interesting finding.
      • The Misalignment: The AI models, when searching the web on their own, tended to visit different websites than the human volunteers did. They were looking in the wrong "neighborhoods."
      • The Fix: When the researchers pointed the AI toward the specific links humans used, the AI's accuracy went up significantly.
      • The Lesson: It's not just about having the internet; it's about knowing where to look. The AI needs to learn to trust the same sources humans trust (like government records or major news outlets) rather than just clicking the first search result.

Key Takeaways in Plain English

  • Static tests are outdated: You can't test a fact-checker with old news. The test needs to be as fast and messy as the internet itself.
  • Access beats memory: An AI with a small brain but internet access will beat an AI with a giant brain but no internet when it comes to checking new facts.
  • AI and Humans look in different places: Even when AI is allowed to search the web, it often picks different sources than humans do. If we want AI to be a good fact-checker, we need to teach it to follow the same "trail of evidence" that human experts follow.
  • The "Community" is the teacher: The paper suggests that instead of just using Community Notes to fix posts after they happen, we should use them to train AI to know where to look for the truth in the first place.

What the Paper Does Not Say

  • It does not claim that AI is now perfect at stopping misinformation.
  • It does not say that we should replace human fact-checkers with AI.
  • It does not claim that this test works for every type of lie (like deepfake videos), as this specific test focused on text-based claims.

In short, COMMUNITYFACT is a new, dynamic scoreboard that shows us exactly where AI fact-checkers are failing and how giving them the right "map" (human-vetted sources) can help them find the truth.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →