CORE: Conflict-Oriented Reasoning for General Multimodal Manipulation Detection
The paper proposes CORE, a conflict-oriented reasoning framework that leverages a newly constructed Conflict Attribution Corpus to enhance multimodal large language models' ability to detect generalizable fake news by identifying intrinsic semantic or physical inconsistencies, thereby achieving superior performance in few-shot and zero-shot scenarios compared to existing state-of-the-art methods.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the internet is a giant, bustling marketplace where people trade stories, pictures, and news. Lately, a new kind of merchant has arrived: Generative AI. These merchants are so skilled at painting pictures and writing stories that they can create "fake news" that looks and sounds completely real. It's like a master forger who can paint a portrait of a President winning a football trophy, even though that never happened.
The problem is that the old security guards (current detection tools) are trained to spot specific types of fakes. If they know how to spot a "crooked smile," they might catch that. But if the forger changes the trick to "weird lighting," the old guards are confused. They need to be retrained from scratch every time the forger changes their style, which takes too much time and data.
Enter CORE: The "Conflict Detective"
The authors of this paper, Jinjie Shen and his team, realized that all fake news has one thing in common: it doesn't make sense. It has an internal argument.
Think of it like a detective solving a mystery.
- The Old Way: The detective memorizes the face of every known criminal. If a new criminal shows up with a different face, the detective fails.
- The CORE Way: The detective learns the logic of a crime. They ask: "Does this story contradict what I know about the world?"
CORE (Conflict-Oriented REasoning) teaches a smart computer brain (called a Multimodal Large Language Model) to be that logic-based detective. Instead of memorizing specific fakes, it learns to spot conflicts.
How CORE Works: The Three-Step Training Camp
To teach this computer brain, the researchers built a special training program with three parts:
1. The "Conflict Library" (The Conflict Attribution Corpus)
First, they needed a textbook. They created a massive library of 14,000 fake news stories. But they didn't just label them "Fake." They broke them down like a forensic expert:
- The Conflict Factor: What is the contradiction? (e.g., "The text says 'President' but the image shows a football award.")
- The Conflict Source: Where did the lie come from? (e.g., "The text is lying," or "The image is lying," or "This contradicts real-world facts.")
- Analogy: It's like giving a student a test where they don't just mark "Wrong," but write down exactly why the answer is wrong and which part of the question caused the error.
2. The "Bridge Builder" (Modality Bridging Pre-Training)
Computers often struggle to connect what they "see" (images) with what they "read" (text). Sometimes, the computer sees a picture of a football field but doesn't really "understand" the text saying "President."
- CORE builds a bridge between these two worlds. It trains the computer to look at a picture and immediately understand the text description, and vice versa, so they speak the same language.
3. The "Logic Gym" (Conflict Perception Training)
Now, the computer goes to the gym. Using the "Conflict Library," it practices spotting contradictions.
- The computer learns to push "conflicting ideas" far apart in its mind. If it sees "President" and "Football Award" together, it learns to scream, "Wait! These two things don't belong together!"
- It learns to say: "This is fake because the text claims X, but the image shows Y, and my knowledge says X and Y can't happen at the same time."
Why This is a Game-Changer
The paper shows that this new approach is incredibly fast and flexible.
The "Few-Shot" Superpower: Imagine you want to teach a guard to spot a new type of fake news (e.g., deepfakes of a specific new celebrity).
- Old Methods: Need thousands of examples of that specific celebrity to learn.
- CORE: Only needs a handful of examples (sometimes just 100, or even zero!). Because it already understands the logic of conflict, it can instantly apply that logic to the new celebrity. It's like teaching someone to play chess by explaining the rules, rather than making them memorize every possible game.
The "Zero-Shot" Miracle: In some tests, CORE could spot brand-new types of fake news it had never seen before, just by reading the news and asking, "Does this make sense?"
The Bottom Line
The paper claims that by focusing on the intrinsic conflict (the logical contradiction) rather than the specific visual trick (the forgery technique), CORE creates a detector that is robust, general, and doesn't need a massive library of data to learn new tricks. It turns the computer into a critical thinker that asks, "Does this story hold up?" rather than just "Have I seen this picture before?"
The researchers have made their "Conflict Library" and their code available to the public, hoping others can use this "logic-based detective" to keep the internet honest.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.