← Latest papers
🤖 AI

DECEIVE-AFC: Adversarial Claim Attacks against Search-Enabled LLM-based Fact-Checking Systems

This paper introduces DECEIVE-AFC, an agent-based adversarial framework that effectively degrades the performance of search-enabled LLM fact-checking systems by manipulating claims to disrupt search behavior and reasoning without requiring access to internal model components or evidence sources.

Original authors: Haoran Ou, Kangjie Chen, Gelei Deng, Hangcheng Liu, Jie Zhang, Tianwei Zhang, Kwok-Yan Lam

Published 2026-03-17
📖 5 min read🧠 Deep dive

Original authors: Haoran Ou, Kangjie Chen, Gelei Deng, Hangcheng Liu, Jie Zhang, Tianwei Zhang, Kwok-Yan Lam

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a super-smart, all-knowing librarian named AI. This librarian doesn't just sit in a dusty library; they have a magical ability to instantly run to the internet, read thousands of news articles, and come back to tell you if a statement you made is True or False. This is what modern "Fact-Checking Systems" do. They are our digital guardians against fake news.

But what if someone could trick this librarian?

This paper, titled DECEIVE-AFC, introduces a new way to "hack" these smart librarians. Instead of trying to break the librarian's brain (which is hard), the hackers learned how to confuse the librarian's search engine and trip up their reasoning using clever word games.

Here is the story of how they did it, explained simply:

1. The Setup: The "Truth Detective"

Think of a Fact-Checking System as a Truth Detective.

  • Old Detectives (DNNs): These used to work like a detective with a single, old file cabinet. If the answer wasn't in the cabinet, they were stuck.
  • New Detectives (Search-Enabled LLMs): These are like detectives with a super-fast car and a direct line to the whole internet. They can drive out, find the latest evidence, and write a report. They are much smarter and harder to fool.

2. The Problem: The "Trojan Horse" Claim

The researchers asked: "Can we trick this super-smart detective?"

They realized that if you just say something obvious and wrong (like "The sky is green"), the detective will quickly find a photo of the blue sky and say, "Nope, that's false."

So, the researchers (the attackers) didn't try to lie. Instead, they tried to speak in riddles. They took a true statement and slightly twisted the wording so that:

  1. It still meant the same thing to a human.
  2. But it sounded like a completely different search query to the AI.

3. The Attack: DECEIVE-AFC (The "Deceiver")

The researchers built a robot agent called DECEIVE-AFC. Think of this robot as a Master of Disguise. Its job is to take a normal sentence and put on a "linguistic costume" to confuse the Truth Detective.

The robot uses three main tricks (analogies included):

  • Trick 1: The "Search Engine Misguidance" (The Wrong Map)

    • The Analogy: Imagine you tell a GPS, "Take me to the Big Apple." It knows you mean New York. But if you say, "Take me to the fruit that is red and round," the GPS might get confused and take you to an orchard instead.
    • The Attack: The robot swaps common words for weird, rare synonyms or adds extra, confusing background details. This tricks the AI into searching for the wrong things on the internet. Instead of finding a clear article about the event, it finds irrelevant, low-quality, or misleading articles.
  • Trick 2: The "Reasoning Disruption" (The Foggy Glasses)

    • The Analogy: Imagine reading a sentence that is grammatically correct but so complicated and full of double-negatives that your brain gets tired and makes a mistake. "It is not untrue that it is not the case that..."
    • The Attack: The robot makes the sentence structurally complex. It forces the AI to work harder to understand the logic. In the process of trying to untangle the knot, the AI's "brain" gets foggy, and it starts drawing the wrong conclusion.
  • Trick 3: The "Structural Complexity" (The Maze)

    • The Analogy: Instead of asking, "Is the door open?", you ask, "If the key is in the box, and the box is under the table, and the table is in the room, is the door open?"
    • The Attack: The robot turns a simple fact into a multi-step puzzle. The AI has to search for three different pieces of evidence and connect them. If it misses just one tiny piece in the middle of the chain, the whole answer collapses.

4. The Result: The Librarian Gets It Wrong

The researchers tested this on real-world Fact-Checking systems.

  • Before the attack: The AI was correct 78.7% of the time.
  • After the attack: The AI's accuracy dropped to 53.7%.

That's a huge drop! The AI started believing lies or rejecting truths. Even worse, when the AI got it wrong, it wrote a justification (an explanation) that sounded very confident and logical. It wasn't just saying "I don't know"; it was confidently saying, "This is false, because I found this article..." (even though the article was irrelevant).

5. Why This Matters

This paper is a wake-up call. It shows that even our most advanced, internet-connected AI fact-checkers are vulnerable. They aren't just vulnerable to obvious lies; they are vulnerable to cleverly phrased truths that confuse their search tools.

The Good News:
The researchers aren't trying to spread fake news; they are trying to break the system so we can fix it. By understanding how the "Master of Disguise" tricks the AI, developers can build better defenses. They can teach the AI to:

  • Ignore weird synonyms and focus on the core meaning.
  • Double-check its search queries.
  • Be more careful when the sentence structure gets too complicated.

In a nutshell:
The paper teaches us that in the battle against misinformation, how you ask a question is just as important as the answer. If you can phrase a question just right, you can trick even the smartest AI into getting the facts wrong. The goal now is to teach the AI to see through the disguise.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →