← Latest papers
🤖 AI

When AI Persuades: Adversarial Explanation Attacks on Human Trust in AI-Assisted Decision Making

This paper introduces adversarial explanation attacks (AEAs) that manipulate LLM-generated explanations to preserve human trust in incorrect AI predictions, revealing through a study of over 200 participants that users are highly vulnerable to such cognitive manipulation, particularly when explanations mimic expert communication or target less educated and more trusting individuals.

Original authors: Shutong Fan, Lan Zhang, Xiaoyong Yuan

Published 2026-05-18
📖 4 min read☕ Coffee break read

Original authors: Shutong Fan, Lan Zhang, Xiaoyong Yuan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are asking a very smart, well-spoken robot for advice on a difficult problem, like choosing a medical treatment or an investment. You trust the robot because it explains its reasoning clearly.

This paper reveals a scary new trick: A bad actor doesn't need to break the robot's brain to fool you; they just need to change how the robot talks.

Here is the breakdown of the study, using simple analogies:

1. The New "Hacker" Move

Usually, when we think of hacking AI, we imagine someone sneaking into the computer code to make the robot give the wrong answer.

  • The Old Way: Breaking the lock on the front door.
  • The New Way (Adversarial Explanation Attacks): The robot still gives the wrong answer, but the hacker changes the script the robot reads. The robot now sounds like a calm, confident, highly educated expert. It uses fancy words, cites fake statistics, and speaks in a neutral, professional tone.

The paper calls this an Adversarial Explanation Attack (AEA). The attacker isn't changing the math; they are just "dressing up" the wrong answer in a tuxedo to make it look trustworthy.

2. The "Trust Miscalibration" Trap

The researchers wanted to see if this "fancy suit" actually works on humans. They set up a game with over 200 people.

  • The Setup: People were asked to solve problems with AI help. Sometimes the AI was right and explained it simply. Other times, the AI was wrong, but the explanation was crafted to sound like a top-tier expert (using citations, neutral tones, and logical steps).
  • The Result: The people couldn't tell the difference. They trusted the "wrong but fancy" answer almost as much as the "right and simple" one.
  • The Metaphor: It's like a magician. If a magician gives you the wrong card but explains the trick with such smooth confidence and scientific jargon that you feel silly for doubting him, you will still believe him. The study found that persuasion is more powerful than truth for many users.

3. Who Gets Fooled the Most?

The study found that not everyone falls for the trick equally. It's like a lock that is easier to pick for some keys than others.

  • The "Hard" Tasks: When the problem is really hard (like advanced physics or complex business strategy), people are more likely to trust the "expert" voice because they don't feel confident enough to check the work themselves.
  • Fact-Heavy Fields: In fields like medicine or business, where facts and numbers seem to rule, the fake "expert" voice works best. People assume, "If it sounds like a doctor or a banker, it must be true."
  • The Vulnerable Groups:
    • Younger people were more easily swayed than older adults.
    • People with less formal education trusted the fancy explanation more than those with advanced degrees (who tended to double-check the logic).
    • People who already loved AI were the most vulnerable. If you already trust the robot, you are less likely to question its fancy speech.

4. The "Expert" Costume

What made the trickiest explanations work? They didn't sound like a salesperson trying to hype a product. They sounded like a boring, neutral professor.

  • The Winning Formula: A calm tone + real-sounding citations (fake or real) + step-by-step logic.
  • The Losers: Explanations that were too emotional, too agreeable ("You're absolutely right!"), or too flashy actually made people distrust the AI more.

5. Does It Wear Off?

The researchers watched what happened over time.

  • Short Term: If you see one fake explanation, you might still trust it.
  • Long Term: If you see a lot of fake explanations in a row, your trust starts to crumble. It's like a "Boy Who Cried Wolf" scenario. Eventually, people start to realize, "Wait, this robot keeps getting things wrong even when it sounds smart."
  • However, if the robot starts giving correct answers again, trust bounces back quickly.

The Bottom Line

This paper argues that security isn't just about protecting the code; it's about protecting the conversation.

If an attacker can control how an AI explains its answer, they can make you trust a wrong decision without ever touching the AI's actual brain. The study concludes that we need to be careful not to be fooled by a fluent, confident voice when the underlying facts are shaky. We need to look at the content, not just the style.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →