← Latest papers
💬 NLP

Arguments that Alter Minds: LLM Rationales Sway Human (and LLM) Notions of Plausibility

This study demonstrates that both humans and LLMs are significantly swayed in their plausibility judgments of commonsense answers by LLM-generated arguments, highlighting a novel method for studying human cognition while raising concerns about the potential for AI to unduly influence human beliefs even in domains of common sense.

Original authors: Shramay Palta, Peter Rankel, Sarah Wiegreffe, Rachel Rudinger

Published 2026-06-23
📖 4 min read☕ Coffee break read

Original authors: Shramay Palta, Peter Rankel, Sarah Wiegreffe, Rachel Rudinger

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are playing a game where you have to guess what happens next in a silly story. For example: "If a person drops a glass, what happens?" You might guess, "It breaks," or "It bounces." Usually, "It breaks" is the obvious answer, and "It bounces" is a silly guess.

In this study, researchers asked a simple but scary question: Can a robot (an AI) talk its way into changing your mind about what is true, even when you know the answer yourself?

Here is how they did it and what they found, explained through a few simple analogies.

The Setup: The "Debate Club" for Common Sense

The researchers took 100 common-sense questions (like the glass example) and picked two answers for each:

  1. The Gold Answer: The one most people agree is correct.
  2. The Distractor: The silly, usually wrong answer.

Then, they used a powerful AI (GPT-4o) to write two types of arguments for every answer:

  • The "Pro" Argument: A speech explaining why this answer makes perfect sense.
  • The "Con" Argument: A speech explaining why this answer is a terrible idea.

They showed these questions and arguments to 3,000 real humans and 13,600 different AI models, asking them to rate how "plausible" (likely) the answer was on a scale of 1 to 5.

The Results: The AI's Magic Trick

1. The "Silly Answer" Gets a Boost

When the AI gave a "Pro" argument for the silly answer (the Distractor), humans and other AIs started thinking, "Hey, that actually makes sense!"

  • Analogy: It's like a smooth-talking friend convincing you that a trampoline is a great place to drop a glass. Even though you know glass usually breaks, the friend's logic makes you pause and give the "bounce" idea a higher score.

2. The "Obvious Answer" Gets a Knock-Down (For Humans)

This is the weirdest part. When the AI gave a "Pro" argument for the obviously correct answer (the Gold), humans actually lowered their ratings.

  • Analogy: Imagine you are 100% sure the glass will break. Then, a robot comes in and says, "Well, it is plausible that the glass breaks." You might think, "Wait, why are you even arguing this? It's obvious! If you have to explain it, maybe I'm wrong."
  • The researchers call this "underselling." By trying to prove the obvious, the AI accidentally made humans doubt their own confidence.

3. The "Con" Argument is a Sledgehammer

When the AI gave a "Con" argument (arguing against an answer), it was very effective at lowering scores.

  • Analogy: If the robot says, "The glass won't break because it's made of rubber," people immediately stop believing the "glass breaks" answer. The negative argument was stronger than the positive one.

4. Robots vs. Humans

The AI models (the robots) were even more easily swayed than the humans.

  • The "Echo Chamber" Effect: The AI models seemed to really like the arguments written by their "big brother" (GPT-4o). When GPT-4o wrote a "Pro" argument, other AI models believed it even more than humans did.
  • The Difference: Humans got confused when the AI tried to explain the obvious (lowering their score), but the other AIs just got more confident in the obvious answer when the AI explained it.

The "Anchoring" Rule

The study found a rule about how stubborn our minds are: The more sure you are to begin with, the harder it is to change your mind.

  • Analogy: If you are standing on a high cliff (a very high rating), it takes a huge wind (a strong argument) to push you off. If you are standing on flat ground (a low rating), a gentle breeze can knock you over easily.
  • The data showed that the initial rating was the strongest predictor of how much the rating would change.

The Big Takeaway

The paper concludes that LLMs (AI) can act like persuasive lawyers. They can talk humans into believing a silly idea is true, or talk humans into doubting a true idea.

Even in a field where humans are supposed to be the "experts" (everyday common sense), an AI's words can shift what we believe is plausible. The researchers warn that if an AI can trick us about a glass dropping, it might be able to trick us about much more serious, complex topics where we don't know the answers as well.

In short: Don't just listen to the answer; listen to who is arguing for it. A robot's smooth talking can make the wrong thing sound right, and the right thing sound doubtful.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →