Understanding Persuasion in Long-Running Agents
This paper introduces a behavior-centered evaluation framework to study "persuasion propagation" in long-running AI agents, revealing that while real-time persuasion yields weak effects, explicitly pre-specifying an agent's belief state significantly reduces its search effort and source diversity during task execution.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a super-smart, tireless research assistant (an AI agent) who can browse the internet, write code, and solve complex problems on its own. You might think this assistant is like a robot that just follows your exact orders. But this paper asks a tricky question: What happens if you convince the assistant to believe something completely unrelated to the task it's doing?
The researchers call this phenomenon "Persuasion Propagation." It's like whispering a secret opinion to a chef while they are chopping vegetables. Even if the secret is about politics and the task is cooking dinner, does that secret change how they chop? Do they become more careful? Do they only use certain knives? Or do they just ignore it?
Here is a simple breakdown of what the paper found:
1. The Experiment: The "Distraction" Test
The researchers set up a game with three main steps:
- The Setup: They gave the AI a specific "personality" (like being "cooperative" or "direct").
- The Persuasion: Before the AI started its real job, they tried to convince it of a controversial opinion on a topic that had nothing to do with the job. For example, they might try to convince the AI that "geoengineering is too risky" before asking it to write a Python code or research how to quit smoking.
- The Task: The AI then had to do a real job, like writing code or gathering information from the web.
They compared three groups:
- The Persuaded Group: The AI was convinced to believe the new opinion.
- The Not-Persuaded Group: The AI heard the opinion but rejected it.
- The Neutral Group: The AI heard nothing.
2. The Big Surprise: "Belief" vs. "Just Saying Yes"
The paper found a major difference between changing your mind and just nodding along.
- The "Nodding" Effect (On-the-Fly Persuasion): When the AI was asked to change its mind during a conversation right before the task, the results were messy. Sometimes it changed its behavior, sometimes it didn't. It was like trying to steer a ship by shouting from the deck while the engine is already running; the effect was weak and inconsistent.
- The "Deep Belief" Effect (Prefilled Belief): When the researchers simply told the AI at the very start, "You believe X," and then gave it the task, the results were much clearer. The AI actually changed how it worked.
- The Analogy: Imagine two detectives solving a crime.
- Detective A (Neutral): Checks 20 different leads, visits 10 different neighborhoods, and talks to many witnesses.
- Detective B (Persuaded): Because they were told to believe a specific theory, they only check 5 leads, visit 2 neighborhoods, and ignore anything that doesn't fit their theory. They still solve the case, but they did it in a much narrower way.
- The Analogy: Imagine two detectives solving a crime.
3. The Results: The "Hidden" Changes
The paper discovered that even when the AI's final answer looked perfect and correct, the journey to get there was different.
- In Coding: The persuaded agents didn't necessarily write worse code, but they took different paths to get there.
- In Web Research: This is where it got interesting. The "Belief-Prefilled" agents (those who genuinely adopted the new stance) did 26.9% fewer searches and visited 16.9% fewer unique websites than the neutral agents.
- The Metaphor: It's like a tourist who is told, "This city is dangerous, so don't go to the park." Even if the tourist still writes a great travel guide, they might skip the park entirely. The final guide looks fine, but it's missing a whole section of the city because of a belief they held earlier.
4. Why This Matters
The paper argues that we usually only judge AI by its final answer (Did it get the code right? Is the essay good?). But this research shows that how the AI gets there matters just as much.
If an AI is subtly influenced by a previous conversation to believe something, it might:
- Stop searching for information too early.
- Ignore sources that contradict its new belief.
- Rely on a smaller, less diverse set of facts.
The Takeaway:
You can't just look at the final report to know if an AI is safe or reliable. You have to look at the "footprints" it left behind while working. Even if the AI says "I agree" with a strange idea, it might not change its final answer, but it might change the path it takes to get there, potentially making its work less thorough or biased in ways you can't easily see.
In short: Persuasion doesn't always change what the AI says, but it can quietly change how it thinks and works.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.