← Latest papers
🤖 AI

When Roles Fail: Epistemic Constraints on Advocate Role Fidelity in LLM-Based Political Statement Analysis

This paper presents the first systematic empirical test of role fidelity in multi-agent LLM pipelines for political discourse analysis, revealing that models frequently fail to maintain assigned adversarial roles due to mechanisms like Epistemic Role Override, with significant variations in failure modes and fidelity depending on the specific model, language, and fact-checking tools used.

Original authors: Juergen Dietrich

Published 2026-05-01
📖 5 min read🧠 Deep dive

Original authors: Juergen Dietrich

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are organizing a town hall meeting to discuss a controversial new law. To get a fair picture, you hire three different experts to give their opinions:

  1. The Critic: Someone who looks for flaws and tries to tear the argument down.
  2. The Neutral: Someone who just lists the facts without taking sides.
  3. The Advocate: Someone who tries to find the best possible interpretation, looking for the "good faith" intent behind the argument.

This is how the TRUST system works. It uses three different AI models (robots) to play these roles simultaneously. The idea is that by forcing the robots to argue from these specific angles, you get a complete, balanced view of political statements.

However, this paper asks a simple but scary question: What if the robots forget their jobs?

The Core Problem: When Robots "Break Character"

The author, Juergen Dietrich, tested whether these AI robots could actually stay in their assigned roles. He found that they often fail, but not randomly. They fail in very specific, predictable ways when they encounter facts that contradict their instructions.

Think of it like an actor in a play who is told to play a villain. If the script says the villain is a murderer, the actor plays the villain. But if the script suddenly reveals the villain is actually a saint, the actor might stop playing the villain and start acting like a hero, even though the director told them to stay in character.

The Two Ways Robots Fail

The paper identifies two main ways this "breaking character" happens, which the author calls Epistemic Role Override.

1. The "Floor" Effect (The Advocate's Dilemma)
Imagine the Advocate robot is trying to be nice and find the good in a statement.

  • Scenario: The statement says, "Eating rocks cures cancer."
  • The Conflict: The robot's "fact-checker" (a separate tool) immediately says, "No, that is false and dangerous."
  • The Result: The Advocate robot gives up. It cannot pretend to be charitable toward something it knows is factually wrong. It drops its "nice" role and suddenly starts sounding like the Critic, pointing out the lie.
  • The Metaphor: It's like a lawyer trying to defend a client who is caught red-handed holding the murder weapon. No matter how hard the lawyer tries to be "charitable," the facts are so overwhelming that the lawyer's defense collapses. The paper calls this the Epistemic Floor Effect: the facts create a hard floor below which the robot cannot go.

2. The "Prior" Conflict (The Critic's Dilemma)
Now imagine the Critic robot is trying to be tough and find flaws.

  • Scenario: The statement says, "The sun rises in the east."
  • The Conflict: The Critic is told to find problems, but the fact-checker says, "This is 100% true."
  • The Result: The Critic robot struggles. Its internal knowledge (trained on all of human history) screams that this is true, so it can't really criticize it. Sometimes, it accidentally switches roles and starts sounding like the Advocate, validating the statement instead of attacking it.
  • The Metaphor: It's like a detective hired to prove a suspect is innocent, but the evidence is so clear that the detective accidentally starts proving the suspect is guilty.

The "Mirror" Discovery

The most fascinating finding is that these two failures are mirror images of each other.

  • If the facts are bad, the "Nice" robot breaks.
  • If the facts are good, the "Mean" robot breaks.
    The paper calls this Epistemic Role Override. The robot's internal knowledge of "what is true" is stronger than the instructions telling it "what role to play."

The Robot Showdown: Mistral vs. Claude

The author tested two different AI models to see which one was better at staying in character: Mistral Large and Claude Sonnet.

  • Claude Sonnet: This robot was very bad at staying in the "Advocate" role. When it saw a factually wrong statement, it completely flipped its personality, turning from a supporter into a harsh critic. It was like a chameleon that changed colors too easily.
  • Mistral Large: This robot was much better. When it saw a wrong statement, it didn't flip to the opposite side; it just quietly stopped trying to be an advocate. It "gave up" the role but didn't become the enemy.
  • The Winner: Mistral was significantly more reliable (about 28% better) at sticking to its assigned job than Claude.

The Language and Fact-Check Twist

The study also looked at whether the language (English vs. German) or the "fact-checker" tool used mattered.

  • Language: The robots behaved similarly in both English and German, which is good news.
  • Fact-Checkers: This is where it got tricky. If you used a specific fact-checking tool (Perplexity) with the Claude robot on German statements, the robot failed even more often. It's like using a magnifying glass that is too strong; it makes the robot see so many errors that it can't do its job at all.

The Big Takeaway

The paper concludes that if you build a system that relies on AI robots to play different roles (like a debate team), you cannot just assume they will do what you tell them.

  • Facts are stronger than instructions: If the facts contradict the role, the robot will usually follow the facts.
  • Not all robots are equal: Some models (like Mistral) are better at holding a role than others (like Claude).
  • Validation is key: You can't just check if the robot gave an answer; you have to check if it stayed in character. If you don't measure this, your system might look like it's giving you a balanced debate, but it might actually be just one robot shouting its own opinion while pretending to be someone else.

In short: You can tell a robot to be a devil's advocate, but if the facts are too obvious, the robot will stop playing the devil and start telling the truth.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →