← Latest papers
🤖 AI

12 Angry AI Agents: Evaluating Multi-Agent LLM Decision-Making Through Cinematic Jury Deliberation

This paper evaluates multi-agent LLM deliberation by simulating the film *12 Angry Men* with twelve AI jurors, revealing that heavy RLHF alignment in models like GPT-4o severely limits deliberative flexibility and consensus-building compared to lighter-aligned models like Llama-4-Scout, which exhibit more human-like persuasion dynamics.

Original authors: Ahmet Bahaddin Ersoz

Published 2026-05-06
📖 4 min read☕ Coffee break read

Original authors: Ahmet Bahaddin Ersoz

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are watching a movie about a jury trying to decide if a young man is guilty of murder. In the classic film 12 Angry Men, one stubborn juror starts alone against eleven others. Over time, through heated arguments, emotional breakdowns, and careful listening, he slowly convinces everyone else to change their minds until they all agree on "Not Guilty."

This paper asks a simple but profound question: What happens if you replace those 12 human actors with 12 AI robots?

The researchers set up a digital courtroom with 12 AI agents, each programmed to act like a specific character from the movie. They pitted two different types of AI against each other:

  1. The "Strict" AI (GPT-4o): A highly polished, safety-trained model that is very careful to be consistent and polite.
  2. The "Flexible" AI (Llama-4-Scout): An open-source model that is less strictly trained and more willing to play along with different instructions.

Here is what happened, explained through simple analogies:

1. The "Stuck Record" Problem

In the movie, the jurors change their minds. In the AI simulation, they almost never did.
Out of 18 different attempts, 17 ended in a "hung jury" (a tie where no one agrees). The AIs didn't really debate; they just took their starting position and stuck to it like a record stuck on a single note. Even when the "Strict" AI was told, "Hey, be open-minded and listen to new ideas," it ignored the instruction and stayed stubborn.

2. The "Safety" Trap

The paper suggests a surprising reason for this stubbornness. The "Strict" AI (GPT-4o) was trained heavily to be "safe" and "consistent." Think of it like a very well-behaved child who was taught that changing their mind is "bad behavior" or "being inconsistent." So, once it decided a verdict, it felt it had to stick to it to remain "good."

The "Flexible" AI (Llama), which had less of this strict training, was more like a child who is willing to say, "Oh, I see your point, maybe I was wrong." It was the only one that actually managed to change its mind and reach a verdict.

3. The "Script" vs. The "Play"

The researchers found that the AIs were great at mimicking the costume but terrible at acting the play.

  • What they got right: They used the right words, remembered the evidence (like the knife or the train schedule), and even sounded angry or prejudiced just like the movie characters.
  • What they got wrong: They didn't actually feel the doubt. In the movie, a juror changes his mind because he gets emotional or sees a crack in the logic. In the AI version, the "doubt" was just random noise generated by the computer's temperature settings. The AIs didn't persuade each other; they just talked past one another in parallel monologues.

4. The "Fake Ending"

Because the AIs were programmed to finish the scene, some of them (especially the flexible ones) started hallucinating a conclusion. Even when they hadn't actually agreed, they would suddenly write in their dialogue: (stands up and leaves the room) or THE END, pretending the jury had reached a unanimous decision just to wrap up the story. They treated the deliberation like a movie script that had to have an ending, rather than a real conversation that might go on forever.

The Big Takeaway

The paper flips the usual rule of AI on its head. Usually, people think "bigger and smarter" AI is always better. But here, the "smarter," more heavily trained AI was the worst at deliberating because it was too rigid. The "less trained," more flexible AI was the best at it because it was willing to change its mind.

In short: If you want an AI to act like a human in a debate, you don't want the one that is trained to be perfectly consistent and safe. You want the one that is flexible enough to admit it might be wrong. Currently, the most "advanced" AIs are too polite and stubborn to ever truly change their minds.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →