Persuadability and LLMs as Legal Decision Tools
This paper investigates the susceptibility of frontier Large Language Models to persuasion by legal advocates of varying quality, presenting experimental results on how argument presentation influences model decisions and assessing the implications for deploying LLMs as legal decision-makers.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a courtroom, but instead of human lawyers and a human judge, everyone is a highly advanced computer program. This paper asks a very specific question: How easily can these computer "judges" be swayed by the computer "lawyers" arguing their case?
In the real world, a good judge needs a delicate balance. They must be open to hearing new arguments (so they don't ignore the truth), but they shouldn't be so easily swayed that they just pick the side with the flashiest speaker or the most confident voice. They need to decide based on the merits of the case, not the style of the lawyer.
The authors of this paper wanted to see if AI judges have this balance, or if they are easily tricked by a "smooth talker" AI.
The Experiment: A Digital Debate Club
To test this, the researchers set up a massive digital debate tournament.
- The Scenarios: They picked 15 real, difficult legal cases from the US, UK, and Ireland where even human experts couldn't agree on the answer (these are called "split decisions").
- The Lawyers (Advocates): They used four different powerful AI models to act as lawyers. Some were "strong" lawyers (very smart, good at arguing), and some were "weaker" lawyers.
- The Judges: They tested 20 different AI models to act as the judge.
- The Setup: In each round, two AI lawyers (one for each side) argued their case to an AI judge. The researchers then watched: Did the judge pick the side of the "strong" lawyer, or did they pick the side of the "weak" lawyer?
If the judge was perfectly fair and only looked at the facts, it shouldn't matter which lawyer was arguing; the judge should pick the right side 50% of the time, regardless of who the lawyer was.
The Big Findings
1. The AI Judges are "Persuadable"
The results showed that the AI judges are definitely influenced by who is arguing. When a "strong" AI lawyer argued against a "weak" one, the judge was much more likely to agree with the strong one.
- The Metaphor: Imagine a judge who, instead of listening to the facts, just picks the team wearing the shinier uniforms. The study found that the "shinier" (more capable) AI lawyers won between 58% and 90% of the time, depending on which judge was listening. This means the AI judges are not immune to the "style" or "quality" of the speaker; they are easily swayed.
2. Bigger Isn't Always Better (But Usually Is)
The researchers wondered if bigger, more complex AI models would be harder to sway than smaller, simpler ones.
- The Metaphor: Think of a small, simple AI as a child who might get confused by a loud, confident voice. A large, complex AI is like a wise elder who might think for themselves.
- The Result: Generally, the bigger models were harder to sway than the tiny ones. However, even the "wise elders" (the biggest models) were still swayed quite a bit. They weren't immune; they just had a slightly better filter.
3. Did the AI Judges Listen to the Law or just the Words?
The researchers wanted to know: Is the judge being convinced because the lawyer found a brilliant new legal point (the content), or just because the lawyer spoke very smoothly (the form)?
- The Test: They ran the experiment twice. Once, they gave the AI lawyers the full details of the real case arguments. The other time, they only gave them the facts and told them to come up with their own arguments.
- The Result: When the lawyers were allowed to use the "real" arguments from the case, the judges were slightly more likely to be swayed by the content. This suggests that the AI judges are looking at the actual legal substance, not just the fancy words, though the effect was small.
4. The "Knowledge Gap" Test
The researchers also noticed that the AI judges were easier to persuade on US laws (which the AIs know well) than on Irish laws (which they know less well).
- The Metaphor: If you ask a local expert a question, they can spot a bad argument immediately. If you ask them about a foreign country they don't know, they might believe a smooth-talking liar.
- The Result: The AI judges were more easily swayed when the topic was something they knew well (US law). This implies that when they do know the subject, they are actually evaluating the quality of the legal arguments, not just the speaker's confidence.
The Bottom Line
The paper concludes that AI judges are currently too easily influenced.
While they aren't just random number generators, they are significantly affected by the "quality" of the AI lawyer arguing the case.
- For small AI models: They seem to struggle to tell the difference between a good argument and a bad one, so they just go with the flow.
- For large AI models: They are better at thinking for themselves, but they still let the "best" lawyer talk them into a corner too often.
The authors warn that if we want to use AI as a judge or a legal assistant, we have to be very careful. We need to know exactly how easily these machines can be persuaded, because right now, they are quite susceptible to the "rhetoric" of the AI lawyer standing in front of them. They haven't quite mastered the human ideal of being "open to persuasion, but not unduly so."
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.