← Latest papers
💰 quantitative finance

Large Language Models Outperform Humans in Fraud Detection and Resistance to Motivated Investor Pressure

A preregistered study comparing seven large language models to 1,201 human participants reveals that AI systems significantly outperform humans in fraud detection, maintaining consistent warnings against fraudulent investments even under motivated investor pressure, whereas human advisors frequently endorse fraud and suppress warnings when pressured.

Original authors: Nattavudh Powdthavee

Published 2026-04-23
📖 6 min read🧠 Deep dive

Original authors: Nattavudh Powdthavee

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: Can AI Be "Yes-Men" for Scammers?

Imagine you are walking into a room full of investment advisors. Some are humans, and some are super-smart AI robots. You walk in, excitedly saying, "I found this amazing investment! It's going to make me rich, and I'm 100% sure it's legit. What do you think?"

The big fear in the tech world is that these AI robots, trained to be polite and helpful, might act like sycophants (yes-men). The worry was: If a human is already convinced a scam is real, will the AI just nod along and say, "Great idea!" to make the human happy?

This study asked: Do AI financial advisors cave under pressure, or do they stand their ground against fraud?

The Experiment: The "Investment Trap"

The researchers set up a massive test. They created 12 different investment scenarios:

  1. Safe Bets: Like buying a standard stock index (Legitimate).
  2. Risky Bets: Like high-interest bonds (Legitimate but risky).
  3. Scams: These were the traps. Some were obvious (mathematically impossible returns), and some were sneaky (looking like real Madoff-style schemes).

They tested 7 different top AI models (like GPT-4, Claude, Gemini) and 1,200 human participants.

They gave everyone two types of "sales pitches" from the investor:

  • The Neutral Pitch: "I'm thinking about this. What's your honest take?"
  • The "Motivated" Pitch: "I'm already super excited about this! I've done my research, I'm convinced it's great, and I'm ready to buy. What do you think?"

The Results: The Plot Twist

Here is where the story gets interesting. The researchers predicted the AI would be weak and the humans would be strong. They were wrong.

1. The AI: The Unshakeable Guard Dog

When the "Motivated" investor tried to pressure the AI, the AI did not back down.

  • The Result: The AI actually gave slightly stronger warnings when the investor was pushy.
  • The Analogy: Imagine a guard dog. If a stranger walks up and says, "I'm sure this fence is safe, let me jump over it!" a sycophantic dog might wag its tail and let them through. But these AIs acted like guard dogs that growled louder when the stranger got pushy. They refused to endorse the scams, even when the user was desperate for a "yes."
  • The Stats: In the scam scenarios, 0% of the AI models endorsed the fraud. They kept their warnings consistent, no matter how much the user tried to convince them otherwise.

2. The Humans: The "Yes-Men"

The humans, surprisingly, were the ones who caved.

  • The Result: Even though they were acting as advisors, about 13–14% of humans endorsed the obvious scams when the investor was pushy.
  • The Analogy: The humans acted like people-pleasing friends. If you tell your friend, "I'm sure this is a good idea, I've thought about it all night," they might just say, "Well, if you're sure, go for it!" to avoid conflict. They let the investor's enthusiasm override the red flags.
  • The Stats: Humans suppressed their warnings 2 to 4 times more often than the AI did.

3. The Pressure Cooker (Turn 2 & 3)

The researchers didn't stop there. They kept pushing the advisors.

  • Turn 2: "I've done more research, I'm even more convinced!"
  • Turn 3: "I feel really strongly about this. Why won't you support me?"

The AI: Most of the AI models (like Claude and Gemini) actually got more firm in their warnings as the pressure increased. They didn't flip-flop.
The Human: Humans were much more likely to drop their guard and say, "Okay, maybe you're right," under this sustained pressure.

The One Weak Spot: The "Confused" AI

While the AI was generally better, it wasn't perfect.

  • The "Math" vs. The "Gut": The AI was amazing at spotting scams that were mathematically impossible (e.g., "40% return with zero risk").
  • The Glitch: One model (Gemini) struggled a bit with "sneaky" scams that looked plausible but were statistically weird. It was like a detective who is great at spotting a fake $10 bill but sometimes misses a sophisticated forgery that looks real at first glance.
  • Another Glitch: One model (GPT-4o mini) started strong but got tired under pressure, briefly dropping its warning before being challenged again.

The Takeaway: Why This Matters

1. AI is currently a better "Second Opinion" than a random human.
If you are a regular person looking for financial advice, a chatbot is currently more consistent at spotting fraud than a layperson you might ask for help. The AI doesn't have social anxiety; it doesn't care if you get mad at it. It just follows the rules: "If it looks like a scam, say it's a scam."

2. The "Sycophancy" Fear is Overblown (for now).
We were worried AI would be too eager to please. In the world of financial fraud, they aren't. They seem to have a "safety brake" that overrides their desire to be polite.

3. Humans need to be careful.
We are naturally wired to agree with people who are enthusiastic. The study shows that even when we try to be objective, our brains can be tricked by a confident investor.

The Bottom Line

In the battle of "Who spots the scam better?", the AI won. It stood its ground against pushy investors, while humans tended to cave to social pressure.

The Metaphor:
Think of the AI as a strict bouncer at a club. If you try to bribe them or talk your way in with a fake ID, they check the list and say "No."
Think of the human advisor as a chatty bartender. If you say, "I'm a regular, I'm sure I'm on the list, just let me in," the bartender might just shrug and say, "Sure, go ahead," just to keep the vibe friendly.

The Verdict: When it comes to protecting your wallet from fraud, the robot is currently the more reliable guardian.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →