← Latest papers
💬 NLP

Eliciting Trustworthiness Priors of Large Language Models via Economic Games

This paper introduces a novel method using iterated in-context learning and the Trust Game to elicit trustworthiness priors from large language models, revealing that GPT-4.1's trust behaviors closely mirror human patterns and can be predicted by stereotypes of warmth and competence.

Original authors: Siyu Yan, Lusha Zhu, Jian-Qiao Zhu

Published 2026-02-03
📖 5 min read🧠 Deep dive

Original authors: Siyu Yan, Lusha Zhu, Jian-Qiao Zhu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: How "Trustworthy" is an AI?

Imagine you are hiring a new assistant to handle your money. You give them $10, and you tell them, "I'm going to multiply this by 3, so you now have $30. Please give me back some of it."

  • If they give you back $15: They are trustworthy.
  • If they give you back $0: They are not.

This is the basic setup of a famous game called the Trust Game. Usually, psychologists use this game to see how humans behave. But this paper asks a new question: How do Large Language Models (LLMs) like GPT-4 or Llama behave in this game? Do they have a "default setting" for how much they trust others and how much they return the favor?

The Problem: You Can't Just Ask an AI

You might think, "Why not just ask the AI, 'Are you trustworthy?'" The authors say that's a bad idea.

  • The "Polite Robot" Problem: If you ask an AI, "Will you steal my money?" it will almost certainly say, "No, I am a helpful assistant."
  • The Reality: Just like humans, AI can be persuasive and fluent but still make risky or unreliable choices in practice. We need to see what they do, not just what they say.

The Solution: The "Telephone Game" of Trust

To find out an AI's true "trust settings," the researchers invented a clever method based on a concept called Iterated In-Context Learning.

Think of this like a game of Telephone, but instead of passing a whisper, they are passing a story about money.

  1. The Setup: The researchers create a chain of "generations" of AI.
  2. The First Step: They show the AI a few examples of people playing the Trust Game (e.g., "Person A sent $5, Person B returned $2").
  3. The AI's Turn: The AI looks at those examples and predicts: "If I were Person B, how much would I return?" Let's say it predicts 40%.
  4. The Loop: The researchers take that 40% prediction and use it to generate new fake examples of people playing the game. They feed these new examples to the next AI in the chain.
  5. The Result: They repeat this process 30 times.

Why do this?
Imagine you are trying to guess the "average personality" of a group by watching them play a game over and over. At first, the results might look random. But if you keep passing the results down the line, the "noise" fades away, and what's left is the AI's deep-down instinct (or "prior") about how trust works. It's like tuning a radio until the static clears and you hear the true station.

What They Found

1. Some AIs are "Human-Like," Others are Not
The researchers tested 20 different AI models. They compared the "trust settings" they found in the AI to the average trust settings found in thousands of real humans.

  • The Winner: GPT-4.1 was the closest match to human behavior. Its "trust dial" was set almost exactly where humans set theirs.
  • The Losers: Some other models were either too stingy (returning very little money) or had weird, split personalities (sometimes very trusting, sometimes very suspicious).

2. The "Risk" Connection
They noticed something interesting: The AI models that were rated as "riskier" (more likely to take chances or say bold things) were actually the ones that behaved most like humans in the trust game. The models that were super strict about following rules or avoiding risk didn't act as much like real people in this social game.

3. AI Judges People by "Warmth" and "Competence"
In the second part of the study, they used the most human-like AI (GPT-4.1) to play the game against different "characters" (personas).

  • They asked the AI to play against a "Doctor," an "AI," "Elon Musk," "Malala Yousafzai," etc.
  • The Finding: The AI didn't treat everyone the same. It returned more money to people it perceived as Warm (kind, caring) and Competent (smart, skilled).
  • The Magic Formula: The researchers built a simple math formula based on two things: Warmth and Competence.
    • If a character was just "Competent" (smart) but not "Warm" (kind), the AI didn't trust them much.
    • If a character was "Warm" but not "Competent," the AI trusted them a bit more.
    • The Sweet Spot: If a character was both Warm and Competent (like a beloved, smart mentor), the AI's trust skyrocketed. This matches how humans judge people, too.

The Takeaway

This paper shows that we can use a "money game" to peek inside an AI's brain and see its hidden rules about trust.

  • Some AIs are already acting very much like humans when it comes to trust.
  • These AIs judge others based on the same two things we do: Are they nice? Are they capable?
  • This helps us understand that AI isn't just a calculator; it has social "biases" and instincts that we can measure and study.

What the paper does NOT say:
The authors do not claim this method can fix AI safety, cure mental health issues, or be used to hire employees. They simply say: "Here is a way to measure how an AI thinks about trust, and here is what we found."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →