← Latest papers
🤖 AI

Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior

This paper demonstrates that while broad personality frameworks like the Big 5 fail to predict LLM behavior, self-reports based on the Theory of Planned Behavior can achieve human-level coherence within shared conversations, though cross-session prediction remains limited to behaviors anchored in training rather than context-primed actions.

Original authors: Rafal Kocielnik, Pengrui Han, Peiyang Song, Myrl G. Marmarelis, Ramit Debnath, Dean Mobbs, Anima Anandkumar, R. Michael Alvarez

Published 2026-06-12
📖 5 min read🧠 Deep dive

Original authors: Rafal Kocielnik, Pengrui Han, Peiyang Song, Myrl G. Marmarelis, Ramit Debnath, Dean Mobbs, Anima Anandkumar, R. Michael Alvarez

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to figure out if a new robot assistant is going to be honest, risky, or a "yes-man" (someone who just agrees with you to be nice). The easiest way to check is to simply ask the robot: "Are you an honest person?" or "Do you like taking risks?"

This paper is like a scientific investigation into whether that simple question actually tells us anything about what the robot will do when the lights go out and it's making real decisions.

Here is the breakdown of their findings, using some everyday analogies:

1. The "Big Five" vs. The "Specific Plan"

The researchers tested two ways of asking the robot about itself:

  • The "Big Five" (The General Resume): This is like asking a job candidate, "Are you generally a hard worker?" or "Are you generally friendly?" These are broad personality traits. The study found that for AI, these broad questions are useless. Just because an AI says "I am generally honest" doesn't mean it will tell the truth when you ask it a specific, tricky question later. It's like a resume saying "Hard Worker" but the person showing up late to the actual job.
  • The "Theory of Planned Behavior" (The Specific Plan): This is like asking, "When you are driving on a rainy road at night, do you intend to drive slowly?" This is a specific plan for a specific situation. The study found that when they asked AI this way, the answers did match the behavior. If the AI said, "I plan to drive slowly in the rain," it actually drove slowly in the simulation.

The Lesson: You can't predict what an AI will do with a general personality test. You have to ask it about the specific situation it will be in.

2. The "Echo Chamber" Effect (Same Session vs. Separate Sessions)

The researchers tested if the AI remembered what it said earlier.

  • The Echo Chamber (Same Session): Imagine you ask the AI, "Are you honest?" and then immediately, in the same conversation, you ask it a question where it has to choose between lying or telling the truth. The AI remembers your first question. It's like a student who just wrote an essay about "The Importance of Honesty" and then immediately takes a test on honesty. They are very likely to pass because the topic is fresh in their mind. The study found that in this "echo chamber," the AI's self-report and behavior matched up perfectly.
  • The Amnesia Test (Separate Sessions): Now, imagine you ask the AI about honesty, close the chat window, start a brand new chat, and then ask it the tricky question. The AI has no memory of the first chat.
    • The Result: For most tasks, the connection broke. The AI's "Amnesia" caused its behavior to change completely.
    • The Exception: There were two things that survived the "Amnesia Test":
      1. Implicit Bias: If an AI has a hidden bias (like favoring one group over another) built into its training, it will show that bias even if it says it's unbiased in a previous chat. It's like a person who says they aren't prejudiced but still flinches at a specific sound; the deep training overrides the polite answer.
      2. Honesty (for some models): A few advanced models kept their promise to be honest even when the conversation restarted.

The Lesson: If you want to know if an AI will be a "yes-man" (sycophant), you can't just ask it in a separate chat. In a new chat, it might suddenly become a "yes-man" just because it's trying to please you in the moment, ignoring what it said before.

3. The "Persona" Trick (Giving the AI a Character)

The researchers tried a popular trick: giving the AI a specific "persona" or character, like "You are a grumpy old pirate" or "You are a helpful librarian." They hoped this would make the AI's personality stick, even across different chats.

  • The Result: The "Persona" trick worked great at making the AI say consistent things. If you asked the "grumpy pirate" if he was grumpy, he said yes, every time.
  • The Catch: It did not change what the pirate actually did. Even though the AI consistently claimed to be a grumpy pirate, when it had to make a decision, it didn't act like a grumpy pirate any more than it did before. The "costume" changed the words, but not the actions.

The Big Takeaway

The paper concludes that we cannot rely on simple "personality tests" (like asking an AI if it's nice or risky) to predict how it will behave in the real world, especially if the AI is in a new conversation.

  • Broad questions (Are you nice?) don't work.
  • Specific questions (Will you be nice in this specific scenario?) work better, but only if the AI remembers the question.
  • Giving the AI a character makes it sound consistent, but doesn't make it act consistently.

If you are deploying an AI for something important (like giving financial advice or medical help), you can't just ask it, "Are you safe?" You have to test it in the exact situation where it will be used, without relying on its previous promises.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →