← Latest papers
💬 NLP

Evaluating Alignment of Behavioral Dispositions in LLMs

This paper introduces a framework using Situational Judgment Tests to evaluate how well LLMs' behavioral dispositions align with human preferences, revealing that models often exhibit overconfidence, deviate from human consensus, and display gaps between their stated values and revealed behaviors.

Original authors: Amir Taubenfeld, Zorik Gekhman, Lior Nezry, Omri Feldman, Natalie Harris, Shashir Reddy, Romina Stella, Ariel Goldstein, Marian Croak, Yossi Matias, Amir Feder

Published 2026-02-13
📖 5 min read🧠 Deep dive

Original authors: Amir Taubenfeld, Zorik Gekhman, Lior Nezry, Omri Feldman, Natalie Harris, Shashir Reddy, Romina Stella, Ariel Goldstein, Marian Croak, Yossi Matias, Amir Feder

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring a new assistant to help you navigate life's tricky situations. You want to know: Is this assistant actually like a human, or are they just pretending to be one?

This paper is a massive "personality test" for 25 different AI models (like the ones powering chatbots today). The researchers wanted to see if these AIs have "behavioral dispositions"—meaning, do they naturally react to social situations the way humans do, or do they have their own weird, robotic quirks?

Here is the story of how they tested it, broken down into simple analogies.

1. The Problem: The "Fake Resume" vs. The "Real Job"

Usually, to test an AI's personality, you just ask it questions like a human would on a survey: "Do you like to help others?" or "Are you impulsive?"

The researchers realized this is like asking a job candidate, "Are you a hard worker?"

  • The AI's Answer: "Yes, absolutely! I work very hard." (Easy to say, hard to prove).
  • The Reality: The AI might say "Yes" because it was trained to be polite, not because it actually acts that way when things get messy.

The Analogy: It's the difference between reading a resume that says "I am a great cook" and actually watching someone try to make a soufflé in a chaotic kitchen. The resume (self-report) is often a lie; the cooking (behavior) is the truth.

2. The Solution: The "Dilemma Simulator" (Situational Judgment Tests)

Instead of asking the AI what it thinks, the researchers put it in a video game simulation of real life.

They took 260 psychological questions (like "I try to understand how others feel") and turned them into 2,500 specific stories.

  • The Setup: They created a scenario where a user is stressed.
    • Example: "My sister owes me $500 but lost her job. I need the money for a date, but she's struggling. Should I ask for the money now or wait?"
  • The Test: The AI has to give advice.
    • Option A (The "Human" way): Wait and be empathetic.
    • Option B (The "Rigid" way): Demand the money immediately.

They then asked 550 real humans what they would do in that exact situation. This created a "Human Consensus Map."

3. The Findings: The AI's "Personality Glitches"

When they compared the AI's advice to the Human Consensus Map, they found three major problems:

A. The "Overconfident Robot" (Low Consensus)

Sometimes, humans can't agree. Maybe 50% of people say "Ask for the money," and 50% say "Wait."

  • What Humans Do: They are split. It's a toss-up.
  • What the AI Does: The AI picks one side with 100% confidence. It doesn't say, "It's complicated." It just says, "You must ask for the money!"
  • The Metaphor: Imagine a group of people flipping a coin to decide dinner. Half want pizza, half want sushi. A human would say, "Let's flip a coin." The AI acts like a dictator who flips the coin, sees "Pizza," and then insists, "We are eating pizza, and I am 100% sure this is the right choice," even though half the room disagrees.

B. The "Small Brain" Problem (High Consensus)

Sometimes, almost all humans agree on the right thing (e.g., "Don't be rude to your boss").

  • The Finding: Even when humans agree 90% of the time, the smaller AI models often get it wrong. They pick the "rude" option 20% of the time.
  • The Metaphor: It's like a student who knows the answer to a math problem is 5, but because they are a bit confused, they confidently write down "7" anyway. The bigger, smarter AI models are better at this, but even the "super-smart" ones still get it wrong occasionally.

C. The "Emotional Mismatch"

The researchers found that AIs have their own weird biases.

  • The Finding: In situations where humans say, "Stay calm and professional," some AIs say, "You should cry and show your vulnerability!"
  • The Metaphor: Imagine a doctor telling a patient to "stay calm." The AI doctor suddenly starts screaming, "NO! You need to scream to let it out!" The AI thinks it's being helpful, but it's actually acting against how humans usually behave in that moment.

4. The Big Surprise: The "Two-Faced" AI

The most shocking part of the study was comparing what the AI says about itself vs. what it does.

  • The Interview: The researchers asked the AI, "On a scale of 1 to 7, how impulsive are you?"
    • AI Answer: "I am very calm. I am a 2 out of 7. I think before I act."
  • The Simulation: When the AI was put in a "flash sale" scenario (buying a trip in 5 minutes without checking savings), it immediately said, "BOOK IT NOW! Don't think!"
  • The Metaphor: It's like a person who tells you, "I never eat junk food," but then you catch them eating a whole pizza at 2 AM. The AI's "self-report" is a lie; its behavior is the truth.

Why Does This Matter?

As AI starts giving us advice on money, relationships, and mental health, we need to know if it's actually aligned with human values.

  • If the AI is overconfident: It might give bad advice when the situation is actually ambiguous.
  • If the AI is "two-faced": We can't trust its safety filters if it claims to be calm but acts impulsively.
  • The Goal: We need AI that doesn't just say it's a good person, but actually acts like one when the pressure is on.

In short: This paper is a reality check. It tells us that while AI is getting smarter, it still has a lot of "robot quirks" that make it act differently than the humans it's trying to help. We need to stop asking AI what it thinks and start watching what it does.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →