HEART-Bench: Do LLM Agents Exhibit Human-like Psychology?
This paper introduces HEART-Bench, a novel benchmark designed to evaluate whether LLM agents can simulate coherent, human-like psychology by testing their ability to make consistent behavioral decisions based on Big Five personality traits and extensive autobiographical memories across 64 diverse scenarios.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you've built a robot that can write poetry, solve math problems, and plan a vacation itinerary. It's incredibly smart at tasks. But if you ask it, "How would you feel if your best friend betrayed you?" or "What would you do if you saw someone being treated unfairly?", does it answer like a human with a heart, or like a calculator spitting out the most logical response?
The paper HEART-BENCH asks this exact question: Do AI agents actually have a "personality" that feels real, or are they just pretending?
Here is a simple breakdown of what the researchers did, using some everyday analogies.
1. The Problem: The "Empty Suit" AI
Right now, most AI tests check if a robot can remember facts or hold a conversation. It's like testing a mannequin to see if it can stand up straight. But a real human isn't just a list of facts; we are shaped by our past, our moods, and our core values.
The authors argue that current AI is like an actor who has memorized a script but hasn't lived the life of the character. They can say the right words, but they don't feel the way the character should feel.
2. The Solution: Building "Digital Humans" with Backstories
To test this, the researchers didn't just give the AI a name and a job. They built 11 distinct "Digital Humans" (like creating 11 different actors for a play).
- The Blueprint (Big Five): Each character was built on a strict psychological framework called the "Big Five" (Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism). Think of this as their DNA. One character might be extremely shy and organized; another might be chaotic and very outgoing.
- The Life Story (1,000 Memories): This is the magic ingredient. Instead of just a bio, each character was given 1,000 detailed memories from childhood to age 50.
- Analogy: Imagine you are writing a biography for a character. You don't just say "He was sad." You write a 2,000-character story about the specific day he dropped his ice cream, the smell of the rain, the sound of his mother's voice, and how that moment made him afraid of failure for years.
- The AI was fed these 1,000 "life moments" for each of the 11 characters.
3. The Test: The "What Would You Do?" Game
Once the characters were built, the researchers created 64 different life scenarios (like a choose-your-own-adventure book). These scenarios covered everything from a workplace argument to a family dinner or a moral dilemma.
- The Setup: They took a specific character (e.g., the shy, organized one) and dropped them into a specific situation (e.g., "Your neighbor is complaining about noise, but you know they are lying").
- The Question: The AI had to choose the best response from four options.
- Option A: Yell at them.
- Option B: Quietly file a formal report alone.
- Option C: Try to make a joke to diffuse tension.
- Option D: Ignore them completely.
The Catch: All four options might look like "doing something," but only one fits the specific personality and life history of that character. The test wasn't about being "right"; it was about being consistent.
4. The Results: The "Uncanny Valley" of Personality
The researchers tested 12 of the world's smartest AI models (including GPT, Claude, Gemini, and others) on this game.
- The Score: Even the best AI (Gemini-3.1-Pro) only got about 63% correct. Most others scored below 40%.
- What this means: The AI is still struggling to "get" the character. It often picks the response that sounds polite or logical, but fails to pick the response that a shy, traumatized, or overly anxious person would actually choose based on their 1,000 memories.
- The Memory Issue: Interestingly, giving the AI more complex memory systems didn't help much. The bottleneck wasn't that the AI forgot the memories; it was that the AI couldn't reason with them to make a human-like decision. It's like having a library of books but not knowing how to write a story using them.
5. The Takeaway
HEART-BENCH is a new ruler for measuring "emotional intelligence" in AI.
- Before: We asked, "Can the AI remember what I said 10 minutes ago?"
- Now: We ask, "Does the AI remember who it is, and does it act like that person when things get tough?"
The paper concludes that while AI is getting better at talking, it is still very far from having a consistent, human-like soul. It's like a very good mimic that hasn't yet learned how to truly be someone else.
In short: We built 11 digital people with full life histories, threw them into 64 tough situations, and found that even the smartest AIs are still mostly guessing how those people would react, rather than truly understanding them.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.