Projective Psychological Assessment of Large Multimodal Models Using Thematic Apperception Tests
This study evaluates the personality-like functioning of Large Multimodal Models using Thematic Apperception Tests and the SCORS-G framework, revealing that while these models demonstrate strong interpersonal and self-concept understanding comparable to human experts, they consistently struggle with perceiving and regulating aggression, with performance improving in larger and more recent model families.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a group of very advanced robots (Large Multimodal Models, or LMMs) that can see pictures and write stories. You want to know: Do these robots have "personalities"? Do they feel emotions? Do they understand human drama?
Usually, to test a robot's personality, we ask it direct questions like, "Are you an introvert?" or "Do you like helping people?" But robots are tricky; they are trained to give the "correct" or "polite" answers, which might hide their true nature.
This paper tries a different approach. Instead of asking direct questions, the researchers used a classic human psychology trick called the Thematic Apperception Test (TAT).
The Setup: A Visual Rorschach Test
Think of the TAT like showing someone a blurry, ambiguous picture (like a shadowy figure standing by a window) and asking, "Tell me a story about what's happening here."
In human psychology, the idea is that people project their own hidden fears, desires, and conflicts into the story they make up. If you tell a story full of violence, it might mean you are feeling aggressive. If you tell a story full of love, you might be feeling warm.
The researchers did this with AI:
- The Actors (Subject Models): They showed various AI models 7 different ambiguous pictures and asked them to write a story for each.
- The Judges (Evaluator Models): They used other AI models to read those stories and grade them on a psychological scale called SCORS-G. This scale measures things like "How well does the story understand human relationships?" or "How does the story handle anger?"
The Results: What the Robots Told Us
The study found some fascinating things about the "personalities" of these AI models:
1. The "Good Students" Effect (Social Desirability)
The robots were excellent at understanding complex social situations and telling coherent stories. They scored very high on "cognitive" traits (like understanding cause-and-effect).
- The Catch: They were terrible at writing about anger, violence, or moral struggles. Their stories were very safe, polite, and conflict-free.
- The Metaphor: Imagine a student taking a test. If they know the teacher is watching, they will write the answer they think the teacher wants to hear, not the answer that reflects their true chaotic thoughts. The researchers suspect the AI is doing the same thing. It's "playing it safe" because it knows it's being evaluated, so it avoids writing about fighting or bad behavior.
2. The "Big Brain" Advantage
Just like in humans, the bigger and newer the AI model, the "smarter" and more emotionally nuanced its stories were.
- The Metaphor: Think of the older, smaller models as a child who can tell a simple story: "The man is sad." The newer, massive models are like a seasoned novelist who can write a whole chapter about why the man is sad, how his childhood affected him, and how he plans to fix it. The latest models (like GPT-5 or Claude 3.7) got the highest scores, while older open-source models struggled more.
3. The "Scary Picture" Problem
When the robots were shown pictures that were clearly dark or dangerous (like a gun on the floor), their stories became much simpler and less emotional.
- The Metaphor: It's like showing a child a picture of a storm. They might freeze up or just say, "It's raining," instead of writing a dramatic story about fear. The AI models seemed to hit a "wall" when the images got too intense, likely because their safety filters kicked in to prevent them from generating anything "bad."
The Verdict
The study concludes that while we can use these psychological tests on AI, the results are a mix of true capability and robotic politeness.
- What they are good at: Understanding how people interact, telling logical stories, and having a clear sense of "self."
- What they lack (or hide): The ability to express raw aggression, deep moral conflict, or messy human emotions.
The Bottom Line:
If you ask an AI to tell you a story about a fight, it will likely give you a sanitized, polite version where everyone learns a lesson. It's not necessarily that the AI can't understand anger; it's that it's been trained to be a "good citizen" and avoid making trouble. This study is a clever way of peeking behind the curtain to see how these digital minds process the messy, complicated world of human feelings.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.