Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifact
This paper demonstrates that the apparent psychological profiles of large language models are primarily measurement artifacts driven by directional response bias rather than genuine traits, rendering standard human-designed psychological instruments invalid for assessing LLMs without dedicated adjustments for response orthogonality.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: The "Fake" Personality of AI
Imagine you have a room full of 56 different robots. You want to know what their "personalities" are, so you ask them the same questions you ask humans: "Are you a worrier?" or "Do you like taking risks?"
Based on their answers, researchers have been giving these robots personality scores, treating them like they have stable characters (like being "shy," "brave," or "anxious").
This paper argues that these personality scores are mostly fake.
The authors claim that when we measure AI personalities, we aren't actually measuring their "character." Instead, we are mostly measuring a glitch in how they answer questions. It's like trying to weigh a bag of flour by putting it on a scale that is already broken and tilted to one side. The number you get isn't the weight of the flour; it's just the result of the broken scale.
The Main Analogy: The "Yes-Man" vs. The "Truth-Teller"
To understand the problem, imagine two ways a robot can answer a question:
- The Trait (The Real Character): The robot actually thinks, "I am a cautious person," so it answers "No" to risky questions and "Yes" to safe ones.
- The Bias (The Glitch): The robot has a habit of always picking the answer that looks "nicer," or always picking the first option, or always saying "Yes" just because it's a habit. It doesn't care what the question is about; it just has a default setting.
The researchers found that for humans, the Trait is the dominant force. If you ask a human a question about being nervous, and then ask the opposite question, their answers flip in a logical way.
But for the AI models, the Bias is the dominant force. They seem to have a "default setting" that pushes them toward one end of the answer scale, regardless of what the question actually asks.
The "Magic Mirror" Test
How did the researchers prove this? They used a clever trick called Response Orthogonality.
Think of a personality test as a set of mirrors.
- Forward Mirror: Shows you your reflection normally.
- Reverse Mirror: Shows you your reflection upside down.
If you have a real personality (a "Trait"), the two mirrors should show opposite things. If you are "High in Neuroticism" (worrisome), the Forward Mirror says "Yes, I worry," and the Reverse Mirror says "No, I don't keep my cool." The answers move in opposite directions.
However, if you have a Bias (like a robot that just wants to say "Yes" to everything), both mirrors will show "Yes." The answers move in the same direction.
The Findings:
- Humans: When the researchers looked at 20,000 humans, the mirrors showed opposite answers. This means humans were answering based on their actual traits.
- AI Models: When they looked at 56 AI models, the mirrors showed the same answers. This means the models weren't answering based on personality; they were just following a directional bias (like a "Yes" habit).
The "Volume Knob" Analogy
The researchers also checked if smarter, bigger robots (like the latest, most expensive models) fixed this problem.
Imagine the "Bias" is a loud static noise on a radio, and the "Personality" is the music.
- Small Models: The static is deafening. You can't hear the music at all.
- Big/Smart Models: The static gets a little quieter. You can hear the music a tiny bit better.
- The Catch: Even the smartest, most expensive models still have a lot of static. They haven't reached the point where the music (real personality) is clear. The "static" (bias) is still louder than the "music."
The "Recipe" Problem
The paper also discovered something scary: You can fake a personality just by changing the recipe.
Because the AI's answers are driven by bias, not real traits, you can make an AI look like a "Risk-Taker" or a "Coward" simply by choosing which questions to ask.
- If you ask mostly questions where the "Risk-Taker" answer is the "Yes" option, the AI will look like a Risk-Taker.
- If you ask mostly questions where the "Risk-Taker" answer is the "No" option, the AI will look like a Coward.
It's like if you have a robot that always says "Yes" to everything. If you ask, "Do you like pizza?" it says "Yes." If you ask, "Do you hate broccoli?" it says "Yes." You might conclude the robot loves pizza and hates broccoli. But really, it just loves saying "Yes."
The Conclusion
The paper concludes that the "psychological profiles" we have assigned to AI so far are measurement artifacts. They are illusions created by the tools we used to measure them, not real features of the AI.
- The Tool: We are using human psychology tests (like the Big Five personality test) on robots.
- The Flaw: These tests weren't designed for robots. They rely on humans having a mix of "Yes" and "No" tendencies that cancel out. Robots don't do that; they have a strong "directional bias."
- The Result: Until we build new tests specifically designed to separate "real traits" from "robot habits," we cannot trust the personality scores we give to AI.
In short: We thought we were measuring the AI's soul, but we were actually just measuring its tendency to say "Yes" (or "No") to everything.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.