Inferring personality from a live interaction with an LLM: Predictions are internally consistent and show convergent validity
This preregistered study demonstrates that large language models can reliably infer Big Five personality traits from live conversations with high internal consistency and convergent validity, particularly for neuroticism and agreeableness, though performance varies by trait observability and prompting strategy.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
For decades, psychologists have relied on a simple tool to understand who we are: a questionnaire. People sit down and rate themselves on a series of statements, answering whether they agree or disagree with ideas about being outgoing, organized, or anxious. This method is efficient and has given us a standard map of human personality, often called the Big Five. These five traits—openness to new experiences, conscientiousness, extraversion, agreeableness, and neuroticism—act as a framework for describing how people think, feel, and behave. Yet, this approach has a flaw. It asks people to look inward and judge themselves, a process that can be clouded by how we wish to be seen or by simple mistakes in self-reflection. It also misses the rich, messy details of how we actually speak and interact in the real world.
Recently, a new kind of technology has emerged that might change how we measure these traits. Large language models are computer programs trained on vast amounts of text, capable of holding conversations that feel surprisingly human. Because these models have read so much human writing, they can sometimes spot patterns in how people use words that reveal their underlying character. The question researchers began to ask was whether a computer could listen to a real-time conversation and accurately guess a person's personality, just by hearing them talk, without them ever filling out a form.
A team of researchers at the University of Bonn set out to test this idea. They invited 117 people to have a ten-minute conversation with an artificial intelligence chatbot. The chatbot was given a specific instruction: to ask open-ended questions about the participant's past week, keeping its own answers brief, all while trying to gather enough information to understand the person's character. The goal was to see if the computer could extract a personality profile from this brief exchange. After the conversation ended, the researchers asked the chatbot to evaluate the participant's personality on the five main traits. To ensure the results were solid, the researchers also asked the participants to fill out a standard personality questionnaire about themselves, providing a benchmark to compare against the computer's guesses.
The researchers tested two different ways of asking the computer to make its judgment. In the first approach, they simply asked the model to rate the person on the five traits based on the conversation. In the second, more complex approach, they fed the computer the exact list of questions from the standard personality test and asked it to guess how the person would have answered each one. The results showed that the computer was surprisingly good at the task. When the researchers checked the consistency of the computer's ratings, they found that the model understood the structure of personality just as well as a human psychologist would. The scores the computer generated were internally consistent, meaning it didn't give contradictory answers for the same person.
More importantly, the computer's assessments matched the participants' own self-reports. The strongest matches were found for neuroticism, which relates to emotional stability and anxiety, and agreeableness, which relates to kindness and cooperation. The computer was particularly good at spotting signs of depression and compassion in the way people spoke. However, the technology struggled with other traits. It had a harder time predicting how outgoing or open-minded a person was. This difficulty likely stems from the nature of the conversation itself. Traits like extraversion often show up in how people act in groups or how energetic they are in person, things that are hard to capture in a text chat with a machine. Similarly, creativity and openness are often internal processes that do not always leave a clear trace in a short conversation.
The study also looked at whether the computer was biased. There was a concern that the model might rely on stereotypes, such as assuming women are more emotional than men, simply because it had read those ideas in its training data. However, because the researchers hid the participants' gender from the chatbot, the computer did not amplify these stereotypes. The differences it found between men and women were consistent with what the participants said about themselves, suggesting the model was reacting to the actual people rather than a preconceived notion. The researchers did find, however, that the computer tended to be overly kind in its judgments. It gave higher scores for traits like agreeableness and conscientiousness and lower scores for neuroticism, perhaps because it was programmed to be polite or because it defaulted to a "nice" version of a person.
While the computer's predictions were not perfect, the study suggests that we are moving toward a future where personality can be understood through natural conversation rather than just checklists. The researchers found that the simpler method of asking for a general rating worked better than the complex method of feeding the computer the test questions, which caused the system to fail more often. This indicates that the model has an intuitive grasp of personality that does not always need a rigid framework to work. The findings suggest that while these tools are not yet ready to replace human psychologists or make life-altering decisions, they can successfully pull meaningful signals about who we are from the words we choose to say. As the technology improves, and as conversations become longer and more detailed, the ability to read personality from speech may become a powerful new way to understand the human mind.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.