← Latest papers
💻 computer science

A Survey of Large Language Models for Perception and Measurement of Human Psychology

This paper presents a systematic review of Large Language Models as instruments for human psychological measurement, proposing a three-dimensional framework to analyze their theoretical plausibility, methodological approaches, and practical effectiveness in assessing constructs like personality and mental health.

Original authors: Yudong Li, Xiaoyi Chen, Jiawei Cai, Zehao Zhong, Haoyang Yang, Huajin Tang, Linlin Shen

Published 2026-06-23
📖 5 min read🧠 Deep dive

Original authors: Yudong Li, Xiaoyi Chen, Jiawei Cai, Zehao Zhong, Haoyang Yang, Huajin Tang, Linlin Shen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a super-smart, incredibly well-read robot librarian. This librarian has read almost every book, article, and tweet ever written. Now, imagine you want to use this librarian to understand the human mind—specifically, to figure out someone's personality, their mood, or if they are struggling with mental health issues, just by reading what they write.

This paper is a big review of that idea. It asks: Can this robot librarian actually do a good job of measuring human psychology?

Here is a breakdown of what the paper says, using simple analogies:

1. The Big Question: Is the Librarian "Real"?

Before we let the robot measure us, we have to ask: Does the robot actually understand us, or is it just a master of mimicry?

  • The Debate: Some people think the robot is just a "statistical parrot." It doesn't have feelings or a soul; it just predicts the next word based on patterns it saw in its training data.
  • The Evidence: However, the paper shows that when you give the robot a "Theory of Mind" test (like asking it to guess what a character in a story is thinking), it often gets it right, performing as well as a 6-to-9-year-old child or even an adult.
  • The Verdict: The robot has developed "functional" understanding. It acts like it understands human thoughts, even if we don't know how it does it. This makes it a promising tool, but we shouldn't treat it as a human yet.

2. How Do We Use the Robot? (The Three Tools)

The paper organizes the ways researchers are using these models into three main "tools":

  • Tool A: The Active Interviewer (The Chatbot)
    • How it works: The robot asks you questions, just like a therapist or a job interviewer. It can ask follow-up questions if your answer is vague.
    • The Analogy: Think of it as a very fast, tireless interviewer who can talk to thousands of people at once. It uses "prompts" (instructions) to decide what to ask next.
  • Tool B: The Passive Observer (The Social Media Stalker)
    • How it works: The robot reads things you've already written—like your diary, tweets, or Facebook posts—without you talking to it directly.
    • The Analogy: Imagine a detective reading your old letters to figure out your personality. It looks for hidden patterns in your word choices to guess if you are happy, sad, or stressed.
  • Tool C: The Multi-Sensory Detective (The Fusion)
    • How it works: The robot doesn't just read text; it looks at pictures, listens to voice recordings, and checks heart rate data all at once.
    • The Analogy: This is like a doctor who listens to your voice, looks at your face, and reads your medical chart simultaneously to get the full picture, rather than just reading a form.

3. What Has the Robot Measured So Far?

The paper reviews where this technology has been tested:

  • Personality: The robot is getting pretty good at guessing the "Big Five" personality traits (like whether you are outgoing or shy) from your writing. In some tests, its guesses match what humans say about themselves quite well.
  • Mental Health: It can spot signs of depression, anxiety, or even suicidal thoughts in text. It's becoming a useful "early warning system" that can flag people who might need help.
  • The "Virtual Subject": Interestingly, researchers are also using the robot to pretend to be people. They tell the robot, "Act like a 30-year-old woman who is stressed," and the robot generates responses. This helps scientists run experiments without needing real human volunteers, saving time and money.

4. The Catch: Why We Can't Trust It Yet

The paper is very careful to point out that while the robot is impressive, it is not ready to replace human doctors or psychologists. Here are the big problems:

  • The "Fickle Friend" Problem (Reliability): If you ask the robot the same question twice, it might give you two different answers. It is sensitive to how you phrase your question. A tiny change in wording can change the result completely. This is bad for medical tests, which need to be consistent.
  • The "Hallucination" Problem: Sometimes the robot makes things up. It might sound confident and logical, but it could be inventing facts or misinterpreting a crisis. In a mental health setting, a wrong guess could be dangerous.
  • The "Bias" Problem: The robot learned mostly from English-speaking, Western internet data. If you use it on someone from a different culture or background, it might misunderstand them or judge them unfairly because it doesn't "get" their cultural context.
  • The "Black Box" Problem: When a human doctor makes a diagnosis, they can explain their reasoning. When the robot makes a guess, it's often hard to know why it decided that. Doctors need to know the "why" to trust the result.

5. The Bottom Line

The paper concludes that Large Language Models are like powerful, high-tech assistants, not replacements for human experts.

  • What they are good at: They are fast, cheap (relatively), and can handle huge amounts of data. They are great for screening large groups of people or helping researchers run experiments.
  • What they are bad at: They are not stable enough, not transparent enough, and not culturally sensitive enough to make life-or-death medical decisions on their own.

The Final Analogy:
Think of the robot as a very talented apprentice doctor. It has read all the textbooks and can spot symptoms quickly. But it lacks the experience, the consistency, and the ethical judgment of a senior doctor. For now, the robot should be used to help the senior doctor do their job better, not to take over the job entirely.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →