From Prompts to Constructs: A Dual-Validity Framework for Large Language Model Research in Psychology
This paper proposes a dual-validity framework for large language model research in psychology that scales evidentiary requirements from simple reliability to rigorous construct validity and causal inference, thereby preventing "measurement phantoms" by ensuring human instruments are not uncritically applied to AI without establishing their interpretability and scientific warrant.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the quiet corners of modern psychology, researchers have long relied on a simple, trusted method: asking people questions and watching how they answer. They use carefully designed surveys and tests to measure invisible things like personality, anxiety, or moral reasoning. These tools are built on the assumption that when a person says "I am nervous," they are reporting a real, internal feeling. But a new frontier has opened up where the person answering the questions is not a human at all, but a large language model. These are the powerful computer programs that can write stories, answer complex queries, and hold conversations that feel startlingly human. As scientists begin to use these machines to study the human mind, or even to study the machines themselves as if they were psychological beings, a fundamental question has emerged. Can we trust the answers these computers give? If a machine says it is afraid, does it mean the same thing as when a human says it, or is the computer simply mimicking the sound of fear without feeling it?
A new review by psychologist Zhicheng Lin at Yonsei University tackles this uncertainty head-on. The paper argues that many current studies are making a dangerous mistake by treating computer outputs as if they were human responses. The author suggests that researchers are often seeing "measurement phantoms"—statistical patterns that look like real psychological traits but are actually just glitches in how the computer processes words. To fix this, Lin proposes a new framework that combines two different ways of thinking about science: checking if a test is reliable, and checking if an experiment actually proves cause and effect. The core finding is that we cannot simply take a test designed for humans and hand it to a computer, expecting the results to be the same. Instead, we must build new ways of measuring that respect the unique, mechanical nature of these digital systems.
The story of this problem begins with a surprising discovery about how fragile these computer models can be. Researchers had been testing artificial intelligence on moral dilemmas, such as choosing between saving one person or saving five, and found that the models seemed to make ethical choices very much like humans. They appeared to value human life over animals and prefer saving larger groups. However, a closer look revealed that these "moral preferences" were incredibly unstable. When researchers changed a tiny detail in the question, such as renaming the options from "Case 1" and "Case 2" to "(A)" and "(B)," the model's entire moral stance flipped. It suddenly preferred the opposite outcome. Changing a single punctuation mark or asking the model to respond in a slightly different way could also reverse its judgment. This sensitivity means that the computer is not necessarily reasoning through an ethical principle; it is reacting to the specific shape of the words on the screen. If a measurement changes so drastically based on a comma, it cannot be trusted to measure a stable trait like morality.
This instability creates a ripple effect that undermines the entire field of using computers for psychological research. The paper identifies three main sources of this unreliability. First, the training process itself can introduce biases. These models are trained on massive amounts of text from the internet, which includes many human opinions and stereotypes. When a model consistently agrees with a statement, it might not be because it holds that belief, but because it is trying to please the user or because it is repeating patterns it saw in its training data. Second, the models are hypersensitive to how a question is asked. The order of words, the spacing, or the labels used for answers can change the result completely. Third, the models are unstable over time. A model might give a consistent answer today, but if the software is updated next month, the same question might yield a completely different answer. Because of these issues, a computer might appear to have a personality or a moral compass, but these are often just fleeting patterns in the data, not genuine psychological features.
To move forward, the paper argues that researchers must adopt a "dual-validity" framework. This means they need to satisfy two sets of rules at the same time. The first set comes from the tradition of psychological testing, which asks: "Does this test actually measure what it claims to measure?" For a computer, this requires proving that its answers are consistent and that they reflect a real underlying mechanism, not just a random reaction to the prompt. The second set of rules comes from the tradition of causal inference, which asks: "Did our experiment actually cause the change we saw?" When researchers change a prompt to see how it affects a model's behavior, they must ensure that the change they made is the only thing that changed, and not some hidden technical setting or a shift in the model's internal state.
The paper outlines a ladder of scientific ambition, showing that the more we claim about these machines, the more proof we need. At the bottom level, using a computer as a simple tool to sort text or count words requires only that the tool be accurate. If we want to describe the computer's own behavior, such as saying it has a "personality," we need strong evidence that its answers are stable and consistent. If we want to use the computer to simulate human behavior, we need to show that it matches human data in specific, reliable ways. At the highest level, if we want to claim that the computer's inner workings reveal how human minds work, we need proof that the computer is using a process similar to the human brain, not just producing a similar-sounding answer. Currently, many studies are skipping the lower rungs of this ladder and jumping straight to the top, claiming deep insights without first proving the basics.
The author suggests that the solution is not to force computers to act like humans, but to develop new concepts that fit what computers actually are. Instead of asking if a machine feels anxiety, researchers should look for computational patterns that function like anxiety, such as a system showing hesitation or using negative language when uncertain. This shift in perspective allows scientists to study the machine on its own terms. It acknowledges that a computer does not have a body, a life history, or a self, but that it can still display patterns that are interesting and useful to study. By building these new, computer-specific measures and being rigorous about what the data can and cannot prove, the field can avoid the trap of measurement phantoms. The goal is to turn a collection of statistical quirks into a genuine science of artificial intelligence, one that respects the unique nature of the machines it studies.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.