Validating LLMs in social science: Epistemic threats and emerging norms
This paper systematically analyzes validation practices in social science research using large language models (LLMs) as measurement instruments, revealing that while LLM-generated data is central to empirical analyses, current validation methods are inconsistent and limited, prompting the authors to propose emerging norms and strategies for more robust validation.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine social scientists are like detectives trying to solve mysteries about human behavior. For years, they've had to hire teams of human helpers to read thousands of documents, sort them into categories, or guess how people would answer survey questions. It's slow, expensive, and tiring.
Enter the Large Language Model (LLM): a super-smart, digital robot assistant that can read and write faster than anyone else. Researchers are now asking these robots to do the heavy lifting, turning messy text into neat numbers and charts. It sounds like a magic wand, right? But a new study by Meera Desai, Dallas Card, and Abigail Z. Jacobs suggests that while the wand is shiny, the magic might be a bit trickier than we thought.
The Big Discovery: The "Black Box" Problem
The authors looked at 27 papers published in the top social science journals between 2023 and 2025. They found 50 specific tasks where researchers used these AI robots to measure things like "offensiveness," "political ideology," or "sentiment."
Here is the twist: In almost all these cases, the robot's measurements were the main event. They were the core evidence used to make big claims about how society works. Yet, when the authors peeked behind the curtain to see how the researchers checked if the robots were actually telling the truth, they found a mess.
Think of it like baking a cake. You can use a fancy, high-tech oven (the LLM) to bake a cake. But if you don't check the temperature, don't taste the batter, and don't know exactly what recipe you used, you can't be sure the cake is actually a cake and not a brick. The study found that researchers are often baking with these high-tech ovens but skipping the taste tests.
The Three Big Mistakes
The paper highlights three main ways researchers are getting lost in the sauce:
The "Vague Recipe" (Conceptualization):
Imagine you tell a robot, "Find me a 'bad' post." But you never tell the robot what "bad" means. Is it rude? Is it sad? Is it wrong?
The study found that in 13 papers (covering 29 tasks), researchers didn't define their concepts at all. They just used a single word or a short phrase. It's like telling a chef, "Make me a 'delicious' meal" without saying if you want spicy, sweet, or savory. Without a clear definition, the robot might be measuring something totally different than what the researcher intended.The "Random Settings" (Operationalization):
Using an LLM is like driving a car with a million knobs and dials. You have to choose the model, the "temperature" (how creative or random the answers are), and how to grab the answer from the robot's long, chatty response.
The authors found that researchers were turning these knobs in wildly different ways, often based on a hunch or a guess. Some used zero-shot prompts (just instructions, no examples), while others used few-shot prompts (instructions plus examples). Some asked the robot to answer with a number; others asked for a sentence.
The scary part? The paper suggests that tiny changes in these settings can completely change the result. It's like turning the radio dial slightly and suddenly hearing a different station. Yet, many researchers didn't explain why they chose their specific settings, making it hard for others to repeat the experiment.The "One-Note" Check (Validation):
This is the biggest issue. When you use a robot to measure something, you need to prove it's accurate. The gold standard is comparing the robot's work to a human's work (the "gold standard").
The study found that 38 out of 50 tasks relied almost entirely on one type of check: convergent validity. This means they just asked, "Does the robot agree with the human?"
But the paper argues this isn't enough. It's like checking if a new thermometer works by comparing it to an old one that might also be broken. The authors suggest researchers should be using a whole toolbox of checks:- Face validity: Does the answer look reasonable at a glance?
- Hypothesis validity: Does the data help answer a real, interesting question?
- Predictive validity: Does the data predict future events correctly?
- Discriminant validity: Is the robot measuring only what we asked it to, or is it picking up on other stuff too?
Shockingly, 8 tasks in the study had no validation at all. They just took the robot's word for it.
What the Paper Rules Out
The authors are very clear about what this study is not saying. They are not saying LLMs are useless. They are not saying we should stop using them.
However, they explicitly rule out the idea that we can just "plug and play" these models. You cannot treat an LLM like a standard ruler or a thermometer that gives the same result every time you use it. The paper argues against the idea that these models are "universal instruments" that elicit meaningful measurements automatically. Instead, they are custom-built tools that require careful design, precise definitions, and rigorous testing.
They also rule out the idea that the current way of doing things is "good enough." The fact that researchers are skipping steps like defining concepts or checking for bias isn't just a minor oversight; it's a threat to the whole field. If the measurements are shaky, the conclusions about society are shaky too.
How Sure Are We?
The authors are very confident in their findings because they didn't just guess; they did the work. They collected a comprehensive list of 2,143 papers from top journals and manually coded 27 papers with 50 specific tasks. They didn't simulate this; they looked at real, published research.
They found that while LLMs are becoming central to social science, the "norms" (the rules of the road) for using them safely haven't caught up yet. The paper suggests that without better rules—like always defining your terms, reporting every setting you used, and checking your work with multiple methods—we risk building a house of cards on a shaky foundation.
The Takeaway
The paper ends with a call to action. It's not a "stop" sign; it's a "proceed with caution" sign. Social scientists need to start treating these AI tools with the same respect and rigor they give to human surveyors. They need to write down exactly what they asked the robot, how they asked it, and how they checked the answer.
Until then, the robot might be a powerful assistant, but it's still a bit of a mystery. And in science, mysteries are fun, but they don't make for solid facts.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.