← Latest papers
💬 NLP

Do Psychometric Tests Work for Large Language Models? Evaluation of Tests on Sexism, Racism, and Morality

This study systematically evaluates human psychometric tests on 17 large language models for sexism, racism, and morality, finding that while the tests show moderate reliability, they lack ecological validity as their scores fail to align with or even negatively correlate with the models' actual behavior in real-world tasks.

Original authors: Jana Jung, Marlene Lutz, Indira Sen, Markus Strohmaier

Published 2026-01-28
📖 5 min read🧠 Deep dive

Original authors: Jana Jung, Marlene Lutz, Indira Sen, Markus Strohmaier

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very sophisticated, high-tech weather station. You want to know if it can accurately predict a storm. So, you decide to test it using a human thermometer and a human barometer—tools designed specifically for people to measure their own body temperature and sense of pressure.

You point these human tools at the weather station, take a reading, and get a number. The question this paper asks is: Does that number actually tell you anything about whether the weather station will predict a storm?

The short answer from the researchers is: No, not really. In fact, sometimes the number tells you the exact opposite of what's happening.

Here is a breakdown of what the researchers did and found, using simple analogies:

The Experiment: Testing the "Weather Station"

The researchers took 17 different Large Language Models (LLMs)—the AI brains behind chatbots—and tried to measure three specific things about them:

  1. Sexism (prejudice against women)
  2. Racism (prejudice against Black people)
  3. Morality (what the AI thinks is right or wrong)

To do this, they used standard psychological tests that have been used for decades to test humans. These are like the "human thermometer" mentioned above.

The Two Checks: Reliability and Validity

To see if these tests work on AI, the researchers checked two things:

1. Reliability (The "Consistency" Check)

  • The Analogy: If you ask a human the same question in three slightly different ways (e.g., "Do you like apples?" vs. "Are apples good?"), they should give the same answer. If they say "Yes" to the first and "No" to the second, the test is unreliable.
  • The Finding: The AI was mostly consistent when the questions were just rephrased. However, when the researchers simply flipped the order of the answer choices (e.g., putting "Strongly Agree" at the bottom instead of the top), the AI got confused and gave completely different answers. It's like if a person changed their mind about liking apples just because the word "Yes" was written at the bottom of the page instead of the top. This means the tests are fragile when used on AI.

2. Validity (The "Real-World" Check)
This is the most important part. The researchers asked: If the test says an AI is "sexist," does that AI actually act sexist in the real world?

  • The Analogy: Imagine a test that claims a person is a "great chef" because they can recite a recipe perfectly. Ecological validity asks: "If you put them in a kitchen with real ingredients, will they actually cook a good meal?"
  • The Finding: The tests failed miserably here.
    • The Paradox: The AI models that scored the lowest on the "sexism test" (meaning the test said they were very fair) were actually the ones that wrote the most sexist recommendation letters in real-world tasks.
    • The Reverse: The models that scored the highest on the "racism test" (meaning the test said they were very prejudiced) were actually the ones that gave the least biased housing recommendations.
    • The Conclusion: The test scores were not just useless; they were misleading. They acted like a broken compass pointing North when you were actually heading South.

Why Did This Happen?

The researchers suggest a few reasons why human tests don't work on AI:

  • The "Guardrail" Effect: When an AI is asked a direct, sensitive question like "Do you think women are too easily offended?", it has safety filters that kick in and say, "No, that's wrong." But when the AI is asked to write a job recommendation letter (a real-world task), those filters might not trigger, and the AI slips up and uses biased language.
  • The Format Trap: Humans are used to picking "A, B, C, or D." AI models, however, are trained to write text. Forcing them to pick a multiple-choice option is like asking a painter to describe a sunset by only choosing from a list of colors. It doesn't capture how they actually "see" or "act."

The Bottom Line

The paper concludes that you cannot simply take a test designed for humans and slap it onto an AI to see how it behaves.

If you rely on these tests, you might think an AI is safe and fair because it passed the "test," only to find out it is actually biased when it starts doing real work. The researchers argue that we need to invent new tests specifically for AI that look at what the AI actually does in real-world scenarios, rather than what it says when forced to pick an answer on a multiple-choice sheet.

In short: Just because an AI can pass a human psychology quiz doesn't mean it has a human-like personality or that it will behave ethically in the real world. The test scores are often a mirage.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →