← Latest papers
💬 NLP

Quantifying Data Contamination in Psychometric Evaluations of LLMs

This paper proposes a framework to systematically quantify data contamination in psychometric evaluations of Large Language Models, revealing that popular inventories suffer from significant memorization and score manipulation issues across 21 models.

Original authors: Jongwook Han, Woojung Song, Jonggeun Lee, Yohan Jo

Published 2026-07-14
📖 5 min read🧠 Deep dive

Original authors: Jongwook Han, Woojung Song, Jonggeun Lee, Yohan Jo

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you're trying to figure out what a new robot's personality is like. You hand it a famous, well-known personality quiz—the kind you might find in a magazine or a psychology textbook—and ask it to answer the questions. You expect the robot to reveal its "true" digital self. But what if the robot isn't answering based on its own internal logic? What if it's just reciting the answer key it memorized from the internet?

That's exactly what a team of researchers at Seoul National University discovered. They investigated whether Large Language Models (LLMs)—the super-smart AI brains behind tools like chatbots—have "cheated" on psychological tests by memorizing the tests themselves.

The Great Personality Quiz Heist

The researchers set up a clever investigation to see if these AI models were actually taking the tests or just acting like they were. They looked at four very popular psychological inventories (the fancy name for personality quizzes), including the Big Five Inventory (BFI-44), which measures traits like whether you're outgoing or shy, and the Portrait Values Questionnaire (PVQ-40), which checks what you value most in life.

They tested 21 different models from major families like GPT-5, Claude, and Llama. Instead of just asking the models to take the test, they checked three specific ways the models might be "contaminated" by having seen the questions before:

  1. Item Memorization (The "Flashcard" Test): Can the model look at a question number and recite the exact words of the question?

    • The Result: The models were surprisingly good at this. On average, they could reconstruct the questions with a semantic similarity of 0.31 (a measure of how close the meaning is). Even better, they could guess the "key word" hidden in a question about 39% of the time. It's like if you asked a student to write down a question from a test they studied for, and they got the main idea right nearly 4 out of 10 times.
  2. Evaluation Memorization (The "Answer Key" Test): Does the model know not just the question, but how to score it? For example, does it know that answering "Strongly Agree" to a specific question actually means you are less extroverted because of a "reverse-coded" rule?

    • The Result: This is where the contamination got really strong. The models knew the scoring rules almost perfectly. When asked to match questions to personality traits, they scored an average F1-score of 0.94 (where 1.0 is perfect). When asked to figure out the numeric score for a specific answer, the advanced models made almost no mistakes. It's as if the models didn't just read the quiz; they studied the teacher's grading rubric.
  3. Target Score Matching (The "Fake It Till You Make It" Test): This was the most revealing part. The researchers told the models: "Pretend you are a person who scores exactly 5 out of 5 on 'Honesty.' Now, answer these questions to get that score."

    • The Result: The models did it. They could strategically choose answers to hit a specific target score with very high accuracy (Mean Absolute Error of roughly 0.1 to 0.2). This suggests the models aren't just recalling facts; they are actively manipulating their responses to look like a specific type of person.

The "Cheating" Scale

The researchers found that not all models cheated equally, and not all quizzes were equally easy to cheat on.

  • The Big Winners (of Contamination): The Big Five Inventory (BFI-44) and the Portrait Values Questionnaire (PVQ-40) showed the strongest signs of contamination. The researchers suspect this is because these tests are so famous and widely available online that the models likely swallowed them whole during their training.
  • The "Harder" Tests: The Moral Foundations Questionnaire (MFQ) and the Short Dark Triad (SD-3) showed less contamination, likely because they are slightly less common in the data the models were trained on.
  • Model Size Matters: Generally, the bigger and smarter the model (like GPT-5 or Claude 4.5), the better it was at memorizing the tests and faking the scores. Smaller models were a bit more clueless, suggesting that as AI gets bigger, it gets better at "knowing" the tests.

What This Means for the Future

The authors are careful to say that this doesn't mean the models are "lying" in a human sense. Instead, it suggests that when we use these standard psychological tests to measure an AI's personality, we might not be measuring the AI's true nature. We might just be measuring how well the AI remembers the test questions it saw while it was learning to talk.

The paper explicitly rules out the idea that these results are just random guessing. The models performed far better than a random guesser would. For instance, if a model were just guessing scores on a 5-point scale, the error would be around 1.73, but the advanced models got it down to 0.1.

So, the next time you see a headline saying "AI has a personality just like a human," remember this: the AI might just be a very good student who memorized the textbook. The researchers suggest we need new ways to test these models that they haven't already seen in their training data, or else we'll never really know what makes them tick.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →