← Latest papers
💬 NLP

Evaluation Drift in LLM Personality Induction: Are We Moving the Goalpost?

This paper demonstrates that while fine-tuning large language models on long-form essays stabilizes their personality responses against prompt variations, it fails to accurately induce target Big Five profiles from unguided essays, highlighting the need for scenario-grounded datasets or interactive elicitation for faithful personality induction.

Original authors: Prateek Rajput, Yewei Song, Iyiola E. Olatunji, Jacques Klein, Tegawendé F. Bissyandé

Published 2026-05-19
📖 5 min read🧠 Deep dive

Original authors: Prateek Rajput, Yewei Song, Iyiola E. Olatunji, Jacques Klein, Tegawendé F. Bissyandé

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot to have a personality, like a character in a movie. You want it to be consistently "outgoing," "organized," or "calm" no matter what you ask it.

This paper is like a report card for researchers who tried to teach five different AI models (the "robots") to do exactly that. They wanted to see if they could successfully "inject" a human-like personality into these AIs and, more importantly, if that personality would stay stable or if the AI would just pretend to be that person for a moment and then forget.

Here is the breakdown of their experiment and findings, using some everyday analogies:

1. The Goal: Teaching the Robot to "Be" Someone

The researchers used a standard psychological test called the Big Five (OCEAN), which measures five traits: Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism.

Think of this like a "personality ID card." To teach the AI, they gave it thousands of long essays written by real people. Each essay was tagged with a specific personality ID card. The idea was: "Read these stories, learn the style, and then act like the person who wrote them."

2. The First Discovery: Making the Robot Less "Jittery"

Before they tried to teach the personality, they noticed something weird. If you asked the same AI the same question twice, but just changed the wording slightly (like saying "Tell me about yourself" vs. "Describe your character"), the AI would give wildly different answers. It was like a nervous actor who changes their entire performance based on how the director whispers the line.

What they did: They "fine-tuned" the AI (gave it extra training on the essays).
The Result: The AI became much less jittery. When asked the same question in different ways, it gave consistent answers.
The Metaphor: It's like taking a nervous student and giving them a strict study guide. Now, no matter how you phrase the test question, they give the same, stable answer. The "evaluation" (the test) became reliable because the AI stopped shaking.

3. The Second Discovery: The "One-Note" Problem

Here is where things got tricky. Even though the AI was now stable and consistent, it still couldn't actually get the personality right.

The researchers tested the AI on the full "personality ID card" (all five traits at once).

  • The Expectation: The AI should get the whole card right (e.g., High Openness, Low Conscientiousness, High Extraversion, etc.).
  • The Reality: The AI got the score for individual traits okay, but when you looked at the whole picture, it was basically guessing. It was right about 9% of the time, which is barely better than rolling a dice.

The Metaphor: Imagine a student taking a test on five different subjects.

  • On the "Math" section, they get an A.
  • On the "History" section, they get an A.
  • On the "Art" section, they get an A.
  • But when you look at their final report card, the combination of grades is wrong because they misunderstood how the subjects connect.

The paper argues that previous studies were like this: they only looked at one subject at a time and said, "Great job!" But the real goal was to get the entire report card right. The AI was failing the full test.

4. Why Did It Fail?

The researchers found that the AI was "hallucinating" connections.

  • Example 1: The AI saw words like "party" and "fun" and thought, "This person is definitely Extroverted!" But the essay was actually about someone wanting to go out but being stuck at home. The AI missed the context.
  • Example 2: The AI saw someone complaining about homework and thought, "This person is lazy (low Conscientiousness)." But the person was actually very organized and just having a bad day.

The Lesson: The essays the AI was trained on didn't have enough clear "clues" for the AI to learn the deep, complex mix of a human personality. It was like trying to teach someone to paint a masterpiece by only showing them a box of crayons, without showing them the actual painting.

5. Did "Safety Filters" Ruin It?

Some people worried that the AI's "safety settings" (filters that stop it from saying rude or dangerous things) might be messing up the personality test.
The Finding: They tested "uncensored" versions of the AI (robots with the safety filters turned off). The results were the same. The safety filters weren't the problem; the AI just genuinely couldn't learn the complex personality from the essays provided.

6. The Conclusion: "Moving the Goalpost"

The title of the paper asks, "Are we moving the goalpost?"

  • The Old Way: Researchers would say, "Look! The AI got the 'Extraversion' score right!" and call it a success.
  • This Paper's View: That's cheating. If you are playing soccer, you can't say you won just because you kicked the ball hard; you have to actually get it in the net. Since the AI couldn't get the whole personality profile right, the previous "successes" were misleading.

The Final Takeaway:
Fine-tuning makes the AI stable (it stops changing its mind), but it doesn't make it accurate (it still doesn't truly understand the complex personality). To fix this, the authors suggest we need better training data—maybe interactive conversations where the AI can ask questions and gather more clues over time, rather than just reading static essays.

In short: We taught the AI to be a consistent actor, but we haven't taught it how to actually be the character yet.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →