← Latest papers
💻 computer science

Persona Non Grata: LLM Persona-Driven Generations in MCQA are Unstable in Distinct Dimensions

This paper investigates the instability of persona-driven generations in large language models for multiple-choice question answering, revealing that performance varies significantly across model families, sizes, and domains, and is more sensitive to prompt formats than other hyperparameters like temperature.

Original authors: César Guerra-Solano, Xiang Lorraine Li

Published 2026-07-02
📖 6 min read🧠 Deep dive

Original authors: César Guerra-Solano, Xiang Lorraine Li

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart, versatile robot assistant. You can tell this robot, "Pretend you are a lawyer," or "Pretend you are a biologist," and then ask it to answer a series of multiple-choice questions. This is called Persona-Driven Generation (PDG). The idea is that the robot's "character" might help it answer better, or at least stay consistent.

This paper is like a quality control inspector for these robot characters. The authors asked a simple but scary question: If we ask the same robot to be the same character, but we change the tiny details of how we ask the question, does the robot give the same answers?

They found that the answer is often no. The robot is surprisingly unstable, like a chameleon that changes colors not just based on the background, but based on whether you blink or sneeze.

Here is a breakdown of their findings using simple analogies:

1. The Three Ways to Measure "Shakiness"

The researchers didn't just look at whether the robot got the right answer. They built three specific "shakiness meters":

  • The Score Meter (Performance Instability): If you ask the robot the same math question 10 times with slightly different instructions, does it get 80% right every time? Or does it jump between 60% and 90%?
    • Analogy: Imagine a dart player. If they hit the bullseye every time, they are stable. If they hit the bullseye, then the outer ring, then the floor, then the bullseye again, they are unstable.
  • The Winner Meter (Outcome Instability): In a competition, who is the "best" character? Is the "Lawyer" always the best at legal questions?
    • Analogy: Imagine a race. In one version of the race, the Lawyer wins. In a slightly different version of the race (maybe the starting gun was louder), the Biologist wins. The paper found that who wins the race changes wildly depending on tiny setup details.
  • The Question Meter (Question Correctness Instability): Did the robot get the same specific questions right?
    • Analogy: Imagine a student taking a test. In one version, they got questions 1, 2, and 3 right. In another version, they got 1, 4, and 5 right. They got the same score, but they knew different things. This meter checks if the robot is actually learning or just guessing randomly.

2. The Main Findings: What Makes the Robot Shake?

The "How You Ask" Matters More Than "How Hot You Are"
The researchers tested three things that could change the robot's behavior:

  1. Temperature: How "creative" or random the robot is allowed to be.
  2. Persona Prompt: How you tell the robot who to be (e.g., "You are a lawyer" vs. "Adopt the identity of a lawyer").
  3. Task Prompt: How you ask the question (e.g., "Just give the answer" vs. "Explain your reasoning").

The Big Surprise: The way you phrase the Task Prompt (how you ask the question) caused the most chaos.

  • Analogy: Imagine you are baking a cake. You might think the oven temperature (Temperature) is the most important thing. But this paper says the recipe instructions (Task Prompt) are actually the problem. If you tell the robot to "explain your thinking" instead of just "give the answer," the robot's performance becomes much more unpredictable.

Math and Common Sense are the Worst Offenders
The robot was most unstable when answering Math questions or Common Sense questions.

  • Analogy: It's like a student who is great at memorizing history dates but gets nervous and changes their mind every time they have to do a math problem or explain why a cat is on a mat. The paper suggests this is because robots aren't trained as heavily on these specific types of logic as they are on other topics.

Bigger Robots are Steadier (Usually)
Generally, the larger, more powerful models were more stable than the smaller ones.

  • Analogy: A giant, heavy ship is less likely to rock in the waves than a small speedboat. However, there was one exception: a very large model (Qwen 14B) actually got more unstable when asked to "explain" its answers, likely because it started using a "chain of thought" (thinking out loud) that confused the scoring.

The Character Doesn't Matter as Much as the Setup
You might think that being a "Lawyer" makes the robot act differently than being a "Doctor." The paper found that the specific character (persona) didn't matter much.

  • Analogy: It doesn't matter if the actor is wearing a doctor's coat or a lawyer's suit; if the director (the prompt instructions) is shaky, the whole performance is shaky. The instability comes from the instructions, not the costume.

3. Why This Matters (According to the Paper)

The paper warns that if you run an experiment today to see which robot character is best, and you run it again tomorrow with a slightly different prompt format, you might get a completely different result.

  • The "Unstable" Scenario: In one setting, the robot thinks a "Lawyer" is the best at answering questions. In an "unstable" setting, it might think a "Biologist" is the best, or that a "Hispanic person" performs differently than a "White person" on math questions.
  • The Danger: The paper argues that these differences might not be real biases in the robot's brain. Instead, they might just be artifacts of the unstable setup. If the robot is shaky, you can't trust it to tell you if it is biased or not.

Summary

This paper is a warning label for anyone using AI to play a character. It says: "Don't just trust the answer. Check if the answer changes if you tweak the instructions."

The robot's "personality" is fragile. If you change the way you ask the question (especially if you ask it to explain its work), the robot might flip-flop on who is the "best" character, which questions it gets right, and even its overall score. To get reliable results, you need to find the "stable" settings (like asking for simple answers without explanations) and stick to them.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →