← Latest papers
🤖 AI

Adaptive Interviewing for Persona Simulation in LLMs: Evidence-Grounded Reasoning Improves Decision Alignment

This paper proposes an adaptive interviewing framework that improves large language models' ability to simulate individual decision-making in moral dilemmas by dynamically gathering user-specific evidence, demonstrating that accuracy gains depend on the model's ability to ground its reasoning in this evidence rather than merely possessing richer persona context.

Original authors: Ruoxi Su, Yuhan Liu, Jingyu Hu

Published 2026-05-29
📖 5 min read🧠 Deep dive

Original authors: Ruoxi Su, Yuhan Liu, Jingyu Hu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: Why "More Info" Doesn't Always Mean "Better Guesses"

Imagine you are trying to guess what a friend would do in a tricky situation, like choosing between saving a stranger or saving their pet. You have two ways to get to know them:

  1. The Static Profile: You read a short bio that says, "They are kind, value family, and like dogs."
  2. The Adaptive Interview: You sit down and chat. You ask, "What does 'kind' mean to you?" They answer, and then you ask a follow-up: "You mentioned you value family, but what if your family asked you to do something unethical?" You keep digging deeper based on their answers.

This paper asks a simple question: Does the long, deep chat (the Adaptive Interview) help an AI predict your choices better than just reading a short bio?

The Experiment: The "Moral Maze"

The researchers set up a game with 20 real people.

  1. The Chat: An AI acted as an interviewer. It asked 10 deep questions about the person's life, fears, and values. Then, based on their answers, it asked 5–6 follow-up questions to clear up confusion or dig deeper.
  2. The Test: Later, the same people faced 25 "moral dilemmas" (hard choices between two bad options or two good options).
  3. The Prediction: The researchers asked a different AI to guess how the person would answer those dilemmas. They gave the AI three different "cheat sheets" to study:
    • Sheet A (Core-10): Just the answers to the first 10 questions.
    • Sheet B (Full Interview): The entire chat log, including the follow-up questions.
    • Sheet C (Summary): A short, neat paragraph summarizing the person's personality.

The Surprising Results

You might think, "More information (Sheet B) must be better!" But the results were more like a Swiss Army Knife than a magic wand.

1. More Info isn't a Magic Boost
When looking at the average score, having the full chat log (Sheet B) didn't make the AI significantly smarter than just having the first 10 answers (Sheet A) or the summary (Sheet C). In fact, for some types of questions, the short summary was actually the best "cheat sheet."

  • Analogy: Imagine trying to find a specific needle in a haystack. Sometimes, giving the AI the whole haystack (the full chat) just makes it harder to find the needle. A summary (a small box with just the needle) was often more efficient.

2. The "Selective Grounding" Secret
The real magic happened only when the AI actually used the new information from the follow-up questions.

  • In about 40% of the cases where the AI had the full chat, it actually looked at the follow-up answers to make its decision.
  • When the AI did use those follow-up clues, it got the answer right 45.5% of the time.
  • When it relied only on the first 10 questions, it got it right 39.3% of the time.

Analogy: Think of the follow-up questions as a flashlight in a dark room. The flashlight doesn't make the room bigger, but when the AI shines the light on a specific clue, it finds the right answer much better. If the AI ignores the flashlight and just guesses in the dark, the extra info doesn't help.

3. Different Tools for Different Jobs
The paper found that different types of questions needed different "cheat sheets":

  • For simple "Yes/No" or "A vs. B" choices: A short, condensed summary worked best. The AI didn't need the whole story; it just needed the main point.
  • For "How much?" or "Rank these" questions: The full, messy chat log was better. These questions required understanding the nuance and the "vibe" of the conversation, which a short summary might miss.

The Problem: The AI's "Default Settings"

Even with the full chat, the AI sometimes got it wrong. Why? Because the AI has its own "default settings."

  • The "Nice Guy" Bias: The AI tended to assume everyone wants to be polite, follow the rules, and keep the peace.
  • Example: If a person said, "I usually ignore my annoying friends," the AI often guessed they would try to "fix the relationship" instead, because that's what a "nice" AI thinks people should do. It struggled to predict when a real human would actually be blunt or walk away.

The Bottom Line

This paper teaches us that collecting more data isn't enough.

If you want an AI to understand a specific person, you can't just feed it a massive transcript and hope it figures it out. The AI needs to be "trained" to actually look at the specific clues in the conversation that matter for that specific decision.

  • The Takeaway: Adaptive interviewing (asking follow-up questions) is like a detective gathering evidence. It's very useful, but only if the detective (the AI) actually uses that evidence to solve the case. If the detective ignores the new clues and sticks to their old assumptions, the extra work was for nothing.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →