PICon: A Multi-Turn Interrogation Framework for Evaluating Persona Agent Consistency
This paper introduces PICon, a multi-turn interrogation framework that evaluates the internal, external, and retest consistency of LLM-based persona agents, revealing that even advanced systems fail to match human baselines when subjected to logically chained questioning.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are hiring a new employee, but instead of a real person, you are interviewing a digital twin—a computer program pretending to be a human. You need to know: Is this program actually "being" that person, or is it just making things up as it goes along?
This paper introduces PICON, a new way to test these digital twins. Think of PICON not as a friendly chat, but as a detective's interrogation room.
Here is the breakdown of how it works, using simple analogies:
1. The Core Idea: The "Lie Detector" Interview
The authors realized that current ways of testing AI characters are too easy. It's like asking a suspect, "Do you like pizza?" and "What's your favorite color?" If they answer correctly, we assume they are telling the truth. But a liar can easily answer simple, unrelated questions without getting caught.
PICON changes the game by using interrogation logic (inspired by old police manuals). The idea is simple: If you make up a story, the more details you add, the more likely you are to trip over your own lies.
2. The Three Pillars of Truth
PICON checks the AI's "honesty" in three specific ways:
Internal Consistency (The "Memory Test"):
- The Analogy: Imagine the AI is a character in a movie. In Scene 1, it says, "I was born in 1990." In Scene 10, it says, "I was born in 1985."
- The Test: PICON asks a chain of connected questions. If the AI says it went to "Berkeley," the next question asks about the specific professor there. If the AI later claims it never went to Berkeley, the system catches the contradiction. It's like a detective saying, "Wait, you said you were at the bank at 2 PM, but you also said you were at the movies at 2 PM. Which one is it?"
External Consistency (The "Google Check"):
- The Analogy: The AI claims, "I work at a company called 'Flubber Corp' in 'Topeka'."
- The Test: PICON has a built-in "Google Search" tool. It instantly looks up "Flubber Corp" and "Topeka." If the company doesn't exist, or if the city is in a different country, the AI fails. It's like a background check that happens in real-time.
Retest Consistency (The "Amnesia Test"):
- The Analogy: You ask the AI, "What is your dog's name?" It says "Buddy." You ask 20 questions later, "What is your dog's name?" It says "Spot."
- The Test: PICON asks the exact same basic questions at the beginning and the end of the interview. If the AI's "memory" changes, it fails. Real humans might forget a small detail, but they don't usually change their entire life story in one sitting.
3. The Experiment: Humans vs. Robots
The researchers didn't just test robots; they tested 63 real humans using the same tough interrogation. They wanted to see how "human" the humans are compared to the robots.
The Shocking Results:
- The Humans: Even real people made small mistakes or got a little confused when asked the same question twice. But overall, they were consistent.
- The Robots: Even the most advanced AI "personas" (like Character.ai or specialized research models) failed miserably compared to real humans.
- Some AIs were great at not contradicting themselves but were terrible at providing real facts (they made up fake companies).
- Some were great at facts but would suddenly forget their own name or age halfway through the chat.
- The Big Takeaway: No current AI is good enough to perfectly replace a real human in a complex, multi-hour conversation. They are like actors who forget their lines if the play goes on too long.
4. Why This Matters
We are starting to use AI to simulate humans for things like:
- Training doctors (simulating patients).
- Testing new products (simulating customers).
- Running social science experiments.
If the "simulated patient" forgets they have a drug allergy halfway through the simulation, the doctor trainee might make a dangerous mistake. If the "simulated customer" changes their mind about what they want, the product design will be wrong.
PICON is the "stress test" we need. It tells us: "This AI is okay for a 5-minute chat, but if you try to use it for a 1-hour deep conversation, it will fall apart and start lying."
Summary
PICON is a rigorous, multi-turn interrogation framework that treats AI personas like suspects in a detective story. By chaining questions together, checking facts against the real world, and asking the same questions twice, it exposes the cracks in the AI's "fabricated identity." The study shows that while AI is getting better, it still cannot perfectly mimic the consistent, reliable memory of a real human being.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.