Can LLMs Simulate Personas with Reversed Performance? A Systematic Investigation for Counterfactual Instruction Following in Math Reasoning Context
This paper introduces a benchmark for "counterfactual instruction following" to demonstrate that state-of-the-art Large Language Models, including reasoning models, fundamentally struggle to simulate personas with reversed performance levels (such as low-proficiency students) in mathematical reasoning tasks, a limitation that is further exacerbated when intersecting performance with demographic attributes like race.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a super-smart, all-knowing robot assistant. You've trained it on every math book ever written, so it's basically a genius at solving problems. Now, you ask this robot to play a game: "Pretend you are a student who is really struggling with math. Make mistakes, get confused, and show frustration."
You might think, "Easy! Just tell the robot to act dumb." But this paper reveals a surprising twist: The robot refuses to play dumb. Even when you give it very specific instructions to be a "bad" student, it keeps trying to be a "good" student.
Here is a breakdown of what the researchers found, using some everyday analogies.
1. The "Actor Who Can't Forget Their Lines"
Think of Large Language Models (LLMs) like method actors who have spent years rehearsing the role of "The Perfect Expert." They are so good at this role that when you ask them to play "The Struggling Student," they can't break character.
- The Goal: The researchers wanted to see if these AI models could simulate a "counterfactual" persona—someone who performs the opposite of their natural talent.
- The Reality: When asked to solve a math problem as a "low-performing student," the AI often still got the right answer. It was like asking a professional chef to cook a burnt, under-salted meal, but the chef kept accidentally making a Michelin-star dish anyway.
2. The Two Ways to Measure "Bad Acting"
The researchers looked at the AI's performance in two ways:
- The Score (Accuracy): Did the AI get the right math answer?
- Result: Most AIs still got the right answer, even when told to fail. They couldn't bring themselves to get the math wrong.
- The Vibe (Reasoning Style): Did the AI sound like it was struggling?
- Result: This is where it got interesting. Some AIs (like OpenAI's o1) got the right answer but changed their "voice." They started saying things like, "Hmm, I'm not sure..." or "Let me count on my fingers..." even though the final number was correct. It was like an actor who says the right line but stutters and sweats while doing it.
3. The "Double Trouble" Experiment
The researchers then added a second layer: Race. They asked the AI to pretend to be a struggling student and to be a specific racial group (e.g., African American, White American, or Hispanic).
- The Analogy: Imagine asking the actor to play a struggling student and to speak with a specific cultural accent.
- The Result: Adding the race element made it even harder for the AI to act "bad" at math. The AI seemed to get confused by the extra instructions and actually performed better (got more correct answers) than when it was just asked to be a struggling student.
- The Bias: The AI also treated different racial groups slightly differently. For some groups, it was easier to make the AI "fail" at math; for others, it was harder. This suggests the AI has hidden biases about who is "good" or "bad" at math based on their background.
4. Why Does This Matter? (The "Virtual Classroom" Problem)
Why would anyone want an AI to pretend to be bad at math?
Imagine a virtual classroom where AI students help real students learn.
- The "Good" Student AI: Helps by solving problems perfectly.
- The "Struggling" Student AI: Helps by making common mistakes. Real students can look at these mistakes, spot the errors, and learn why they are wrong. This is called "learning by teaching."
The Problem: If the AI cannot convincingly pretend to be a struggling student, the "virtual classroom" loses its diversity. It's like a play where every single actor, no matter the script, insists on being the hero. You don't get a realistic story, and you don't get a realistic learning environment.
5. The Takeaway
The paper concludes that current AI models are too "helpful" to be "unhelpful." They are so trained to be accurate and logical that they struggle to simulate human error, hesitation, or confusion, even when explicitly told to do so.
- The Good News: The AI is safe and reliable; it won't accidentally give you wrong math answers.
- The Bad News: If we want to build diverse, realistic simulations (like virtual classrooms, therapy bots, or social experiments), we need to teach these robots how to be imperfect, confused, and human-like in their mistakes. Right now, they are just too perfect.
In short: You can tell a super-genius robot to "act dumb," but it will likely just act like a very nervous genius who still gets the answer right.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.