Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models
This paper introduces a benchmark using the Moral Foundations Questionnaire to quantify moral susceptibility and robustness in large language models under persona role-play, revealing that model family primarily determines robustness (with Claude being the most robust) while model size drives susceptibility, and that these two properties are positively correlated.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a group of very smart, digital "actors" (these are the Large Language Models, or LLMs). Usually, these actors try to be helpful, polite, and neutral. But what happens if you tell them, "Okay, today you are not a helpful assistant. Today, you are a grumpy old pirate," or "Today, you are a strict school principal"?
This paper asks a simple but deep question: How much does an AI's sense of "right and wrong" change when it puts on a different costume?
To answer this, the researchers created a "Moral Suitability Test" using a famous psychology tool called the Moral Foundations Questionnaire (MFQ). Think of the MFQ as a personality test for morality. It asks questions about five core values:
- Care/Harm: Do you care about hurting others?
- Fairness: Do you believe in equal treatment?
- Loyalty: Do you stick with your "team"?
- Authority: Do you respect rules and leaders?
- Purity: Do you care about cleanliness or sacredness?
The researchers tested 15 different AI models (like Claude, Gemini, GPT-4, and Grok) by asking them to answer these moral questions while pretending to be 100 different characters (from a "disappointed factory worker" to a "religious family member").
Here are the two main things they measured, explained with simple analogies:
1. Moral Robustness: The "Steady Hand" Test
What it is: If you ask the AI the same moral question 10 times while it's playing the same character, does it give the same answer every time?
- High Robustness: The AI is like a lighthouse. No matter how the waves (random computer noise) hit it, the light stays steady. It knows exactly what "Captain Jack Sparrow" believes about stealing, and it sticks to that belief consistently.
- Low Robustness: The AI is like a weathervane in a storm. Even if you ask the same question twice while it's playing the same role, it might flip-flop between answers. It's confused or unstable.
The Finding: The "family" the AI belongs to matters most here. Claude models were the most "steady hands" (very robust). Grok models were the most "wobbly." Interestingly, making the AI "smarter" (bigger size) didn't necessarily make it steadier.
2. Moral Susceptibility: The "Chameleon" Test
What it is: If you change the character from a "Pirate" to a "Teacher," how much does the AI's answer change?
- High Susceptibility: The AI is a chameleon. It changes its colors (moral views) instantly to match the environment. If you tell it to be a pirate, it suddenly thinks stealing is okay. If you tell it to be a nun, it suddenly thinks stealing is terrible. It has no core self; it just mirrors the role.
- Low Susceptibility: The AI is a rock. It has a strong internal compass. Whether you tell it to be a pirate or a nun, its core beliefs about fairness or harm stay mostly the same. It says, "I'm playing a pirate, but I still think stealing is wrong."
The Finding: This is where size matters. Larger, more complex AI models were actually more susceptible. They were better at "acting" and shifting their moral views to fit the character perfectly. Smaller models were more stubborn and kept their original views.
The Big Surprises
- The "Acting" Paradox: The models that were best at staying consistent (Robust) were also the ones that changed the most when you gave them a new role (Susceptible). It's like a great actor who is very consistent in their performance, but also very good at transforming into completely different people.
- The "Grok" Outlier: The Grok models were the worst at both. They were shaky when playing the same role, and they changed their minds wildly when the role changed. They were the most unpredictable.
- The "Claude" Champion: The Claude family was the most reliable. They were the most consistent actors and the most stable when the role changed.
Why Does This Matter?
Imagine you are using an AI to help mediate a conflict between two people.
- If the AI has low robustness, it might give you different advice every time you ask, making it useless.
- If the AI has high susceptibility, it might adopt the moral views of the "villain" in the story you are telling it, potentially justifying bad behavior just because the character said so.
In short: This paper gives us a new way to measure how "stable" and "flexible" an AI's conscience is. It tells us that while AI is getting better at acting like humans, we need to be careful that they don't become too good at pretending, to the point where they lose their own moral compass.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.