Framing Instability in LLM Ethical Stance: Auditing Negation Sensitivity in Moral Dilemmas
This paper audits 16 language models across ethical dilemmas and reveals that their moral stances are highly unstable and prone to flipping based on whether questions are framed as affirmations or negations, a vulnerability that is particularly severe in smaller models and often masked by standard binary evaluation methods.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a super-smart robot friend who is really good at giving advice on tough moral questions, like "Should I tell the truth if it hurts someone?" or "Is it okay to break a rule to save a life?" You'd expect this robot to have a solid moral compass, right? If you ask, "Is stealing bad?" it should say "Yes." If you ask, "Is not stealing good?" it should also say "Yes." The answer should stay the same because the situation hasn't changed, only the words you used.
But here's the twist: a new study by Katherine Elkins and Jon Chun found that for many of these AI robots, the answer changes completely just because you flipped the sentence around. It's like asking a friend, "Do you like pizza?" and getting a "Yes!" Then asking, "Do you dislike pizza?" and getting a "No!" (which means they like it), but then asking, "Do you not dislike pizza?" and suddenly getting a "Yes!" again, even though the robot is acting like it's confused about what pizza is.
The Great Flip-Flop
The researchers tested 16 different AI models (some made by big US companies, some by Chinese companies, and some smaller, open-source ones) with 14 tricky moral dilemmas. They asked each model the same question in two ways:
- The "Should" way: "They should rob the store."
- The "Should Not" way: "They should not rob the store."
If the robot is stable, it should say "No" to the first one and "Yes" to the second one (meaning it agrees that not robbing is the right choice). But for the smaller, open-source models, the results were wild. When asked "They should rob the store," these models only agreed about 24% of the time. But when asked "They should not rob the store," they suddenly agreed to the action 77% of the time!
That's a swing of 76 percentage points. It's as if the robot looked at a red light and said "Stop," but when you said "Don't go," it said "Go!" The study found that for some of these smaller models, the flip was even more extreme, reaching 100% agreement under certain tricky phrasings.
The "Yes-Man" Problem
Why does this happen? The authors suggest the robots aren't actually thinking through the logic. Instead, they might be playing a game of "keyword matching." If the prompt says "rob," the robot sees the word "rob" and thinks, "Oh, I need to talk about robbing!" and accidentally agrees with the action, even if the sentence says "should not." It's like a student who sees the word "not" in a math problem but ignores it because they're too focused on the numbers.
The study also checked if the robots were just being "yes-men" (agreeing with whatever the user said). If they were, they would agree with "Don't rob" by saying "Okay, I won't rob." But they didn't. They agreed with "Don't rob" by saying "Okay, I will rob." So, it's not just being a yes-man; it's a genuine glitch in how they handle negative words.
The "Fake" Agreement
Here's another sneaky part: the way we usually test these robots is flawed. Most tests just ask, "Do you agree or disagree?" and count the "agree" answers. The researchers found that this simple "agree/disagree" button is a liar. It counts 38% more "agreements" than actually exist.
Imagine a robot says, "I'm not sure, but maybe it's okay," and the test counts that as a "Yes." The researchers had humans read the answers and found that the robots often wrote reasons that contradicted their own "Yes" or "No" buttons. About 13% of the time, the robot's written explanation said the opposite of what its button said. It's like a student writing an essay arguing against stealing, but then checking the box that says "I support stealing."
Who's the Most Stable?
Not all robots are equally confused.
- The Small, Open-Source Models: These were the most fragile. They flipped their answers wildly depending on how the question was asked.
- The Big Commercial Models: The models from big companies (like GPT-5, Claude, and Gemini) were better, but still not perfect. They flipped less, but they still changed their minds. For example, the agreement between different models dropped from 73% when the question was positive to 59% when it was negative. This means if you asked two different "smart" robots the same question in a negative way, they were much more likely to disagree with each other.
- The "Reasoning" Models: Some models that were told to "think step-by-step" before answering did better. One model, Grok-4.1-reasoning, reduced its confusion by 57% compared to its non-reasoning version. This suggests that if the robot slows down and actually reads the whole sentence instead of just skimming for keywords, it gets it right more often.
Does It Matter?
The researchers warn that this isn't just a fun party trick. If a hospital uses an AI to help decide patient care, or a bank uses one to approve loans, and the AI's answer changes just because the doctor or banker phrased the question differently, that's a huge problem.
They found that the robots were especially shaky in financial and business scenarios (with a "sensitivity score" of 0.63–0.65) compared to medical ones (0.36). This means the robots might be giving the most unreliable advice exactly when people are dealing with money or debt.
The Bottom Line
The paper doesn't say these robots are "broken" forever, but it does say they are currently unreliable for high-stakes decisions if you don't check them carefully. The authors propose a new test called the Negation Sensitivity Index (NSI) to measure how much a robot's answer flips when you flip the words.
They suggest that until robots can pass this test, we shouldn't let them make big decisions on their own. If a robot's answer depends on whether you say "should" or "should not," it's not really making a judgment about the situation; it's just reacting to the sentence structure. And in the real world, we need our AI friends to be consistent, not confused.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.