The yes-no bias of large language models reflects answer order and wording, not shifts in moral judgment
This paper demonstrates that the apparent "yes-no bias" in large language models' moral judgments is not a shift in ethical reasoning but a superficial artifact driven by answer order and lexical wording, which disappears when models are evaluated using a psychometric battery that separates logical verdicts from surface-level formatting.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to understand the moral compass of a very smart robot. You ask it a difficult question: "Is it okay to sacrifice one person to save five?"
If you ask the robot, "Do you say Yes or No?" it might seem to have a strong opinion. But this paper reveals a surprising trick: the robot's answer often depends less on what it thinks is right, and more on how you asked the question.
Here is the breakdown of what the researchers found, using simple analogies.
1. The Robot Has a Secret "Moral Scale"
First, the researchers wanted to know: Does the robot actually have a consistent set of values, or is it just guessing?
To find out, they didn't just ask the robot "Yes or No." Instead, they asked the same 20 moral dilemmas in 48 different ways. They changed the scale (0 to 10 vs. 0 to 100), flipped the wording, and changed which option they asked the robot to rate first.
The Finding: When you look past the surface, the robot's "inner voice" is actually quite consistent. It has a stable moral scale. If it thinks an action is "mostly bad" in one format, it will likely think it's "mostly bad" in another format, too. The smartest models (the "frontier" models) are surprisingly coherent in their internal logic.
2. The "Yes/No" Trap: The Magic of the Last Word
So, if the robot has a consistent inner voice, why do its "Yes/No" answers change so wildly when you reword the question?
The researchers discovered that the "Yes/No" format acts like a distorting lens. When forced to pick "Yes" or "No," the robot gets confused by two specific tricks:
- The "Last Word" Bias (Order Bias): Humans usually pay attention to the first thing they read. These robots, however, seem to pay extra attention to the last thing they read. If you write "Answer: No, Yes," the robot is more likely to pick "Yes" simply because it's at the end. If you write "Answer: Yes, No," it leans toward "No."
- The "No" Magnet (Lexical Bias): The robot seems to have a strange attraction to the specific word "No." It's not necessarily rejecting the idea logically; it's just drawn to that specific word on the page.
The Analogy: Imagine a person taking a test where the answers are printed on a card.
- Human: Reads the question, thinks about the answer, and picks the right one.
- Robot: Reads the question, thinks about the answer, but then gets distracted by the fact that the word "No" is printed in bold at the bottom of the card, or that "Yes" is the very last word they see. It picks the answer based on where the word is, not just what the word means.
3. The "No" Bias Isn't About Saying "No"
A common assumption was that these robots are just "negative" or "pessimistic" by nature—they just want to say "No" to everything.
The Finding: This is false. The researchers proved this by swapping the words. Instead of asking "Yes or No," they asked the robot to choose between "Option A" and "Option B."
When they did this, the "No" bias vanished. The robot didn't suddenly start saying "Yes" more often; it just stopped being pulled toward the word "No."
- Conclusion: The robot isn't biased against rejecting things. It is biased against the surface appearance of the question. It follows the printed label, not the logical verdict.
4. Thinking Harder Helps (But Doesn't Fix Everything)
The researchers tested what happens when they tell the robot to "think step-by-step" before answering (a process called deliberation).
The Finding: When the robot takes its time to think, the "Yes/No" bias gets smaller. The robot becomes less distracted by the order of the words and the specific word "No." However, the bias doesn't disappear completely for all models. Some models (like the "Claude" family) still show a strong pull toward the last option or the word "No," even when thinking hard.
5. The "Small" Robots Are Different
The study also looked at smaller, open-source models. These models were much messier. They didn't have a consistent internal moral scale at all. Their answers jumped around wildly depending on the question format, and they didn't get better even when they "thought" longer. They were essentially guessing based on the surface of the question.
The Big Takeaway
If you want to know what an AI truly values, you cannot just ask it "Yes or No" once.
- The Problem: A single "Yes/No" question mixes the robot's actual opinion with a "format artifact" (a glitch caused by how the question is written).
- The Solution: To measure an AI's true moral stance, you must ask the same question in many different ways (crossing the frames) and average the results. This cancels out the "last word" and "No word" distractions, revealing the robot's true, consistent inner scale.
In short: The robot isn't necessarily a "No" person. It's just a person who gets easily distracted by the font, the order, and the last word on the page. If you clean up the page, its true opinion shines through.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.