← Latest papers
💻 computer science

Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments

This paper demonstrates that while large language models often achieve high agreement with human moral judgments on final labels, they systematically diverge in the underlying moral principles and rationales used to reach those conclusions, revealing that label-based agreement is an insufficient proxy for true ethical alignment.

Original authors: Octavian M. Machidon, Alina L. Machidon, Vojko Strahovnik, Mateja Centa Strahovnik, Jonas Miklavčič, Marko Robnik Šikonja

Published 2026-08-14
📖 6 min read🧠 Deep dive

Original authors: Octavian M. Machidon, Alina L. Machidon, Vojko Strahovnik, Mateja Centa Strahovnik, Jonas Miklavčič, Marko Robnik Šikonja

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Great Moral Mimicry: When Robots Agree but Don't Understand

Imagine you are teaching a robot how to be a good person. You show it thousands of stories about people doing right or wrong things, and you ask it to give a simple grade: "Good" or "Bad." If the robot starts giving the exact same grades as your human friends, you might think, "Great! The robot has learned our values!" This is the world of AI alignment, a field where scientists try to make Artificial Intelligence (specifically Large Language Models, or LLMs) behave in ways that match human expectations.

For a long time, the main way to check if an AI is "aligned" has been to look at the final answer. If a human says, "That action is wrong," and the AI also says, "That action is wrong," we count it as a success. It's like a student getting the right answer on a math test; we assume they did the math correctly. But what if the student got the right answer by guessing, or by using a completely different formula that just happened to work this one time? In the world of AI ethics, getting the right label isn't enough. We need to know why the AI thinks something is wrong. If the AI's internal reasoning is totally different from ours, it might act strangely when faced with a new, tricky situation. This paper dives deep into that exact question: Do AI models really understand our moral reasons, or are they just really good at pretending?

The Great "Right Answer" Illusion

The researchers behind this study, a team from the University of Ljubljana, decided to put this "right answer" assumption to the test. They created a special challenge using 500 moral scenarios, ranging from everyday dilemmas (like "Is it okay to lie to a friend?") to more complex philosophical puzzles about justice and promises. They asked both human volunteers and several advanced AI models to do two things for every single scenario: first, give a final verdict (the "label"), and second, explain their reasoning (the "rationale").

Think of it like a game of "Two Truths and a Lie," but instead of guessing which statement is fake, you are trying to see if the AI and the human are thinking the same thoughts. The team looked at two different groups of models: the "Frontier" models (the super-smart, top-of-the-line ones like Claude and ChatGPT) and "Open" models (slightly smaller, open-source versions).

Here is the twist they found: The AI models were incredibly good at getting the final grade right, but they were often wrong about why it was right.

When the researchers looked only at the final "Good" or "Bad" labels, the AI models agreed with the human majority about 88% to 89% of the time. That's a huge score! If you were just checking the final answers, you'd say, "Perfect! The AI is aligned with us." But then, the researchers peeked under the hood at the explanations.

The "Why" Matters More Than the "What"

The study revealed that while the AI and humans often landed on the same destination, they took completely different roads to get there. The authors call this "divergent moral grounds."

To use a playful analogy: Imagine a human and a robot are both judging a person who stole a car.

  • The Human might say, "This is bad because they broke a promise to the owner and showed a lack of trust." They are focusing on the relationship and the broken promise.
  • The Robot, however, might say, "This is bad because it causes harm and is disrespectful." The robot is focusing on harm and politeness.

Both say "Bad," so the final score is a match. But their internal logic is totally different. The human is thinking about a specific social contract (promises), while the robot is thinking about general safety and manners.

The paper found that this mismatch happens all the time. In the "Deontology" section (which deals with duties and rules), humans often looked at whether an excuse was relevant to the rule. For example, if someone says, "I can't walk the dog because I walked him yesterday," a human might say, "That excuse is irrelevant; you still have a duty today." But the AI models often ignored the logic of the excuse and just said, "The whole situation is morally wrong." The AI was so busy seeing the "moral wrongness" of the act that it forgot to check if the excuse actually made sense in that specific context.

The Numbers Don't Lie (But They Don't Tell the Whole Story)

The data was quite clear. When the researchers measured how much the AI's reasoning matched the humans' reasoning, the scores dropped significantly.

  • For the "Commonsense" moral rules, the AI and humans only agreed on the specific reasons about 35% of the time (a Jaccard overlap of 0.350).
  • For the "Deontology" rules, it was even worse, with only 23.5% agreement on the reasons.

Even though the "Frontier" models (the smartest ones) agreed with humans on the final answer about 88.9% of the time, their reasons were often a different flavor entirely. They tended to overuse words like "harm" and "disrespect" and underuse words like "promise-keeping" or "trust." It's as if the AI has a vocabulary that is too narrow; it sees almost every problem as a safety issue or a rude gesture, missing the subtle, relational nuances that humans care about.

Why This Changes Everything

The main takeaway from this paper is a gentle warning: Agreement is not Alignment.

Just because an AI gives you the same answer you would give, doesn't mean it understands the world the way you do. If we only check the final answers, we might be lulled into a false sense of security. We might think our AI is a wise moral companion, when in reality, it's just a very good mimic that uses a different dictionary.

The authors suggest that we need to stop just looking at the "A" or "F" on the test and start reading the essay. We need to ask the AI why it made a choice. If the AI says something is wrong because it's "unsafe," but you think it's wrong because it's "unfair," that's a problem. In the real world, these differences matter. If an AI is moderating a chat room or giving legal advice, and it's using "harm" as its only lens when humans are thinking about "fairness" or "promises," it might make decisions that feel right on the surface but feel deeply wrong to the people involved.

So, the next time you see an AI that seems to agree with you perfectly, remember the lesson from this study: It might just be playing a very convincing game of "Guess the Answer," without truly understanding the story behind it. The path to truly safe and helpful AI isn't just about getting the right grade; it's about sharing the same reasons.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →