← Latest papers
🤖 AI

Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture

This paper argues that current evaluations of AI moral reasoning overly focus on value alignment while neglecting context-sensitive norm application, and it proposes a research agenda to address this gap through standardized formal representations, expert-annotated datasets, and distinct evaluation protocols for normative competence.

Original authors: Aidan Kierans, Ritam Dutt, Kaley Rittichier, Shiri Dori-Hacohen, Avijit Ghosh

Published 2026-08-18
📖 5 min read🧠 Deep dive

Original authors: Aidan Kierans, Ritam Dutt, Kaley Rittichier, Shiri Dori-Hacohen, Avijit Ghosh

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

When people ask a computer for advice on a difficult choice, they are often asking a question that has two very different parts. The first part is about what we care about: do we value kindness over strict rules, or fairness over loyalty? The second part is about how we use those cares to make a decision in a specific situation. Imagine you are trying to decide whether to tell a painful truth to a friend. Knowing that you value honesty is one thing; knowing exactly how that value applies when the truth might hurt someone you love is another. This second part requires following a set of principles that tell you how to weigh different concerns against each other. For years, scientists studying artificial intelligence have focused almost entirely on the first part, checking if a machine's answers match what humans generally prefer. But a new analysis suggests that by stopping there, researchers have missed the most important test of whether a machine can actually think like a moral agent.

A team of researchers from the University of Connecticut, Carnegie Mellon University, and Hugging Face has examined how we currently test the moral reasoning of large language models. These are the powerful computer programs that can write essays, answer questions, and offer advice. The authors argue that the field has become too focused on a "moral value problem," which simply asks if the computer's answers line up with human opinions. They contend that this approach leaves the "moral norm problem" largely unexplored. The moral norm problem is the harder challenge of seeing if a machine can identify the right principles for a situation and apply them correctly to reach a conclusion. The researchers found that current tests are like checking if a student has memorized a list of values, but never asking if they can use those values to solve a complex problem.

To understand why this distinction matters, consider how scientists have traditionally measured morality in both humans and machines. They have relied heavily on frameworks like Moral Foundations Theory, which breaks down human morality into broad categories such as care, fairness, and loyalty. These tools are excellent for describing what people tend to care about. If you ask a computer to choose between saving one person or five people, and it chooses the five, researchers can say the computer aligns with the human value of caring for the greater number. However, the researchers point out that caring about a value is not the same as knowing how to apply it. A machine might agree that "harm is bad" but fail to understand that in some ethical systems, causing harm to one person to save five is a violation of a different rule. The current tests often mistake a machine's ability to say the right words for its ability to do the right reasoning.

The authors reviewed dozens of existing studies and benchmarks used to evaluate AI morality. They discovered that almost all of these tests cluster around the moral value problem. These tests ask the computer to pick an answer from a list or to describe a preference, and then compare that choice to what a group of humans would choose. While this tells us what the machine prioritizes, it does not tell us if the machine can build a valid argument based on a specific ethical theory. The researchers found that very few tests actually check if a machine can take a set of principles, apply them to a new and tricky situation, and explain its reasoning step by step. When tests do try to look at reasoning, they often use the same broad value categories as a shortcut, assuming that if a machine cares about "fairness," it must know how to apply the rules of fairness. The authors argue this is a mistake, because different ethical theories can all claim to care about fairness while demanding completely different outcomes.

This confusion creates a gap in our understanding of what these machines can actually do. The researchers identified three specific missing pieces in the current evaluation landscape. First, there is no high-quality, shared set of correct answers for how moral rules should be applied in different situations. Without a standard reference, every research team builds its own test, making it impossible to compare results or see real progress. Second, the tests do not look closely enough at the steps a machine takes to reach a conclusion. A machine might get the right answer by luck or by guessing, but current methods often fail to catch if the reasoning behind that answer is flawed. Third, the tests do not check if the machine can spot the most important details in a story. Before a machine can apply a rule, it must first figure out which parts of a situation matter morally, and current tools do not measure this skill well.

To fix these issues, the authors propose a new path forward for the field. They suggest that scientists need to create a common language for describing ethical theories, so that different researchers can build tests that speak to the same principles. They call for the creation of new datasets where experts in ethics annotate how specific rules should be applied to real-world scenarios. Finally, they urge the community to stop mixing up the test of what a machine values with the test of how it reasons. A machine should be judged on its ability to follow a logical path from a principle to a decision, not just on whether its final choice matches human opinion. The researchers conclude that while we have made good progress in understanding what machines care about, we have only just begun to understand if they can truly think through a moral problem. Until we build the tools to test the application of rules, we will not know if these systems are capable of the kind of moral reasoning that humans rely on in their most difficult moments.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →