← Latest papers
🤖 AI

Incoherent by Design? On the Moral Self-Consistency of LLMs

This paper demonstrates that large language models exhibit significant moral self-inconsistency, with contradiction rates reaching up to 78% across ethically equivalent scenarios, thereby challenging the epistemic integrity and value alignment of AI systems used in morally sensitive contexts.

Original authors: Pegah Nokhiz, Aravinda Kanchana Ruwanpathirana, Helen Nissenbaum

Published 2026-08-18
📖 5 min read🧠 Deep dive

Original authors: Pegah Nokhiz, Aravinda Kanchana Ruwanpathirana, Helen Nissenbaum

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the modern world, we increasingly turn to artificial intelligence to help us navigate difficult choices. These systems, known as large language models, are trained on vast amounts of human text and can discuss complex topics, from history to science, with surprising fluency. When asked about right and wrong, they can often recite moral principles that sound very much like our own. We might ask them whether it is acceptable to lie to protect someone's feelings, or how to divide limited resources fairly. The hope is that these tools can offer consistent, reliable guidance in situations where human judgment might falter. However, for a system to be truly trustworthy in matters of ethics, it must do more than just sound reasonable; it must be consistent. If a person or a machine changes its mind about a moral rule simply because the question was asked in a slightly different way, its advice becomes unreliable. This raises a fundamental question about the nature of these digital minds: do they hold a stable set of beliefs, or are they merely mimicking the tone of the question they receive?

A team of researchers at Cornell University and Nanyang Technological University set out to test this stability. They wanted to see if these artificial intelligence systems could stick to a single moral philosophy when faced with nearly identical situations. To do this, they did not ask the models to choose between different ethical schools of thought, such as comparing a rule-based approach against a consequence-based one. Instead, they asked the models to adopt one specific viewpoint and see if they could apply it consistently. The researchers focused on three major ways humans have historically thought about morality. The first is a rule-based approach, where actions are judged as right or wrong based on whether they follow a universal law, regardless of the outcome. The second is a consequence-based approach, where the morality of an action depends entirely on whether it produces the best overall result for the most people. The third focuses on character, judging actions based on whether they reflect the virtues of a good person, such as honesty or courage.

The researchers designed a controlled experiment to isolate the models' reasoning. They created four specific moral scenarios, each highlighting a different tension, such as the conflict between telling the truth and preventing harm, or the trade-off between fairness and efficiency. For each scenario, they wrote three slightly different versions of the question. These versions were crafted to subtly encourage the same ethical perspective without explicitly naming it. For instance, in a scenario about a mother with dementia who has forgotten her husband's death, one question might emphasize the duty to tell the truth, while another might emphasize the duty to protect a vulnerable person from pain. Both questions were designed to elicit a rule-based response, but they framed the dilemma differently. The researchers then asked three different large language models to answer these questions. After receiving the responses, they translated the natural language answers into a structured logical format to check for contradictions. This allowed them to see if the model's reasoning held together when the same underlying situation was presented with a different twist.

The results revealed a startling lack of stability. In the most extreme cases, the models contradicted themselves in up to 78 percent of the scenarios. This means that when the same model was asked to reason from the same ethical standpoint about the same situation, it frequently arrived at opposite conclusions simply because the wording of the prompt had changed slightly. For example, a model might decide that lying to a dying patient is morally wrong when the question focuses on the violation of a rule, but then decide that lying is morally acceptable when the question focuses on the duty to prevent suffering, even though both questions were intended to test the same rule-based philosophy. The inconsistency was not limited to one type of model; it appeared across different systems, including those from major developers. The study found that the models were particularly unstable when dealing with the tension between truth and consequences, and between emotional motivation and cold reasoning.

These findings suggest that the internal logic of these systems is far more fragile than previously assumed. The models appear to be highly sensitive to the specific phrasing of a prompt, shifting their moral stance to match the tone of the question rather than adhering to a coherent set of principles. This is not merely a technical glitch; it points to a deeper issue with how these systems process information. If a system cannot consistently apply its own stated principles, it cannot be relied upon to offer stable guidance in real-world situations where people depend on it for advice. The researchers argue that before we can hope to align artificial intelligence with human values, we must first ensure that the system is internally consistent. Without this foundation of self-consistency, the goal of creating a trustworthy moral agent remains out of reach. The study does not claim that these models are incapable of moral reasoning, but rather that their reasoning is currently too unstable to be considered reliable. As these tools become more integrated into our lives, understanding their limitations is crucial for ensuring they serve us well without leading us astray.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →