← Latest papers
💬 NLP

MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning

This paper introduces MET, a theory-grounded and culture-aware framework featuring the MCLASH benchmark and a self-distillation training method (MET-D) to enhance multilingual moral reasoning by adapting to cultural contexts and leveraging expert-curated ethical principles without requiring external supervision.

Original authors: Ayoung Lee, Ryan Kwon, Yunxiang Zhang, Yuxuan Liu, Peter Railton, Lu Wang

Published 2026-07-14
📖 5 min read🧠 Deep dive

Original authors: Ayoung Lee, Ryan Kwon, Yunxiang Zhang, Yuxuan Liu, Peter Railton, Lu Wang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you're trying to solve a moral puzzle, like deciding whether a firefighter should be punished for taking wood from a forest. Now, imagine you have to solve this puzzle not just in English, but in Korean, Hindi, Spanish, and Malay. If you just take the English puzzle and run it through a translator, you end up with a weird mess. Suddenly, the firefighter is at a "Wendy's" in Oregon during "Thanksgiving"—concepts that might make no sense or feel totally wrong to someone in Seoul or Mumbai.

That's the problem a team of researchers at the University of Michigan tackled in their new paper, MET. They found that current AI models are like tourists who only speak English; they try to apply American rules to the whole world, and it often leads to clumsy, culturally blind decisions.

The Problem: One Size Does Not Fit All

The authors argue that simply translating moral dilemmas is a bad idea. It's like trying to wear a winter coat in the tropics; the fabric is right, but the fit is terrible. Existing AI benchmarks often just translate English scenarios directly, ignoring local customs, currencies, and institutions. Furthermore, the AI's "thinking" process usually relies on static, English-centric rules that don't account for how different cultures actually reason about right and wrong.

The Solution: A Cultural Toolkit

To fix this, the team built a new playground called MCLASH. Instead of just translating old puzzles, they rebuilt them from the ground up for five new languages (Chinese, Hindi, Korean, Malay, and Spanish). They swapped out American landmarks for local ones (like changing "Amazon" to "Naver" in Korea) and adjusted the stories so they felt natural to local speakers. This resulted in 1,852 long, complex scenarios.

Then, they created MET (Multilingual Ethics with Theory-grounded reasoning). Think of MET as giving the AI a massive, organized toolbox of "moral grounds" drawn from philosophy and psychology. These aren't just random ideas; they are organized into six big categories, like "Value Systems" (what people care about) and "Conflict Handling" (how to solve arguments).

Here's how MET works in two steps:

  1. Picking the Tools: Before solving the puzzle, the AI looks at the specific situation and the culture involved, then picks the right tools from its toolbox. For an English speaker, it might pick "First-principles reasoning" (logical, step-by-step thinking). For a Korean speaker, it might pick "Contractarianism" (focusing on mutual social agreements), which fits better with Confucian values.
  2. Solving the Puzzle: The AI then uses those specific tools to reason through the problem in the user's native language.

The Secret Sauce: Self-Training (MET-D)

Just picking the right tools wasn't enough. The researchers noticed that even when the AI picked the right tools, it often ignored most of them, leading to shallow answers. It was like a chef picking ten spices but only using salt.

To fix this, they introduced MET-D (MET-Distillation). This is a clever training trick where the AI teaches itself. They created a synthetic dataset where the "correct" answer is obvious based on the character's values (so no human had to grade every single answer). The AI practiced on this data, learning to actually use the tools it picked, not just list them. It's like a student doing practice problems until they stop guessing and start understanding the logic.

What They Found

The results were promising. When they tested MET-D on three different AI models (Qwen3-4B, Qwen3-8B, and Gemma3-4B), it improved the models' performance significantly.

  • On their new MCLASH benchmark, the average score went up by 3.71 points.
  • On another dataset called MMoralExceptQA, it improved by 4.23 points.
  • The biggest jump was for Malay speakers using the Qwen3-8B model, which saw a massive 12.94 point gain.

But the most surprising discovery wasn't just about scores. The researchers found that the "best" tools for one culture were often different for another. For example, English speakers benefited most from "Probabilistic Moral Uncertainty" (weighing odds), while Spanish speakers did best with "Relation-Based Authority" (focusing on family and social bonds). This proves that moral reasoning isn't universal; it's deeply tied to culture.

Also, MET-D made the AI much better at thinking in the user's native language. Before training, the AI often thought in English even when talking to a Korean user. After training, the AI's internal reasoning chain became 62.13 percentage points more likely to be in the native language. This makes the AI's decisions much easier for people to understand and trust.

What They Didn't Find (and What's Still Unknown)

The paper is careful to note what it didn't do. They didn't find a way to make the AI perfectly reconcile conflicting moral rules (like when two chosen tools fight each other); the AI still sometimes just lists them without fully merging them. Also, while the training helped the AI use the tools better, it didn't necessarily make the AI better at picking the tools in the first place; that part still relies on the base model.

The authors suggest that cross-language learning follows the rules of language structure (like how sentences are built) rather than just cultural similarity. For instance, Korean and Hindi, which share a similar sentence structure, helped each other more than Korean and Chinese did, even though Korean and Chinese share more cultural history.

In short, the paper suggests that to make AI truly moral across the globe, we can't just translate English rules. We need to build culturally aware toolkits and teach the AI to use them deeply in the local language. It's a step toward AI that doesn't just speak your language, but thinks like you.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →