← Latest papers
💬 NLP

Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability

This paper demonstrates that the Knobe effect, a moral bias in intentionality judgments, emerges in finetuned large language models but can be localized to specific layers and mitigated through targeted layer-patching interventions without retraining.

Original authors: Bianca Raimondi, Daniela Dalbagno, Maurizio Gabbrielli

Published 2026-07-20
📖 4 min read☕ Coffee break read

Original authors: Bianca Raimondi, Daniela Dalbagno, Maurizio Gabbrielli

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot how to be a good judge. You don't just give it a rulebook; you show it thousands of stories about people making choices and ask, "Was that person being good or bad?" Over time, the robot learns patterns. But here's the tricky part: humans aren't perfect judges. We often have hidden biases. For example, if a person accidentally causes a bad result, we are much more likely to say, "They meant to do it!" than if they accidentally caused a good result. This is a famous psychological quirk called the "Knobe effect." It's like our brains have a built-in filter that makes us think bad accidents were actually on purpose, while good accidents were just luck.

Now, scientists are worried that if we teach robots to act like us, they might pick up these same weird filters. If a robot makes moral decisions for us, we need to know: does it think like a human, or is it just a giant calculator? This is where "mechanistic interpretability" comes in. Think of a giant robot brain as a massive city with millions of tiny workers (neurons) passing notes to each other. Mechanistic interpretability is like putting on X-ray glasses to see exactly which workers are passing which notes. Instead of just guessing why the robot gave a certain answer, we can peek inside the machine to see the specific gears turning. This matters because if we can find exactly where a bias lives inside the robot, maybe we can fix just that one gear without having to rebuild the whole city.

The researchers in this paper decided to put this idea to the test. They wanted to see if popular AI models (specifically Llama, Mistral, and Gemma) had picked up the "Knobe effect" and, if they had, whether they could find the exact spot in the AI's brain where this bias was hiding. They started by testing the AI models on the same moral stories used with human volunteers. They found that the "raw" AI models, which hadn't been taught to talk like humans yet, didn't really show this bias. Their answers were pretty balanced. However, once the researchers "finetuned" the models—training them to follow instructions and align with human preferences—the bias suddenly appeared. The AI started acting just like us: it was much more likely to say a bad outcome was intentional than a good one.

But the real magic happened when they looked inside the machine. Using a technique called "Layer-Patching," the researchers treated the AI like a set of stacked layers, like floors in a skyscraper. They suspected that the bias wasn't spread out everywhere, but was stuck in a specific set of floors. To test this, they took the "thoughts" (activations) from the unbiased, raw model and swapped them into the biased, finetuned model, one floor at a time. It was like taking a clean, unbiased blueprint and pasting it over a specific floor of a corrupted building to see if the whole structure became honest again.

The results were surprisingly precise. They discovered that the bias wasn't scattered randomly; it was localized in the middle-to-upper layers of the AI's brain. Even more exciting, they found that swapping just one specific layer from the unbiased version into the biased model was enough to almost completely erase the Knobe effect. For some models, the bias dropped from a strong 1.6 or 3.8 difference down to nearly zero (0.00 or 0.03). And the best part? They did this without retraining the model or changing its weights. It was a surgical fix. They also checked to make sure this "surgery" didn't break the robot's ability to do other things, like answering general questions, and found that the models stayed just as smart as before.

The paper suggests that these social biases aren't some mysterious, unfixable magic that emerges from the whole system. Instead, they are specific, localized patterns that get written into the AI's brain during the training process where we try to make it act like a human. By using these X-ray glasses and surgical swaps, the researchers showed that we can spot these biases and remove them with targeted interventions, keeping the AI helpful without the unwanted human quirks.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →