MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models
This paper introduces MENTIS, a geometry-first framework that reveals preference alignment induces selective, depth-localized geometric reorganization in language models, characterized by larger torsion shifts in normative concepts and negative correlations with contextual entropy, rather than uniform internal changes.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have two versions of a very smart robot assistant.
- Version A (The "Student"): This robot has read a lot of books and learned how to follow instructions. It's smart, but if you trick it with a sneaky question, it might accidentally say something rude or dangerous.
- Version B (The "Graduate"): This is the same robot after a special training camp called "Alignment." It's been taught to be helpful, harmless, and honest. It refuses to answer bad questions and behaves much better.
We know Version B acts better on the outside. But the big mystery is: What actually changed inside its brain? Did it just learn a few new "stop" words? Or did its entire way of thinking get rearranged?
This paper introduces a new tool called MENTIS to look inside the robot's brain and answer that question.
The Core Idea: "Belief" as a Direction
The authors don't think of a robot's "belief" like a human belief (e.g., "I believe the sky is blue"). Instead, they think of it as a compass needle.
Imagine that for every question you ask, the robot has a tiny compass needle pointing in the direction of the answer it wants to give.
- In the Student version, if you ask a tricky question, that needle might wobble or point toward a rude answer.
- In the Graduate version, that needle has been physically rotated to point toward a helpful, safe answer.
MENTIS is a tool that measures how much that needle twists and turns as the robot processes your question from the first layer of its brain to the last.
The Three Big Discoveries
The researchers used MENTIS to compare the Student and Graduate robots. Here is what they found, explained with simple analogies:
1. The Change is Selective, Not Uniform
The Analogy: Imagine you are painting a house. You might think "Alignment" is like painting the whole house a new color (a uniform change).
The Reality: MENTIS found that the paint job is actually very specific.
- When the robot thinks about values (like "Justice" or "Peace"), its internal compass needles twist and turn a lot. The robot reorganizes its thinking significantly to handle these topics.
- When the robot thinks about facts (like "What is the capital of France?"), the needles barely move.
- Conclusion: The training didn't just add a generic "safety filter." It specifically rewired how the robot handles moral and value-based concepts.
2. The Change Happens in "Quiet" Moments, Not Chaos
The Analogy: You might guess that a robot changes its mind most when it is confused or unsure (high entropy).
The Reality: The researchers found the opposite. The biggest "twists" in the robot's internal compass happened when the context was clear and structured.
- When the robot was confused or the situation was chaotic, the internal changes were small.
- When the robot knew exactly what kind of answer was expected, the training camp had the biggest effect on how it oriented its thoughts.
- Conclusion: Alignment works best when the robot is confident enough to have a clear direction, and then it gently steers that direction toward safety.
3. The "Twist" Happens at Specific Floors
The Analogy: Imagine the robot's brain is a 30-story building. You might think the safety training happens on the ground floor (where it starts) or the roof (where it finishes).
The Reality: The "twist" happens on different floors for different robot models.
- For one model (OLMo), the biggest change happened on the 29th or 30th floor.
- For another model (Mistral), the biggest change happened on the 14th floor.
- Conclusion: There is no single "safety layer." Each robot architecture has its own specific "floor" where the training camp did the most work.
A Surprising Twist: Safe Prompts Change More
Usually, we think safety training is all about stopping bad things. You'd expect the robot's brain to change the most when it sees a "bad" (unsafe) prompt.
The Surprise: The robot's brain actually changed more when it was answering safe, helpful prompts.
- Why? The researchers suggest that saying "no" to a bad question is easy and automatic (like a reflex). But saying "yes" to a good question while staying polite, helpful, and safe requires a much more complex internal dance. The robot has to carefully balance being helpful without crossing the line, which causes a bigger reorganization of its internal compass.
What This Means (and What It Doesn't)
- What it means: Alignment isn't just a surface-level patch. It leaves a specific, measurable "geometric signature" deep inside the robot's brain. It changes how the robot turns its thoughts, but it does so selectively (mostly for values) and in specific locations (different floors for different models).
- What it doesn't mean: This paper is a diagnostic tool, like an X-ray. It shows us where the bones are broken or changed, but it doesn't prove that fixing that specific bone is the only reason the robot behaves better. It also doesn't tell us how to use this to build better robots yet; it just tells us what the current ones look like on the inside.
In short: MENTIS shows us that when we teach a robot to be "good," we aren't just adding a stop sign. We are gently reshaping its internal compass, especially when it's thinking about values, and we do it in very specific, unique ways for every different robot model.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.