Shared Semantics, Divergent Mechanisms: Unsupervised Feature Discovery by Aligning Semantics and Mechanisms
This paper introduces a distribution-level unsupervised feature discovery method that clusters LLM continuations by jointly optimizing semantic coherence and mechanistic attribution signatures, thereby revealing diverse continuation modes and actionable internal mechanisms that single-view, target-conditioned analyses often miss.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Problem: The "One-Answer" Trap
Imagine you are a detective trying to understand how a magic box (a Large Language Model) works. Usually, when you ask the box a question, it gives you one answer. To figure out how the box works, you might look at the gears and springs inside only for that specific answer.
The paper argues that this is a trap. If you only look at one answer, you might miss the fact that the box has many different "modes" of thinking. Sometimes, two answers might sound very similar (same meaning) but were built using completely different internal gears. Other times, two answers might sound totally different but were built using the exact same gears.
The authors call this the "Misalignment between semantic and mechanistic spaces." It's like two houses that look identical from the outside (same paint, same windows) but have completely different plumbing systems inside. Or, two houses that look totally different (one is a castle, one is a hut) but share the exact same plumbing.
The Solution: A New Way to Group Answers
Instead of picking one specific answer to study, the authors propose a new method to look at all the possible answers the box could give to a single question at once. They call this "Distribution-Level Unsupervised Feature Discovery."
Here is how their method works, broken down into three simple steps:
1. The "Sample the Crowd" Step
Instead of asking the box for one answer, they ask it for hundreds of different answers to the same question.
- Analogy: Imagine asking a chef, "How do you make soup?" instead of getting one recipe, you get 500 different variations. Some are spicy, some are creamy, some use chicken, some use tofu.
2. The "Double-View" Step
For every single soup recipe, the team creates two different "ID cards":
- The Semantic Card (What it says): This describes the meaning of the soup. Is it spicy? Is it a tomato soup?
- The Mechanistic Card (How it's made): This describes the internal gears the model used to create it. Which specific neurons fired? Which internal "switches" were flipped?
- Analogy: For every soup, you write down the flavor profile (Semantic) and the kitchen equipment list used (Mechanistic). You realize that two soups might taste the same (Semantic match) but one was made with a blender and the other with a mortar and pestle (Mechanistic mismatch).
3. The "Smart Sorting" Step
Now, they have a huge pile of soup recipes, each with two ID cards. They use a special mathematical sorting machine (called Rate-Distortion Clustering) to group these recipes.
- The Goal: They want to find groups where the soups are similar in both flavor and equipment.
- The Trade-off: The machine has a "knob" (called ) that controls how strict the sorting is.
- Loose setting: It groups soups that are just generally "soup-like."
- Strict setting: It separates them into tiny groups like "Spicy Tomato Soup made with a Blender" vs. "Spicy Tomato Soup made with a Mortar."
- The Result: They find hidden "modes" of thinking that a human looking at just one answer would never see.
Why This Matters: The "Steering" Test
The paper doesn't just say, "Look, we found groups." They prove these groups are real by doing a "steering" test.
- The Test: They take a group of soups that the machine identified as "Spicy Tomato made with a Blender." They then try to force the magic box to make more soups like that by turning up the "Blender" switch inside the machine.
- The Result: When they turn up the switch, the box actually starts making more "Spicy Tomato" soups.
- The Comparison: When they tried to do this using only the "flavor" (Semantic) or just picking one random soup, the steering didn't work as well.
- Analogy: It's like realizing that to get a specific type of cake, you don't just need to say "I want cake"; you need to know exactly which mixer and oven settings produce that specific texture. The authors found the "mixer settings" for different types of thinking.
The Main Takeaway
The paper introduces a tool that helps us audit AI models by looking at the whole landscape of possible answers, not just one.
- Old Way: Pick one answer, find the gears that made it. (Like studying one car to understand how all cars work).
- New Way: Look at all the cars the factory could build, group them by how they are built and what they look like, and find the hidden patterns.
This helps researchers understand that an AI model isn't just a single path of logic; it's a complex landscape with many different "paths" (mechanisms) that can lead to similar or different results. By mapping these paths, we can better understand, audit, and control how these models think.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.