Representational Alignment Across Model Layers and Brain Regions with Multi-Level Optimal Transport
This paper introduces Multi-Level Optimal Transport (MOT), a unified framework that overcomes the limitations of standard layer-wise alignment by jointly inferring globally consistent, soft transport plans between neural network layers and brain regions, thereby enabling robust, interpretable comparisons across models of varying depths and architectures.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to compare two different cities: City A (a small, cozy town) and City B (a massive, sprawling metropolis).
Your goal is to understand how similar their "neighborhoods" are. In the world of Artificial Intelligence (AI) and neuroscience, these "neighborhoods" are layers of a computer network or parts of the human brain, and the "people" living there are neurons.
The Old Way: The "Best Match" Game
For a long time, scientists compared these cities using a method called Greedy Pairwise Matching.
Imagine you are a tour guide in City A. You look at your first neighborhood (Layer 1) and ask, "Which neighborhood in City B looks most like me?" You find the best match, say, Neighborhood 5. You write that down.
Then you move to your second neighborhood (Layer 2) and ask the same question. You find its best match, maybe Neighborhood 6.
The Problem:
- The "One-to-One" Trap: This method forces a strict rule: One neighborhood in City A must match exactly one in City B. But what if City A's "Layer 1" is actually a mix of features found in City B's "Layer 5, 6, and 7"? The old method can't see that. It forces a square peg into a round hole.
- The "Selfish" Tour Guide: Because each neighborhood picks its own best match independently, you might end up with 10 neighborhoods from City A all claiming they match Neighborhood 5 in City B, while Neighborhoods 1 through 4 in City B are completely ignored. It's chaotic and unbalanced.
- No Big Picture: You get a list of matches, but you don't get a single score that tells you, "Overall, how similar are these two cities?"
The New Way: Multi-Level Optimal Transport (MOT)
The authors of this paper propose a new framework called Multi-Level Optimal Transport (MOT). Think of this as a Master Logistics Planner who looks at the entire map of both cities at once.
Instead of asking, "Who is my best friend?" the planner asks, "How can we distribute the entire population of City A across City B so that everyone is happy, no one is left out, and the total travel cost is minimized?"
Here is how it works, using our city analogy:
1. The "Mass Distribution" (Soft Matching)
In the old method, a neighborhood had to pick one partner. In MOT, a neighborhood can split its "mass" (its importance or features).
- Example: Neighborhood 1 in City A might send 40% of its "traffic" to Neighborhood 5 in City B, 30% to Neighborhood 6, and 30% to Neighborhood 7.
- Why it helps: This perfectly handles the problem of different city sizes. If City A is small and City B is huge, City A's features can naturally spread out over several layers of City B without breaking the rules.
2. The "Global Harmony" (Consistency)
The planner ensures that the whole system is balanced.
- Every neighborhood in City A must send out 100% of its traffic (nothing is lost).
- Every neighborhood in City B must receive a fair share of traffic (nothing is ignored or overloaded).
- The Result: You get a single, clear score for the whole city comparison, and a map showing exactly how the two cities relate to each other globally.
3. The "Rotation" Trick
Sometimes, two neighborhoods might be doing the exact same thing, but they are just "rotated" (like a map turned sideways). The old methods might say they are totally different because the coordinates don't match.
The authors added a special feature to MOT (called MOT+R) that acts like a compass. It realizes, "Oh, this neighborhood is just the same as that one, but turned 90 degrees." It rotates the map to align them perfectly, revealing hidden similarities.
What Did They Discover?
The researchers tested this on three very different things:
- AI Models: Comparing different sizes of Large Language Models (like LLaMA and Qwen).
- Brains: Comparing the visual cortex of different humans looking at the same pictures.
- AI vs. Brains: Comparing computer vision models to human brains.
The Surprising Findings:
- Natural Order: Without being told to do so, MOT naturally figured out that "Early layers match early layers, and deep layers match deep layers." It found the hidden hierarchy that the old methods missed.
- Depth Differences: When comparing a small AI to a big AI, MOT showed that the small AI's single layer was actually "spread out" across three layers of the big AI. The old method just couldn't see this connection.
- Better Scores: In almost every test, MOT gave a more accurate and meaningful comparison than the old "Best Match" method.
The Big Takeaway
Imagine you are trying to compare two different recipes for a cake.
- The Old Way: You compare the flour in Recipe A to the flour in Recipe B, then the sugar to the sugar. If Recipe A has one bowl of flour and Recipe B has three bowls, you get confused.
- The MOT Way: You look at the entire process. You realize that the single bowl of flour in Recipe A is actually distributed across the mixing, baking, and cooling stages of Recipe B. You see the whole story of how the ingredients flow through the system.
MOT allows scientists to finally compare AI and brains (and different AIs) in a way that respects their complexity, handles different sizes, and reveals the true, hidden structure of how they think. It turns a messy, one-to-one guessing game into a harmonious, global alignment.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.