Beyond Routing: Characterising Expert Tuning and Representation in Vision Mixture-of-Experts
This paper demonstrates that expert specialisation in vision Mixture-of-Experts models extends beyond simple category routing to encompass broader, stable tuning to continuous visual and semantic dimensions, which is best understood through fine-grained expert-level analyses rather than gating statistics alone.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a giant, high-tech kitchen with a head chef (the "gating network") and a team of specialized sous-chefs (the "experts"). When a customer orders a dish (an image), the head chef quickly decides which sous-chefs should work on it.
For a long time, researchers studying these AI kitchens only looked at the order slips. They would say, "Ah, the head chef sends all 'dog' orders to Sous-Chef A and all 'car' orders to Sous-Chef B." They assumed that because the orders were sorted by category, the chefs themselves were only experts in those specific categories.
This paper, however, decides to walk into the kitchen and watch the chefs actually cook. The authors, Gene Tangtartharakul and Katherine Storrs, built a vision AI model and used tools from human brain science to see what these "sous-chefs" were really doing.
Here is what they found, explained simply:
1. The Order Slip Lies (Routing vs. Reality)
The head chef's decisions (routing) do look like they are sorting by simple categories. If you look at the data, it seems like one chef only handles "animals" and another only handles "machines."
But when the researchers looked at what the chefs actually responded to most strongly (the "most exciting inputs"), the story changed.
- The Analogy: Imagine a chef who gets assigned all "deer" orders. You might think they are a "deer expert." But when you look at what makes them cook fastest and best, you realize they aren't actually thinking about "deer." They are obsessed with grainy textures and high-contrast patterns. It just so happens that the pictures of deer in the training data had a lot of grainy texture.
- The Takeaway: The head chef sorts by broad categories, but the individual chefs are actually tuned to specific visual features like texture, shape, color, and "spikiness," which cut across many different categories.
2. The "Alive vs. Not-Alive" Divide
Despite the chefs having different specific tastes (one likes "metallic" things, another likes "furry" things), the researchers found a massive, consistent pattern running through the whole kitchen.
No matter how many chefs they hired (4, 8, or 16), or how they started the training, the team always split the world into two big buckets:
- The "Alive" Team: Experts that responded best to organic, irregular, moving, or biological things.
- The "Not-Alive" Team: Experts that responded best to artificial, geometric, static, or manufactured things.
This is a bit like how the human brain has a specific area that lights up for faces and another for places. The AI naturally discovered this same "Alive vs. Not-Alive" split, even though no one told it to look for it. It's the most fundamental way to sort the visual world.
3. The "Taste" is Different, but the "Menu" is the Same
Here is a surprising twist. Even though the chefs were tuned to different specific features (one loves "metal," another loves "fur"), they all agreed on how to organize the menu.
If you asked any chef, "Is a cat more similar to a dog or a toaster?" they would all say "a dog." Even though they were using different sensory tools to get there, they all built the same mental map of the world. This suggests that while the tools the experts use are unique, the structure of their understanding is shared and stable.
4. Why This Matters
The paper argues that we've been looking at these AI models the wrong way. We've been too focused on the "who gets which order" (routing) and not enough on "what does the expert actually taste?" (tuning).
By using tools borrowed from neuroscience (like looking at what stimuli make a brain cell fire the most), the authors show that these AI experts aren't just simple category folders. They are complex, nuanced detectors for continuous features like "shininess," "roundness," or "organic-ness."
In short: The AI doesn't just sort pictures into "Animals" and "Cars." It sorts them into a rich, continuous spectrum of visual features, with a deep, natural split between things that are alive and things that are not. The "order slip" (routing) is just a rough sketch; the real masterpiece is in the detailed cooking (expert tuning).
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.