Where You Measure Decides What You Measure: Position Selection in Ablation-Based SAE Evaluation
This paper argues that the standard practice of evaluating sparse autoencoders by measuring ablation effects at the token where a latent fires most strongly introduces a critical confounding variable, as this position is determined by the dictionary itself rather than the experimenter, leading to misleading comparisons that can be resolved by enforcing a consistent measurement position across different dictionaries.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to understand how a giant, magical library thinks. Inside this library, the "brain" is made of millions of tiny, glowing switches. When the library reads a story, certain switches light up to help it understand the meaning. Scientists have built a special tool called a Sparse Autoencoder (or SAE) to act like a translator. This tool watches the glowing switches and tries to give them names, like "switch for 'cat'" or "switch for 'future tense'."
Once the translator has named the switches, the next big question is: Which ones actually matter? To find out, scientists usually play a game of "what if." They take a specific named switch, turn it off, and see if the library's story changes. If the story goes weird, that switch was important. If the story stays the same, maybe it wasn't doing much. This is called an ablation test. It's a standard way to figure out what makes the library tick. But here's the tricky part: a single switch doesn't just light up once; it lights up thousands of times in different places in the story. So, when you decide to turn it off, where do you turn it off? Do you pick the first time it lights up? The last time? Or the time it shines the brightest?
This paper asks a simple but startling question: Does it matter where you pick to turn the switch off? The author found that the answer is a resounding "yes," and that the way scientists have been doing this test for years might be leading them to the wrong conclusions.
The "Brightest Spark" Trap
For a long time, the standard rule for these tests has been: "Turn off the switch at the moment it shines the brightest." It feels like a smart instinct. If a lightbulb is super bright in one spot, that's probably where it's doing its most important work, right?
The problem, as this paper explains, is that "where it shines the brightest" isn't a fixed fact about the library. It's a fact about the translator (the SAE) you are using. Imagine two different translators looking at the same library. They both agree that Switch #42 is the "cat switch." But Translator A thinks the cat switch shines brightest at the word "kitty," while Translator B thinks it shines brightest at the word "feline."
If you follow the standard rule, Translator A will turn off the switch at "kitty," and Translator B will turn it off at "feline." Even though they are testing the exact same concept, they are testing it in two completely different places in the story. The paper shows that this isn't a tiny detail; it's a massive source of confusion.
The Great Switch-Off Experiment
To prove this, the researcher didn't just guess; they ran a massive, controlled experiment. They built six different translators (SAEs) that were almost identical twins. They started with the exact same blueprint and trained them on the exact same data, changing only tiny, harmless settings. Because they started the same, every "Switch #42" in all six translators was supposed to mean the exact same thing.
Then, they ran the standard test:
- The Old Way: They let each translator pick its own "brightest spot" to turn the switch off.
- The New Way: They forced all six translators to turn the switch off at the exact same spot in the story.
The Result was Shocking.
When they used the old way (letting each translator pick its own spot), the results were all over the place. The translators seemed to disagree wildly about how important the switch was. One said it was super important; another said it did nothing. The researcher calculated that about 7.6% to 11.9% of the disagreement between the translators was actually just because they were looking at different spots in the story.
But when they forced everyone to look at the same spot, the disagreement almost vanished. The "disagreement" dropped to nearly 0% for one model and 2.4% for another.
It turns out that most of the time scientists thought the translators were arguing about the meaning of the switch, they were actually just arguing about where to look. The "brightest spot" rule was making them look in different places, creating an illusion of disagreement.
The "Where" Matters More Than the "What"
The paper reveals a hidden truth: The location where you measure the effect is more important than the dictionary you use.
In fact, the researcher found that the "noise" caused by picking different spots was huge. In their data, the variation caused by where you measure was 67.4% of the total difference. That means if you change the spot where you turn the switch off, you change the answer more than if you changed the entire translator itself.
This gets even more interesting as the library gets bigger. You might think that if you give the translators more text to read, they would agree better. But the opposite happened. As the amount of text increased, the translators disagreed more about where the "brightest spot" was. With more text, there are more places for the translators to pick different spots, making the problem worse, not better.
Why This Changes Everything
The author checked five major papers that had already published results using this method. They found that none of them reported where they measured the effect. They just said, "We turned off the switch at the brightest spot."
This is like a scientist saying, "We tested the medicine at the perfect temperature," without ever telling you what that temperature was. If two scientists test the same medicine but one uses 98°F and the other uses 102°F, they might get different results and think the medicine is unreliable. But really, they just measured at different temperatures.
The paper concludes that for a result to be trustworthy and comparable across different studies, scientists must report exactly which word (token) they used to turn the switch off. They also need to report how they handled special cases, like the very first word of a sentence, which can mess up the math if not treated carefully.
The Fix is Simple
The good news is that fixing this doesn't require building new libraries or training new translators. The author says the fix is just one line of code. Instead of letting the translator decide where to measure, researchers should agree on a shared list of spots to measure at, or at least report exactly which spot they chose.
By doing this, the "noise" disappears, and we can finally see which switches in the library's brain are truly important, and which ones are just bright because of where we happened to look. The paper doesn't claim to have solved all the mysteries of AI, but it has removed a giant, invisible wall that was blocking us from seeing the truth clearly.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.