Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects
This paper introduces Feature-Effect Geometry Analysis (FEGA) to demonstrate that while sparse autoencoder features can be causally relevant, they rarely function as stable steering directions, instead exhibiting distinct geometric patterns where "value-like" features produce structured effects and "pointer-like" features generate diffuse, context-dependent changes.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to understand a giant, magical library where books write themselves. This library is a "Large Language Model," a type of artificial intelligence that has read almost everything on the internet and learned to predict the next word in a sentence. Inside this library, there are millions of tiny, invisible switches called "neurons" that light up when the model thinks about something. For a long time, scientists tried to understand the model by just watching which switches lit up. They thought, "If a switch lights up when the model talks about 'Paris,' then that switch must be the 'Paris switch.'"
But here's the tricky part: just because a switch lights up doesn't mean it's actually doing the work. It might just be watching. To really know what a switch does, you have to reach in, flip it off, and see if the story changes. This is called "intervention." Recently, scientists built a special tool called a "Sparse Autoencoder" (or SAE) to help them find these switches. Think of an SAE as a super-organized filing cabinet that sorts the messy, glowing switches into neat, labeled folders. The hope was that if we find a folder labeled "Paris," we could just turn that folder on or off to make the model talk about Paris or stop talking about it. This idea is called "steering."
However, there was a nagging doubt. Sometimes, when scientists turned a "Paris" folder on, the model didn't talk about Paris at all. Sometimes it talked about France, sometimes it got confused, and sometimes it did the exact opposite of what was expected. It was like having a remote control where the "Volume Up" button sometimes turned the TV off, or changed the channel instead. The big question was: Why does this happen? Is the remote broken, or is our understanding of how the buttons work wrong?
This paper, titled "Sparse Autoencoders Encode Both Concepts and Functions," goes on a detective mission to solve this mystery. The authors, a team of researchers from Germany and India, decided to stop just looking at the labels on the folders and start watching what happens to the story after they flip a switch. They invented a new method called "Feature-Effect Geometry Analysis" (FEGA). Imagine FEGA as a high-tech camera that takes a snapshot of the library every time they flip a switch, not just once, but in thousands of different stories. They then look at the "cloud" of all these snapshots to see if the switch always pushes the story in the same direction.
What they found was a surprise. They discovered that the switches inside the model fall into two very different categories, which they named "Value-like" and "Pointer-like."
First, there are the Value-like switches. These are like the "facts" in the library. If a switch is tied to the fact that "Paris is in France," it acts like a static piece of information. When you flip this switch, the model's output changes in a somewhat organized way, like a group of people walking in a loose formation. It's not a single straight line, but it's not a total mess either. These switches are reliable for holding specific facts.
Then, there are the Pointer-like switches. These are the real troublemakers for the "remote control" idea. These switches don't hold a specific fact; instead, they hold a job or a function. Think of them like a "Copy" button on a keyboard. A "Copy" button doesn't care what you are copying; it just does the copying. In the model, a pointer-like switch might be responsible for "copying the word from earlier in the sentence." If the sentence says "The cat is big," the switch copies "big." If the sentence says "The cat is small," the switch copies "small." Because the job is to copy whatever is there, flipping this switch doesn't push the story in one single direction. Instead, the effect scatters everywhere, like a cloud of dust.
The paper shows that for these "Pointer-like" switches, the old idea of "steering" is mostly wrong. You can't just turn a "Copy" switch on and expect the model to always say "Copy." The effect depends entirely on what the model is looking at at that exact moment. The authors found that while some switches (the Value-like ones) have a somewhat structured effect, most of the interesting, functional switches (the Pointer-like ones) create a "diffuse" effect. They are essential for the model to work—they are the reason the model can follow rules or copy words—but they don't point in a single, predictable direction.
In short, the paper suggests that we need to change how we think about these AI switches. We can't treat them all like simple on/off buttons for specific topics. Some are like filing cabinets for facts, but others are like dynamic tools that change their effect based on the situation. If we want to control these AI models in the future, we can't just look for a single "Paris" button or a single "Copy" button and hope it works the same way every time. We have to understand that for many of these tools, the effect is a complex cloud that changes shape depending on the story being told. The authors don't claim to have fixed the problem yet, but they have provided a new map that shows us exactly where the confusion lies, proving that the "remote control" for AI is much more complicated than we thought.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.