Polysemanticity or Polysemy? Lexical Identity Confounds Superposition Metrics
This paper demonstrates that a significant portion of apparent neural superposition in language models is actually caused by a lexical confound where neurons activate for shared word forms like "bank" rather than compressed unrelated concepts, a finding that, when corrected, improves word sense disambiguation and the selectivity of knowledge edits.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: The "Bank" Confusion
Imagine you are a detective trying to figure out how a super-smart robot (a Large Language Model) thinks. You notice that one specific "lightbulb" (a neuron) inside the robot's brain lights up whenever the robot reads the word "Bank."
But here's the catch: The robot lights up this same bulb for two very different meanings of "Bank":
- Financial Bank: Where you keep your money.
- River Bank: The muddy edge of a river.
The Old Theory (Superposition):
For a long time, researchers thought this meant the robot was being incredibly efficient. They believed the robot was "compressing" information. Since the robot doesn't have enough lightbulbs to give every single concept its own unique switch, it was forced to stack unrelated ideas (money and rivers) into the same switch to save space. This is called Superposition. It's like trying to fit a sofa and a potted plant into a tiny closet; you have to shove them in together.
The New Discovery (Lexical Confound):
This paper argues that the robot isn't actually compressing unrelated ideas. Instead, it's just reacting to the spelling of the word.
Think of it like a security guard at a club.
- The Old View: The guard lets in both "Rock Stars" and "Astronauts" because he's trying to save space in the VIP list.
- The New View: The guard lets them in simply because they both have the same name tag that says "Bank." He hasn't even looked at what they are doing (saving money vs. fishing) yet. He is reacting to the word form, not the concept.
The authors call this a Lexical Confound. The robot is confusing the word with the meaning.
The Experiment: The "2x2" Test
To prove this, the researchers set up a clever game with four scenarios, like a chef testing ingredients:
- Same Word, Same Meaning: "Bank" (money) vs. "Bank" (money). Result: The lightbulb goes crazy. (Expected).
- Different Word, Same Meaning: "Bank" (money) vs. "Credit Union" (money). Result: The lightbulb lights up a little. (This is true semantic overlap).
- Same Word, Different Meaning: "Bank" (money) vs. "Bank" (river). Result: The lightbulb goes crazy.
- Different Word, Different Meaning: "Bank" (money) vs. "Apple." Result: The lightbulb stays off.
The Shocking Result:
The lightbulb lit up much more for Scenario 3 (Same Word, Different Meaning) than for Scenario 2 (Different Word, Same Meaning).
This proves that the robot is mostly reacting to the letters "B-A-N-K," not the idea of money or rivers. The "compression" we thought we saw was actually just the robot recognizing the word itself before it figured out what the word meant.
The "Sense-Blind" Neurons
The paper finds that a huge chunk of the neurons we thought were "super-efficient" (polysemantic) are actually just Sense-Blind.
- Sense-Selective Neurons: These are the smart ones. They know the difference between a river and a bank account.
- Sense-Blind Neurons: These are the "Word Form Detectors." They fire for any meaning of the word "Bank," but they don't care which one it is. They are like a motion sensor that goes off whenever a human walks in, regardless of whether the human is holding a gun or a bouquet of flowers.
The researchers found that 18% to 36% of the features in these models are actually just these "Sense-Blind" detectors. They aren't compressing concepts; they are just recognizing the word.
Why Does This Matter? (The Real-World Impact)
If we think these neurons are compressing concepts, we might try to fix them or edit the robot's brain in the wrong way.
Analogy: The Surgery
Imagine you want to teach the robot that "Bank" (river) is dangerous, but you want to keep "Bank" (money) safe.
- The Old Way (Standard Editing): You cut out the "Bank" neuron. Disaster! Now the robot forgets both meanings. It can't talk about money or rivers anymore.
- The New Way (Sense-Selective Editing): You realize the "Bank" neuron is just a word detector. You ignore it and only cut the specific "River" neuron. Success! The robot still knows about money, but now it knows rivers are dangerous.
The paper shows that if you filter out these "Sense-Blind" neurons, the robot becomes much better at understanding context and allows for much more precise editing of its knowledge.
The "Hidden Subspace" (The Secret Drawer)
The researchers also found that all these "Sense-Blind" neurons are hiding in a very small, specific corner of the robot's brain (a 20-dimensional subspace). It's like finding that all the "Word Form Detectors" are stored in a single, tiny drawer in a massive library.
Because they are in such a small, organized space, we can actually remove this drawer entirely without breaking the robot. If you take out this "Word Form" drawer, the robot stops confusing the word with the meaning, and the "compression" metrics drop significantly.
Summary in One Sentence
We thought AI models were geniusly compressing unrelated ideas into single neurons to save space, but they were actually just reacting to the spelling of the words; once we separate the "spelling detectors" from the "meaning detectors," the models become clearer, more accurate, and easier to edit.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.