Causal Evidence of Stack Representations in Modeling Counter Languages Using Transformers
This paper provides causal evidence that transformers trained on counter languages learn and rely on a specific stack representation, as ablating the identified principal direction causes sequential accuracy to collapse to near zero.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very smart robot that has learned to play a complex game of "matching parentheses." The game involves strings of opening and closing brackets, like ((())) or mixed-up versions of them. To win, the robot needs to know exactly how many open brackets are waiting to be closed at any given moment. If it loses count, it makes a mistake.
Scientists have long suspected that when these robots (called Transformers) learn this game, they secretly build a mental "stack" in their brains—a mental notepad where they write down the current count of open brackets. But here's the big question: Is this mental notepad just a side effect of learning, or is it actually the engine that makes the robot work?
This paper says: It's the engine. The stack isn't just a habit; it's the critical part of the robot's brain that makes the game possible.
Here is how they proved it, using a simple story:
1. The Setup: Teaching the Robot
The researchers taught a small robot to predict the next move in these bracket games. They used different levels of difficulty (from simple pairs to complex mixes of up to 8 different types of brackets). The robot got really good at it, eventually getting a perfect score.
2. The Detective Work: Finding the "Stack"
First, the researchers wanted to see if the robot actually had a "stack" in its brain. They built a simple tool called a probe (think of it like a metal detector). They pointed this detector at the robot's internal thoughts (its hidden states) at every step of the game.
The metal detector worked! It could tell exactly how many open brackets were currently "in the air" just by looking at the robot's brain activity. This confirmed that the robot had learned to represent the stack depth.
3. The Big Experiment: The "Brain Surgery"
Knowing the robot had this stack representation was good, but it didn't prove the stack was necessary. Maybe the robot was using a secret backup plan, and the stack was just a decoration.
To find out, the researchers performed a tiny, precise "brain surgery" on the robot while it was playing the game.
- The Target: They identified the specific direction in the robot's brain that held the "stack count" information.
- The Action: They effectively "turned off" or erased that specific piece of information from the robot's brain at every step of the game.
- The Control: They also tried erasing random, meaningless directions in the brain to make sure they weren't just breaking the robot by accident.
4. The Result: The Robot Crashes
The result was dramatic.
- When they erased random information, the robot kept playing perfectly.
- When they erased the stack information, the robot's performance didn't just get slightly worse; it collapsed completely.
The robot went from getting 100% of the answers right to getting almost 0% right. It was as if you took the steering wheel out of a car while it was driving; the car didn't just drive slowly, it stopped working entirely.
The Conclusion
This experiment provides strong evidence that the "stack" isn't just a side effect. It is causally necessary. The robot needs that specific mental stack to solve the puzzle. If you remove the stack, the robot loses the ability to understand the language.
In short: The researchers didn't just find a map of how the robot thinks; they proved that the map is the actual terrain. Without the stack representation, the robot's brain literally cannot function for this task.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.