Function-Vector Heads Are Two Populations: Writers and Cancellers in In-Context Learning
This paper challenges the assumption that function-vector heads are a homogeneous group by revealing they consist of two opposing sub-populations—writers that promote correct rules and cancellers that suppress them—which are invisible to magnitude-only ranking but significantly impact model performance when identified and ablated.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a large language model (like the ones that chat with you) as a massive orchestra trying to solve a puzzle based on a few examples. In this orchestra, there are specific musicians called attention heads. For a long time, researchers believed that the most important musicians were simply the ones playing the loudest notes. They assumed that if a musician was playing very loudly, they were definitely helping the orchestra play the right tune.
This paper, however, discovers that this assumption is wrong. It turns out that the "loudest" musicians actually belong to two completely different, opposing groups: The Writers and The Cancellers.
Here is the breakdown of what the paper found, using simple analogies:
1. The Big Misunderstanding: Volume vs. Direction
Previously, researchers looked at these musicians and ranked them by volume (how much they contributed to the answer). They assumed the top 20 loudest musicians were all "helpers."
The authors of this paper realized that volume isn't enough; you need to know the direction.
- The Writers: These musicians play a loud note that pushes the orchestra toward the correct answer.
- The Cancellers: These musicians also play a loud note, but they are pushing the orchestra toward the wrong answer. They are actively trying to cancel out the correct answer.
The Analogy: Imagine a tug-of-war.
- Writers are the team pulling the rope toward the finish line.
- Cancellers are a second team pulling the rope just as hard, but in the opposite direction.
- If you only look at how hard they are pulling (magnitude), you think they are both equally important "pullers." But if you look at the direction, you see they are fighting each other.
2. The Discovery: Two Populations
The researchers tested this on several different models (ranging from small to large) and different types of logic puzzles. They found that in almost every case, the "top" musicians split into these two opposing camps.
- The "Writers" push the probability of the correct answer up.
- The "Cancellers" push the probability of the correct answer down.
When the researchers silenced (ablated) the Cancellers, the model actually got better at solving the puzzles. It was like removing a saboteur from the tug-of-war team; suddenly, the "Writers" could pull the rope to the finish line without resistance.
3. Why Previous Research Missed This
The paper explains that earlier studies (like one by Todd et al., 2024) used a method that only looked at the "loudness" of the contribution. Because of this, they accidentally picked up a mix of Writers and Cancellers, but the mix changed depending on the puzzle.
- On some puzzles, the "Cancellers" were louder, so the old method thought they were the main helpers.
- On other puzzles, the "Writers" were louder, so the old method thought they were the main helpers.
The paper shows that by ignoring the "sign" (direction), previous research was essentially mixing up the heroes and the villains, thinking they were all just "loud helpers."
4. Are the Cancellers Just Mistakes?
The researchers wanted to make sure the "Cancellers" weren't just broken parts of the model or accidental noise. They ran many tests to prove they are real, functional parts of the system:
- They read different things: Writers focus on the "labels" (the answers) in the examples. Cancellers focus on the "format" (the punctuation and spacing). They are looking at different parts of the page.
- They are content-driven: The Cancellers aren't just randomly suppressing answers; they are specifically trying to suppress the answer that was just demonstrated in the examples.
- They are not "Copy Suppressors": The paper ruled out the idea that these are just generic "don't copy" mechanisms. They are specific to the logic of the task.
5. The "L11.H4" Character Study
The paper zooms in on one specific musician, L11.H4, to see how it works.
- This musician is a "Canceller" on the logic puzzles (it tries to push the wrong answer).
- However, if you change the puzzle to a vocabulary task (like matching antonyms), this same musician flips its role and becomes a "Writer" (it helps the correct answer).
- The Lesson: This musician isn't inherently "good" or "bad." Its job is to suppress what was just demonstrated. If the demonstration is the right answer, it suppresses it (Canceller). If the demonstration is the wrong answer (in a different context), it suppresses the wrong answer, effectively helping the right one (Writer).
6. The Takeaway
The main conclusion is that we cannot treat the most "active" parts of an AI model as a single, uniform group.
- Old View: "These are the top 20 most important heads. Let's use them to steer the model."
- New View: "These top 20 heads are actually a civil war. Half are pushing the model to the right place, and half are pushing it to the wrong place. If you want to steer the model, you need to silence the Cancellers, not just amplify the Writers."
The paper does not claim this makes the model "safe" or ready for real-world medical or legal use. It simply states that to understand how these models learn rules from examples, we must recognize that their internal "musicians" are often fighting each other, not working together.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.