← Latest papers
💬 NLP

Priors Persist Through Suppression: A Stroop Paradigm for Lexical Override

This paper demonstrates that in language models, familiar lexical priors persist through explicit override rules rather than being replaced, causing Stroop-like interference that originates from and is mechanistically localized to the specific activation pathways binding definition targets to query words.

Original authors: Han-yu Wang

Published 2026-06-09
📖 5 min read🧠 Deep dive

Original authors: Han-yu Wang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a very smart, well-read robot a new rule for a game. You tell it: "In this game, the word 'doctor' means 'forest'."

You then ask the robot: "Using only this rule, what word is related to 'doctor'?"

Ideally, the robot should say "forest." But because the robot has read millions of books before, it has a deep, automatic habit: when it hears "doctor," it immediately thinks of "hospital."

This paper is a study of what happens when the robot tries to follow your new rule while fighting its old habit. The researchers call this a "Stroop-style" test (named after a famous psychology experiment where people struggle to name the color of a word when the word itself spells a different color, like the word "RED" written in blue ink).

Here is the breakdown of their findings, using simple analogies:

1. The Old Habit Doesn't Just Disappear

The researchers found that when you tell the robot to ignore its old knowledge, it doesn't simply delete the old meaning and replace it with the new one. Instead, the old meaning lingers in the background.

  • The Analogy: Imagine you are trying to drive a car on a new road, but your muscle memory keeps trying to steer you toward your old home. You can force the car to stay on the new road, but your hands are still twitching toward the old direction. The stronger your old habit (the "lexical prior"), the harder it is to stay on the new road, even if you are trying very hard.
  • The Finding: The more strongly the robot usually associates "doctor" with "hospital," the more it struggles to say "forest," even when you explicitly told it the new rule.

2. The "Glitch" Happens in a Specific Spot

The researchers didn't just watch the robot's answers; they looked inside its "brain" (its internal computer code) to see where the struggle happened. They used a technique called "activation patching," which is like temporarily swapping a piece of a robot's brain with a clean version from a different robot to see if it fixes the error.

  • The Analogy: Think of the robot's brain as a factory assembly line. The researchers found a specific three-person team on the line that is responsible for this specific task:
    1. The person who hears the new rule ("Doctor means...").
    2. The person who holds the new meaning ("...Forest").
    3. The person who hears the question word ("...Doctor?").
  • The Finding: When the researchers "patched" (replaced) the work of these three specific people with a clean version, the robot suddenly got the answer right. If they only fixed two of them, or if they swapped the "Forest" person with someone who was holding a different word, the robot still failed. This proves that the robot needs to bind the specific new meaning to the specific word to make it work.

3. The "Suppression" vs. "Preservation" Mystery

One of the most interesting discoveries was how the robot fixes the mistake. The researchers expected the robot to simply "turn down the volume" on the wrong answer ("hospital").

  • The Analogy: Imagine a radio with two stations playing at once: the old habit (Hospital) and the new rule (Forest).
    • What they thought would happen: The robot would just turn the volume down on the "Hospital" station.
    • What actually happened: The robot does turn down the volume on "Hospital," but it does that for almost any reason (even if the rule is slightly messed up). The real magic happens because the robot turns the volume UP on the correct answer ("Forest") specifically when the new rule is perfectly clear.
  • The Finding: The robot's ability to follow the new rule depends on preserving the new meaning, not just suppressing the old one. If the "new meaning" part of the brain gets corrupted, the robot collapses, even if the "old meaning" is still being suppressed.

4. It Happens in All Kinds of Robots

The researchers tested this on 11 different types of AI models (ranging from small to large, and from different companies).

  • The Finding: No matter how big or smart the robot was, or how they phrased the rule (as a game, a legal document, or a technical manual), the old habit always interfered. The stronger the robot's original habit, the more it struggled to follow the new rule.

Summary

The paper concludes that when an AI is asked to change the meaning of a word, it doesn't simply overwrite its old knowledge. Instead, it's a tug-of-war. The old habit keeps pulling, and the new rule has to fight to keep the answer correct. The AI succeeds not by erasing the past, but by successfully "binding" the new meaning to the word in a specific part of its brain, while simultaneously trying to quiet the old habit.

In short: You can tell a robot a new rule, but its old habits are still whispering in its ear, and the robot has to work hard to ignore them.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →