Do Language Models Know What Not to Say? Causal Evidence for Statistical Preemption in LLMs
This paper provides the first causal evidence that large language models acquire knowledge of linguistic unacceptability through statistical preemption, demonstrating that exposure to conventional forms suppresses structurally possible but unattested alternatives via distributional competition rather than mere verb entrenchment.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Question: How Do We Learn What Not to Say?
Imagine a child learning to speak. They hear their parents say, "I donated the books to the library." Eventually, the child figures out that saying, "*I donated the library the books," sounds wrong.
Here is the puzzle: The child has never been told, "Don't say it that way." They have never heard a correction. In fact, the sentence structure they are avoiding (the "double-object" version) works perfectly fine for other words, like "give" ("I gave the library the books").
So, how does the brain know to stop using a pattern that could work, just because it hasn't heard it used for this specific word? This is known as Baker's Paradox.
The Theory: The "Crowded Room" Analogy
The paper tests a theory called Statistical Preemption.
Imagine a crowded room where everyone is trying to get a message across.
- The Theory: If you want to tell someone you are giving them something, and you see that 99% of people in the room say, "I gave the book to the person," your brain learns that this is the "standard" way. Even though you could physically say, "I gave the person the book," your brain realizes, "Oh, everyone else uses the 'to' version for this specific action. If I use the other version, people might think I'm weird or wrong."
- The Competing Idea (Entrenchment): A rival theory suggests that you just stop using new patterns because you've heard the word "donate" so many times in any context. It's like saying, "I've heard 'donate' so much that I'm too set in my ways to try anything new."
The authors wanted to know: Do AI Language Models (LLMs) learn like humans? Do they figure out what not to say just by noticing which version is the "crowded room" favorite?
The Experiment: The AI as a Super-Learner
The researchers treated Large Language Models (like GPT-2, LLaMA, and Pythia) as "super-learners." These models read billions of sentences but were never taught grammar rules. They just absorbed statistics.
They tested the models on 120 different verbs (like donate, give, explain, pour) across three types of sentence structures.
1. The "Surprise" Test (Correlation)
The researchers measured how "surprised" the AI was by different sentence structures.
- The Metaphor: Imagine the AI is a DJ playing music. If a song fits the vibe perfectly, the crowd (the AI) is calm. If a song is weird and out of place, the crowd gets "surprised" (the AI's internal alarm goes off).
- The Result: The AI was highly "surprised" by the unnatural sentences (like "*donated the library the books") and calm about the natural ones. This pattern matched human judgments almost perfectly. When humans said a sentence was weird, the AI was also "surprised" by it.
2. The "Crowded Room" vs. "Famous Name" Test (Dissociation)
This was the most critical part. They needed to prove the AI wasn't just avoiding the sentence because the word was famous (Entrenchment), but because a specific competitor was winning (Preemption).
- The Setup: They picked two groups of verbs.
- Group A (The "Crowded Room"): Verbs where one specific sentence structure dominates (e.g., donate is almost always used with "to").
- Group B (The "Famous Name"): Verbs that are used very often, but in many different ways, with no single "winner."
- The Result: The AI only avoided the weird sentence structure for Group A. For Group B, even though the words were famous, the AI didn't mind trying the weird structure.
- The Conclusion: The AI isn't just avoiding things because a word is common. It's avoiding them because it has learned that another specific way of saying it is the standard. This proves the "Statistical Preemption" theory is correct.
3. The "Size Matters" Test (Scaling)
They tested models of different sizes, from small (like a pocket calculator) to huge (like a supercomputer).
- The Metaphor: It's like training a dog. A puppy might make mistakes, but a highly trained dog knows the rules.
- The Result: As the AI models got bigger, their ability to spot the "wrong" sentences got better. However, it wasn't a magic switch that flipped on suddenly; it was a smooth, steady improvement, like a student getting better grades over time.
4. The "Rewiring" Test (Causal Proof)
To prove this wasn't just a coincidence, they did a "surgery" on a small AI model.
- The Action: They took a model and fed it extra training data.
- Scenario 1: They made the "correct" version of the sentence appear 3 times more often.
- Scenario 2: They made the "wrong" version appear 3 times more often.
- The Result:
- In Scenario 1, the AI became more strict about avoiding the wrong sentence.
- In Scenario 2, the AI became less strict, but not as much as Scenario 1 changed it.
- The Conclusion: This proved that the AI's behavior is directly caused by the frequency of the competing sentences. If you change the input statistics, you change the AI's "grammar" in a predictable way.
The Bottom Line
The paper concludes that neural language models learn "what not to say" in the exact same way humans do: through statistical competition.
They don't need a teacher to say, "That's wrong." They just need to hear the "right" way of saying things enough times that their brain (or code) realizes, "Oh, that's the standard way. The other way is blocked."
This solves a decades-old puzzle about how humans learn language without negative feedback, showing that distributional statistics (counting how often things happen) are powerful enough to teach us the rules of grammar.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.