Large Language Models as Proxies for Theories of Human Linguistic Cognition
This paper explores the potential of using current large language models as proxies for linguistically-neutral theories of human cognition to evaluate linguistic pattern acquisition, while acknowledging that their utility in this role remains quite limited at present.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Question: Are AI Chatbots the Future of Understanding Human Brains?
Imagine linguists (scientists who study how humans learn language) are trying to solve a mystery: How does a human baby learn to speak so perfectly, even though they don't hear enough examples to figure out all the complex rules?
For decades, the leading theory (let's call it Theory A) says: "Humans are born with a special, built-in 'language chip' in their brains. This chip is biased toward specific language rules, like a puzzle piece that only fits one way."
Recently, a new idea has popped up (let's call it Theory B). It suggests: "Maybe we don't need a special chip. Maybe if you just feed a computer enough text (like a baby hears words), it will figure out the rules all by itself. Therefore, current AI chatbots (Large Language Models or LLMs) are actually the best explanation for how humans learn."
The authors of this paper say: "Hold on. That's not quite right."
They propose a middle-ground idea called the Proxy View. They suggest that while AI chatbots aren't perfect models of the human brain, they might act as a "test dummy" or a "proxy" for a third, unknown theory (Theory C). Theory C would be a learning system that is fair and neutral (no special "language chip") but still manages to learn language.
The paper asks: Can current AI chatbots act as a good "test dummy" to prove that a neutral learning system (Theory C) could actually work?
The Two Big Tests
To see if the AI can serve as a good "test dummy," the authors ran two types of experiments. Think of these as two different obstacle courses.
Test 1: The "Baby Book" Challenge (Learning from Limited Data)
The Analogy: Imagine trying to teach a robot to play chess by showing it only 10 games. A human child learns complex grammar from a relatively small amount of speech (about 10 years of talking).
The Experiment: The authors trained AI models on different amounts of text:
- Some models saw as much text as a 10-month-old baby.
- Some saw as much as an 8-year-old.
- Some saw as much as a human would hear in 800,000 years (a massive amount of data).
They tested these models on tricky grammar puzzles that humans solve easily, like:
- The "Who" Puzzle: Figuring out who did what in a sentence like "Which book did you say that Kim hated?"
- The "That" Puzzle: Knowing when to drop the word "that" in a sentence (e.g., "Who did you say loves Sue?" is okay, but "Who did you say that loves Sue?" sounds wrong to a native speaker).
The Result:
The AI models failed miserably.
- Even the tiny models (baby-level data) got it wrong.
- Crucially, even the giant models (trained on 800,000 years of data) still got it wrong. They often preferred the wrong version of the sentence.
What this means: If a computer with infinite time and data cannot figure out these rules, it's unlikely that a "neutral" human learning system (Theory C) could do it either. The AI failed to act as a good "test dummy" for this theory.
Test 2: The "Weird Language" Challenge (Learning the Unusual)
The Analogy: Imagine you have a robot that learns languages. You want to know if it naturally prefers "normal" languages (like English) over "weird" languages (like a language where every sentence is backwards).
The Experiment: The authors took real languages (English, Italian, Russian) and created "weird" versions of them using math tricks:
- Partial Reverse: Reversing half the sentence.
- Full Reverse: Reversing the whole sentence.
- Token Hop: Making the robot count words to know where to put a symbol.
They asked the AI: "Which version is easier to learn? The normal one or the weird one?"
The Result:
The AI didn't care about "weirdness."
- Sometimes, the AI found the weird, backwards languages easier to learn than the normal ones.
- Sometimes, it found the normal languages harder.
What this means: If the AI (our "test dummy") can learn "weird" languages just as easily as real ones, it suggests that a neutral learning system wouldn't naturally explain why humans only speak real languages and never speak backwards ones. The AI failed to show that "real" languages are naturally easier to learn.
The Conclusion: The "Test Dummy" is Broken
The authors conclude that the Proxy View (using AI as a stand-in for a neutral learning theory) is currently not working.
- The LLM Theory (AI = Human Brain): This is wrong because AI lacks the specific "human" quirks (like knowing the difference between a sentence that is grammatically correct but unlikely, versus one that is just wrong).
- The Proxy View (AI = Test Dummy for a Neutral Theory): This is also failing right now. The AI models are so bad at learning the specific rules humans learn easily, and so good at learning "weird" rules humans never use, that they cannot help us prove that a neutral learning system works.
The Final Takeaway:
The authors aren't saying AI is useless. They are saying that right now, AI is too different from the human brain to be used as a reliable "test dummy" for new theories.
They are essentially telling other scientists: "If you want to use AI to prove that humans learn language without a special 'language chip,' you need to build a much better AI or a much clearer theory first. The current chatbots aren't doing the job."
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.