← Latest papers
💬 NLP

Anything Goes? A Crosslinguistic Study of (Im)possible Language Learning in LMs

This crosslinguistic study demonstrates that while GPT-2 small exhibits some human-like inductive biases by distinguishing attested languages from typologically impossible ones, these biases are imperfect and weaker than those found in human learners.

Original authors: Xiulin Yang, Tatsuya Aoyama, Yuekun Yao, Ethan Gotlieb Wilcox

Published 2026-06-16
📖 5 min read🧠 Deep dive

Original authors: Xiulin Yang, Tatsuya Aoyama, Yuekun Yao, Ethan Gotlieb Wilcox

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart, very fast robot that learns to speak by reading millions of books. Scientists have been arguing about whether this robot actually "understands" language the way humans do, or if it's just a statistical parrot that can mimic anything it sees, even nonsense.

Some critics say the robot is too flexible: "It can learn gibberish just as easily as real English, so it doesn't prove anything about how human brains work."

This paper puts that idea to the test. The researchers asked: Can this robot tell the difference between a real language and a made-up, impossible one?

Here is a breakdown of their experiments and what they found, using some everyday analogies.

The Setup: The "Language Gym"

The researchers built a gym for their robot (a small AI model called GPT-2). They didn't just use English; they used 12 different languages from around the world (like Turkish, Chinese, German, and Arabic) to make sure the results weren't just a fluke of one specific language.

They created two types of "workouts" for the robot:

  1. Real Languages: Normal sentences spoken by humans.
  2. Impossible Languages: These were real sentences that had been scrambled. Imagine taking a sentence like "The cat sat on the mat" and shuffling the words until it read "mat on the sat cat the." Or reversing every word. These are "impossible" because no human brain would ever naturally learn or speak them.

They also tested "Unattested" Languages. These are languages that could exist logically but just happen to never be used by humans. It's like a puzzle piece that fits the picture but nobody has ever put it in the frame.

Experiment 1: The "Scrambled Egg" Test

The Question: If you feed the robot a scrambled version of a language, does it realize it's broken?

The Analogy: Imagine you are teaching a dog to fetch a ball. If you throw the ball, the dog fetches it. If you throw a rock, the dog might still fetch it, but maybe not as eagerly. The researchers wanted to see if the robot "prefers" the ball (real language) over the rock (scrambled language).

The Result:

  • Mostly Yes: In almost every language, the robot learned the real sentences much faster and better than the scrambled ones. It got "confused" (measured by a score called perplexity) when the words were in the wrong order.
  • The Catch: The robot wasn't perfect. Sometimes, if the scrambling was only a little bit (like swapping just two words), the robot didn't notice the difference as clearly as a human would. It was a bit like a person who can tell a song is off-key, but might miss it if only one note is slightly wrong.

Experiment 2: The "Global Comparison"

The Question: If we mix all the real languages and all the scrambled languages together, can the robot sort them into two neat piles?

The Analogy: Imagine a librarian trying to sort a pile of real books from a pile of books where the pages have been torn out and glued back together randomly. Can the librarian tell which pile is which just by looking at the books?

The Result:

  • Not Perfectly: The robot could tell the difference most of the time, but not 100%. Some of the scrambled languages looked so much like real ones that the robot got them mixed up.
  • The Takeaway: The robot has a "soft" preference for real languages, but it doesn't have the strict, hard-wired rules that human brains seem to have. It's a bit like a music fan who prefers real songs over noise, but will still enjoy a weird remix if it's catchy enough.

Experiment 3: The "Grammar Puzzle"

The Question: Humans have a natural bias for certain word orders. For example, in English, we say "The red car" (Adjective-Noun). In some other languages, they say "Car red." But there are combinations humans never use, like "Red car the." Can the robot learn these weird, unused patterns just as well as the normal ones?

The Analogy: Imagine teaching a child to stack blocks. Humans naturally prefer stacking them in a stable way (big block on bottom, small on top). If you try to teach them to stack them upside down, they struggle. The researchers wanted to see if the robot would struggle with the "upside-down" stacking or if it would just do it without complaining.

The Result:

  • The Confusing Part: When they just looked at how fast the robot learned (perplexity), it seemed to learn the weird, unused word orders just as easily as the normal ones. It was like the robot didn't care about the "rules" of stability.
  • The Reveal: However, when they tested the robot's ability to generalize (apply what it learned to new situations), a different picture emerged. The robot actually did better with the word orders that humans naturally use. It showed a subtle, human-like bias, but it was much weaker than what we see in actual people.

The Final Verdict: "Anything Goes?"

The paper concludes that the robot is not a "blank slate" that can learn anything equally well.

  • The "Anything Goes" Myth: The idea that the robot can learn gibberish just as easily as real language is false. The robot does show some "human-like" instincts; it prefers real languages and struggles slightly more with impossible ones.
  • The Reality: However, these instincts are weak. They aren't as strong or rigid as the rules inside a human brain. The robot is like a very talented student who has studied the rules of grammar but doesn't have the innate "feel" for language that humans are born with.

In short: The robot isn't magic, and it's not a perfect copy of a human mind. It sits somewhere in the middle: it has learned some of the "rules" of human language, but it's still missing the deep, intuitive understanding that makes human learning so unique.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →