Differences in Typological Alignment in Language Models' Treatment of Differential Argument Marking
This paper demonstrates that while GPT-2 models trained on synthetic corpora successfully replicate the human-like preference for marking semantically atypical arguments in differential argument marking systems, they fail to reproduce the cross-linguistic tendency to preferentially mark objects over subjects, suggesting these typological patterns stem from distinct underlying sources.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a robot to speak a new language. But instead of teaching it a real human language like Spanish or Japanese, you invent 18 different "fake" languages. In each fake language, you have a specific rule about when to add a special "stamp" (a marker) to a word in a sentence.
This paper is about a team of researchers who taught AI models (specifically a version of GPT-2) these 18 fake languages to see if the AI would naturally figure out the rules the same way humans do, or if it would get confused.
Here is the breakdown of what they did and what they found, using some everyday analogies.
The Game: "The Stamp Rule"
In the real world, languages often have a system called Differential Argument Marking (DAM). This is a fancy way of saying: "We only put a special stamp on a word if it's doing something unusual."
The researchers created 18 versions of English where they added a special symbol (like a 🅿️ or an 🅰️) to sentences based on four different rules:
- The Trigger: What makes a word special? (Is it alive? Is it specific? Is it a pronoun?)
- The Target: Do we stamp the doer (Subject) or the receiver (Object)?
- The Direction: Do we stamp the "weird" cases or the "normal" cases?
- The Complexity: Do we look at just one word, or compare two words against each other?
Example:
- Normal English: "The dog chased the cat." (No stamps).
- Fake Rule: "If the thing being chased is alive, put a 🅿️ stamp on it."
- Result: "The dog chased the cat 🅿️." (Because a cat is alive).
- Result: "The dog chased the rock." (No stamp, because a rock isn't alive).
The Two Big Questions
The researchers wanted to see if the AI would learn two specific patterns that humans naturally follow in real languages:
- The "Weirdness" Rule (Markedness): Humans tend to stamp the unusual things. If 90% of the time you chase a rock, but 10% of the time you chase a cat, you might stamp the cat to say, "Hey, this is the special one!"
- The "Receiver" Rule (Object Preference): In human languages, we almost always stamp the receiver (the object) of an action, and rarely the doer (the subject).
What the AI Did (The Results)
1. The AI Got the "Weirdness" Rule Right 🎯
The AI was excellent at learning the Markedness rule.
- The Analogy: Imagine you are wearing a uniform. Usually, everyone wears a blue shirt. But if you are the only one wearing a red shirt, you get a special badge.
- The Result: The AI learned that if a rule says "Stamp the unusual thing," it did it perfectly. It figured out that if the "unusual" configuration happens less often, it needs a stamp to stand out. This matched human language patterns perfectly.
2. The AI Failed the "Receiver" Rule ❌
The AI did not learn the Object Preference rule.
- The Analogy: In human languages, it's like a rule that says, "We only put the badge on the person receiving the gift, never the person giving it."
- The Result: The AI didn't care who was giving or receiving. If the rule said "Stamp the giver," the AI learned that just fine. If the rule said "Stamp the receiver," it learned that too. It didn't show a strong preference for the receiver like humans do.
Why Did This Happen? (The "Why" Behind the "What")
The authors suggest that these two rules come from two different "engines" inside our brains (or in this case, the AI's training).
- The "Weirdness" Engine is Local: The AI is trained to predict the next word in a sentence. It's very good at noticing, "Hey, this word is rare here, so I should expect a stamp to help me predict what comes next." This is a local, mechanical skill that the AI is great at.
- The "Receiver" Engine is Global: The reason humans prefer stamping the receiver has to do with storytelling and conversation. Usually, the "doer" is the main character (the topic), so we don't need to mark them. The "receiver" is often the new information, so we mark them. The AI, however, is just looking at the next word; it doesn't really understand the "big picture" of a story or who the main character is. Because it lacks this "storytelling" context, it didn't develop the human preference for stamping the receiver.
The Takeaway
This paper shows that AI models are like very smart students who are great at spotting patterns and statistics (like "unusual things get stamped"), but they struggle with social and storytelling pressures (like "we usually focus on the person doing the action").
It suggests that the rules of human language aren't all made of the same stuff. Some rules are just about math and efficiency (which AI learns easily), while others are about how we tell stories and manage attention (which AI, at least in this setup, doesn't quite get yet).
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.