Measuring Form and Function in Language Models
This paper introduces the Contextual Alternative Choice (CAC) metric to quantitatively evaluate language models on the formal syntactic and functional discourse properties of English determiners, revealing that while some very large models can match human children's performance on both benchmarks, no current model trained on comparable data achieves this simultaneously.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot to speak like a human. You have two ways to check if it's working:
- The Grammar Test: Does it say "The cat" instead of "Cat the"? (Correct form)
- The Conversation Test: Does it know when to say "The cat" versus "A cat" based on what the other person just said? (Correct function)
This paper is about building a better test for Language Models (LLMs) by looking at how real children learn to talk. The authors, researchers from the University of Pennsylvania, argue that we've been testing AI too much like a grammar quiz and not enough like a real conversation.
Here is the breakdown of their work using simple analogies:
1. The Problem: The "Multiple Choice" Trap
Most AI tests today are like multiple-choice exams. You show the AI two sentences:
- "The dog barked."
- "Dog the barked."
And ask, "Which one is right?"
The AI usually gets this right. But the authors say this is like testing a child by asking them to point to a picture of a dog versus a cat. It doesn't prove they understand how to use the word in a real story. Real children don't take multiple-choice tests; they just talk. To truly know if an AI is "learning" like a human, we need to see how it handles the messy, back-and-forth flow of real conversation.
2. The New Tool: "Contextual Alternative Choice" (CAC)
The authors invented a new way to test AI called CAC. Think of it like a "fill-in-the-blank" game, but with a twist.
- The Setup: They take a real recording of a mother talking to her 2-year-old child.
- The Game: They hide the word the child used (usually a small word like "the" or "a") and ask the AI: "Given what the mom just said, which word should go here?"
- The Goal: They don't just want the AI to pick the right word. They want to see if the AI's pattern of choices matches the child's pattern.
3. The Two Benchmarks (The Scorecard)
The authors created two specific scorecards to see if the AI is acting like a human child.
Scorecard A: The "Mix-and-Match" Test (Formal Productivity)
- The Concept: A smart child knows that "the" and "a" can go with almost any noun. They aren't just memorizing "the dog" and "a cat." They know the rule.
- The Test: The researchers look at how many different nouns the AI uses with "the" and "a." If the AI only uses "the" with specific words it memorized, it fails. If it mixes them up freely with many different words, it passes.
- The Result: Many AI models passed this. They learned the "mix-and-match" rule.
Scorecard B: The "Conversation Flow" Test (Discourse Function)
- The Concept: This is the hard part. In a real conversation, if a mom says, "Look at the dog," the child usually replies, "The dog is big." But if the mom says, "I see a dog," the child might say, "A dog is running."
- Sometimes you keep the same word ("the" -> "the").
- Sometimes you switch ("a" -> "the").
- The switch depends entirely on the context of the story.
- The Test: The researchers measured how often the AI switched words compared to how often real humans switch words.
- The Result: This is where most AIs failed. Even the big, smart models tended to switch words randomly, like a coin flip. They didn't seem to understand the story behind the words. They knew the grammar, but they didn't know the social rules of conversation.
4. The Big Surprise
The authors tested 45 different AI models.
- The "Small" Models: The models trained on about the same amount of data as a 2-year-old child (10–20 million words) generally failed both tests. They hadn't learned enough.
- The "Big" Models: The massive models trained on billions of words (like the size of the entire internet) did much better.
- Two models passed both tests. They sounded like humans in both grammar and conversation flow.
- One model passed the conversation test but failed the grammar test.
- Many others passed the grammar test but failed the conversation test.
5. The Takeaway
The paper concludes that size matters, but it's not everything.
Just because an AI can get a grammar quiz right doesn't mean it understands language the way a human does. It might just be memorizing patterns. To truly be a "cognitive model" of human learning, an AI needs to pass the "Conversation Flow" test, showing it understands why we choose certain words based on the story we are telling.
Currently, only the very largest models are starting to show this human-like understanding, but even they are a mixed bag. The authors suggest that if we want to build better AI, we should stop treating them like students taking a test and start treating them like toddlers in a conversation.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.