← Latest papers
💬 NLP

Beyond Memorization: Assessing Semantic Generalization in Large Language Models Using Phrasal Constructions

This paper introduces a novel evaluation dataset based on Construction Grammar to assess the semantic generalization capabilities of Large Language Models, revealing that even state-of-the-art models struggle to distinguish between syntactically identical phrasal constructions with divergent meanings, exhibiting a performance drop of over 40% compared to human intuition.

Original authors: Wesley Scivetti, Melissa Torgbi, Austin Blodgett, Mollie Shichman, Taylor Hudson, Claire Bonial, Harish Tayyar Madabushi

Published 2026-06-01
📖 5 min read🧠 Deep dive

Original authors: Wesley Scivetti, Melissa Torgbi, Austin Blodgett, Mollie Shichman, Taylor Hudson, Claire Bonial, Harish Tayyar Madabushi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot to understand human language. You give it a library containing almost every book ever written on the internet. Naturally, the robot becomes incredibly good at reading and reciting what it has seen. But does it actually understand the rules of language, or is it just a super-powered parrot that memorized patterns?

This paper, titled "Beyond Memorization," asks that exact question. The researchers built a special test to see if Large Language Models (LLMs) can do something humans do effortlessly: understand the "skeleton" of a sentence even when the "flesh" (the specific words) is totally new or tricky.

Here is the breakdown of their study using simple analogies.

The Core Idea: The "Sentence Skeleton"

In linguistics, there is a concept called Construction Grammar. Think of a sentence not just as a string of words, but as a mold or a cookie cutter.

  • The Mold: The structure of the sentence (e.g., "Subject + Verb + Object + Adjective").
  • The Dough: The specific words you put into that mold.

Humans are great at this. If you know the mold for "making something flat by hitting it" (like hammering the metal flat), you can instantly understand a weird new sentence like brushing the hair smooth, even if you've never heard that specific combination before. You know the "skeleton" implies that the action caused the result.

The Experiment: Two Levels of Difficulty

The researchers created two tests (Experiments) to see if the AI could handle these molds.

Experiment 1: The "Creative Cookie" Test

The Goal: Can the AI understand a familiar mold when filled with unusual ingredients?

  • The Setup: They took common sentence molds (like the "Resultative" mold, where an action causes a change of state) and filled them with words that are rarely used in that specific mold.
  • The Analogy: Imagine a recipe for "Chocolate Cake." The AI has eaten millions of chocolate cakes. Now, they give it a recipe for "Savory Broccoli Cake."
  • The Result: The AI did very well. Even the smartest models (like GPT-4o and GPT-o1) understood that the action caused the result, even with the weird words. They successfully generalized the rule.

Experiment 2: The "Look-Alike Trap" Test

The Goal: Can the AI tell the difference between two molds that look identical but mean totally different things?

  • The Setup: This is where it gets tricky. The researchers found two sentence structures that look exactly the same on paper (same Subject, Verb, Object, Adjective) but have opposite meanings.
    • Sentence A (Resultative): "He hammered the metal flat." (The hammering caused the metal to become flat).
    • Sentence B (Depictive): "He ate the apple fresh." (The apple was already fresh when he ate it; eating didn't make it fresh).
  • The Analogy: Imagine two identical-looking boxes.
    • Box 1 says: "I put the toy inside the box." (The action of putting caused the toy to be inside).
    • Box 2 says: "I carried the toy inside the box." (The toy was already inside when I carried it).
    • To a human, the context tells you which box is which. To the AI, they look like the same box.
  • The Result: The AI failed miserably.
    • When asked to distinguish between these "look-alike" sentences, the performance of even the best models dropped by over 40%.
    • The models kept applying the "caused" meaning to the "already existing" sentences. They couldn't tell that the action wasn't the cause in the second example.

What This Means

The paper concludes that while AI is getting very good at recognizing patterns and applying rules to new words (Experiment 1), it struggles when it has to balance the structure of the sentence with the specific meaning of the words to figure out the cause-and-effect relationship (Experiment 2).

  • Humans use a mix of grammar rules and real-world knowledge (e.g., "You can't make an apple fresh by eating it") to solve the puzzle.
  • AI seems to rely too heavily on the most common pattern it has seen before. When faced with a "look-alike" trap, it defaults to the most frequent meaning it knows, rather than analyzing the specific logic of the new sentence.

The Bottom Line

The researchers built a dataset (a set of test questions) to prove that current AI models are still "memorizing" more than they are "reasoning." They can handle creative new words, but they stumble when the sentence structure tricks them into applying the wrong logic.

The paper makes it clear: We cannot assume that because an AI speaks fluently, it understands the deep, causal logic of language the way humans do. It's a parrot that has learned to sing very well, but it still gets confused when the lyrics change the meaning of the song.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →