← Latest papers
💬 NLP

Language Models Can Resolve Reference Compositionally, But It's Not Their Native Strength: The Case of the Personal Relation Task

This paper demonstrates that while Large Language Models excel at the structured, intensional interpretation of personal relations compared to humans, they underperform in the referential, extensional task, suggesting that a lack of referential grounding is a critical limitation in their ability to achieve human-like language understanding.

Original authors: Bart Evelo, Meaghan Fowlie, Denis Paperno

Published 2026-06-01
📖 4 min read☕ Coffee break read

Original authors: Bart Evelo, Meaghan Fowlie, Denis Paperno

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a giant, complex family tree of six people: Amber, Bryan, Christina, and Dana. They are all connected by relationships like "friend," "enemy," "parent," and "child."

Now, imagine someone asks you a tricky question about this family tree, like: "Who is the friend of Amber's parent?"

This paper is a head-to-head competition between humans and AI language models (LLMs) to see who is better at answering this question. But here's the twist: the researchers asked the participants to solve the problem in two completely different ways.

The Two Ways to Solve the Puzzle

  1. The "Who Is It?" Game (Extensional Task):

    • The Goal: Just give the name of the person.
    • The Answer: "Felicia."
    • What it tests: Can you navigate the map and find the specific destination? This is like looking at a map and pointing to the city you need to visit.
  2. The "How Do I Get There?" Game (Intensional Task):

    • The Goal: Don't give the name. Instead, write down the recipe or the formula to find the person.
    • The Answer: friend(parent(Amber))
    • What it tests: Can you understand the structure of the sentence and write down the steps? This is like writing out the driving directions without actually driving the car.

The Big Surprise: Humans and AI are Opposites

The researchers found that humans and AI are good at completely different things. It's like comparing a human navigator to a super-fast translator.

  • Humans are the Navigators:
    When asked "Who is it?", humans were great at it (83% accuracy). They could look at the family tree, follow the connections, and point to the right person.
    However, when asked to write the "recipe" (friend(parent(Amber))), humans struggled (71% accuracy). It's like being a great driver who suddenly has to write a textbook on how to drive but forgets the specific rules for writing them down.

  • AI is the Super-Translator:
    When asked to write the "recipe," the AI was amazing (95% accuracy). It treated the question like a language puzzle. It saw "Amber's parent's friend" and simply rearranged the words into a formula. It's like a translator who is perfect at converting French to English but has never actually been to France.
    However, when asked "Who is it?", the AI got confused (80% accuracy). It struggled to actually use the family tree to find the specific person. It knew the recipe, but it couldn't follow it to the destination.

Why Does This Happen?

The paper suggests a simple reason: Where they learned their skills.

  • Humans learn language by interacting with the real world. We know what a "friend" or a "parent" is because we have met them. We are grounded in reality. So, finding a specific person (the "who") comes naturally to us.
  • AI learns only from text. It has never met a person named Amber. It has only seen millions of sentences about friends and parents. It is an expert at spotting patterns in words and rearranging them (translation), but it doesn't have a "real world" map to navigate. It's like a chef who has read every cookbook in the world but has never actually cooked a meal or tasted the ingredients.

The "Abstract" Twist

The researchers also tried a harder version where they replaced names like "Amber" with random letters like "h" and "friend" with "x."

  • Humans hated this. It was very hard for them to figure out the recipe with random letters.
  • AI was still okay at the recipe part, but it got very bad at finding the person.

The Takeaway

The paper concludes that while AI is incredibly smart at manipulating symbols and writing formulas (the "Intensional" task), it is not yet as good as humans at using those symbols to find real-world answers (the "Extensional" task).

The authors argue that for AI to truly understand language like a human, it needs more than just reading books; it needs to be "grounded" in the real world, just like we are. Until then, AI is a brilliant translator of ideas, but not yet a perfect navigator of reality.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →