← Latest papers
💬 NLP

Using Perspectival Words Is Harder Than Vocabulary Words for Humans and Even More So for Multimodal Language Models

This paper reveals that while multimodal language models approach human performance on basic vocabulary, they struggle significantly more than humans with perspectival words like possessives and demonstratives due to deficits in perspective-taking and spatial reasoning.

Original authors: Dota Tianai Dong, Yifan Luo, Po-Ya Angela Wang, Asli Ozyurek, Paula Rubio-Fernandez

Published 2026-04-21
📖 4 min read☕ Coffee break read

Original authors: Dota Tianai Dong, Yifan Luo, Po-Ya Angela Wang, Asli Ozyurek, Paula Rubio-Fernandez

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot how to talk to people. You've taught it a massive dictionary: it knows what a "boat," a "cup," and a "red" are. It's a vocabulary genius. But now, you want to teach it how to have a real conversation, where words change meaning depending on who is speaking, who is listening, and where things are located.

This paper is like a report card on how well modern AI (specifically Multimodal Language Models) handles these "perspective words" compared to real humans.

Here is the breakdown of their findings, using some everyday analogies:

1. The Three Levels of Difficulty

The researchers tested the AI and humans on three types of words, which they arranged like a ladder of difficulty:

  • Level 1: The Vocabulary Ladder (Easy)
    • The Task: Pointing at a picture of a boat and saying "Boat."
    • The Result: Both humans and AI are basically perfect here. It's like recognizing a face in a crowd. The AI knows what a boat is.
  • Level 2: The Ownership Ladder (Medium)
    • The Task: Two people are looking at a spoon. One says, "This is mine," and the other says, "That is yours."
    • The Result: Humans are great at this. The AI starts to stumble. It's like the AI gets confused about who actually owns the spoon. It struggles to understand that "mine" changes meaning depending on who is holding the pen.
  • Level 3: The Distance Ladder (Hard)
    • The Task: Two people are looking at two identical cups. One is close, one is far. One says, "I want this one," and the other says, "I want that one."
    • The Result: This is where the AI really falls apart. Humans find this tricky but manageable. The AI, however, gets lost in space. It doesn't fully grasp that "this" means "near me" and "that" means "far from me" in a way that shifts based on where I am standing.

2. The "Blindfolded" Analogy

The researchers did a cool experiment to see why the AI fails. They took away the clues the AI usually relies on.

  • The Human Way: Imagine you are in a room. You see a person, you see a cup, and you know where you are standing. You use your eyes and your brain together to figure out "this cup" vs. "that cup."
  • The AI Way: The AI is like a person who is blindfolded but has a very loud radio in their ear.
    • When the researchers removed the visual clues (the picture of the cup's location), the AI didn't panic; it just relied even more on the radio (the text instructions).
    • When they removed the text clues, the AI got confused.
    • The Discovery: Humans use their eyes and ears in sync. The AI relies heavily on the text instructions and treats the image almost like a secondary decoration. It doesn't naturally "feel" the distance in the picture the way a human does.

3. The "Teacher" vs. The "Student"

The researchers tried to help the AI by giving it "instruction-based prompting." This is like a teacher whispering, "Hey, remember! 'This' means the thing close to you, and 'That' means the thing far away."

  • For Ownership (Mine/Yours): The whisper worked! The AI's performance jumped up to almost human levels. It just needed a reminder of the rules.
  • For Distance (This/That): The whisper helped a little, but the AI still couldn't do it. It's as if the AI can memorize the rule "This = Near," but it can't actually visualize the concept of "near" in a dynamic scene. It lacks the "embodied" experience of moving around a room and seeing things get closer or farther away.

The Big Takeaway

The paper concludes that while AI is getting really good at knowing what things are (vocabulary), it is still very bad at understanding where things are relative to people (perspective).

  • Vocabulary is like learning the names of the players on a soccer team.
  • Perspective is like understanding the game itself: who is passing to whom, who is near the goal, and who is the referee.

Current AI models are great at memorizing the player names, but they are still struggling to understand the flow of the game. To truly communicate like humans, AI needs to learn how to "stand in someone else's shoes" and see the world from their specific spot in the room, not just read a description of the room.

In short: AI is a brilliant encyclopedia, but it's still a bit clumsy at having a casual chat about "this cup" and "that cup" across a table.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →