Can LLMs interpret figurative language as humans do?: surface-level vs representational similarity
While large language models align with humans on the surface-level interpretation of dialogue, they significantly diverge at the representational level, particularly when processing context-dependent and socio-pragmatic figurative language such as idioms, slang, and sarcasm, with GPT-4 demonstrating the closest approximation to human patterns among the tested models.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot how to understand human jokes, sarcasm, and slang. You show it a sentence like, "This food is gas," and ask, "Is this funny?" or "Is this positive?"
A new study by researchers at Texas A&M University asks a tricky question: When the robot gives the same answer as a human, does it actually think the same way a human does?
Here is the breakdown of their findings, using some everyday analogies.
The Experiment: The "Taste Test"
The researchers set up a massive taste test, but instead of ice cream, they used 240 different sentences. These sentences covered six "flavors" of language:
- Conventional: Normal sentences (e.g., "The dog is outside").
- Idiomatic: Phrases where the words don't mean what they say (e.g., "Bite the bullet").
- Emotional: Sentences full of feelings.
- Funny: Jokes and puns.
- Sarcastic: Saying the opposite of what you mean.
- Gen Z Slang: Modern internet speak (e.g., "This food is gas").
They asked 211 humans and four different AI models (including GPT-4, Llama, and Mistral) to rate every sentence on 40 different questions, like "Is this positive?" or "Is this concerning?" on a scale of 1 to 10.
The Two Ways to Compare: The "Score" vs. The "Map"
The researchers looked at the results in two very different ways:
1. Surface-Level Similarity (The Scoreboard)
This is like looking at the final score of a soccer game. If the AI says "8/10" for a joke and a human says "8/10," they look like they agree.
- The Finding: On this level, the AI models, especially the smartest one (GPT-4), looked very human-like. They gave similar scores to humans for most sentences. It looked like they were on the same team.
2. Representational Similarity (The Map)
This is the deeper part. Imagine you and a friend are both drawing a map of a city.
- Human Map: You know that the "Pizza Place" is right next to the "Park," and the "School" is far away from the "Cemetery." Your map has a specific structure based on how you understand the world.
- AI Map: The AI might also put the "Pizza Place" and "Park" close together, but for the wrong reasons. Maybe it thinks they are close because they both have the word "place" in them, not because of how they function in real life.
The researchers used a mathematical tool called Representational Similarity Analysis (RSA) to compare the structure of these maps. They didn't just check if the scores matched; they checked if the relationships between the sentences were the same.
The Big Reveal: The "Imposter" Effect
Here is what they found:
- The Surface is Deceptive: The AI models were great at mimicking human scores. If you asked, "Is this sarcastic?", the AI often said "Yes," just like a human would.
- The Map is Different: When the researchers looked at the underlying "maps," the AI and the humans were actually quite far apart. The AI organized the meaning of sentences differently than humans do.
The Analogy:
Think of the AI as a very talented parrot.
- If you ask the parrot, "Is this sentence funny?" it can say "Yes" perfectly, just like a human.
- But the parrot doesn't actually get the joke. It just knows that when humans hear this specific pattern of words, they usually say "Yes."
- The human, however, understands the joke because they understand the social context, the tone, and the hidden meaning.
The Results by "Flavor"
The study found that the AI's "parrot skills" worked better for some types of language than others:
- Emotional & Conventional Sentences: The AI's map was somewhat similar to the human map. It's easier to guess that "I am sad" is negative than to guess that "I'm so excited I wet my plants" is funny.
- Idioms, Sarcasm, and Slang: This is where the AI failed the "Map Test."
- Idioms: When humans hear "Break a leg," they know it means "Good luck." The AI often struggled to map this correctly to the concept of "good luck" without getting confused by the literal words "break" and "leg."
- Sarcasm & Slang: These rely heavily on social context and shared human experiences (like Gen Z slang). The AI's internal map for these was very different from the human map. It was like the AI was speaking a different language, even though the words looked the same.
The "Big Brain" vs. The "Small Brain"
The study compared different AI models:
- GPT-4: This was the "big brain." It had the most similar map to humans, but it still wasn't a perfect match. It was the closest thing to a human, but it still had gaps.
- Smaller Models (Llama, Mistral): These were further away. Their maps were very different from humans, especially when it came to tricky stuff like sarcasm and slang.
The Bottom Line
The paper concludes that just because an AI gives the same answer as a human, it doesn't mean it understands the language the same way.
The AI is excellent at pattern matching (mimicking the scoreboard), but it lacks the deep, social, and contextual "grounding" that humans have. It organizes meaning in its own unique way, which works well for simple tasks but falls apart when the language gets tricky, emotional, or culturally specific.
In short: The AI is a very good actor that can say the right lines, but it doesn't necessarily feel the emotions behind them.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.