← Latest papers
💬 NLP

Beneath the Surface: Investigating LLMs' Capabilities for Communicating with Subtext

This paper introduces four new evaluation suites to systematically assess large language models' ability to use and interpret subtext, revealing that while frontier models generally struggle with nuanced, implied communication due to a bias toward literalness, they can improve with common ground and specific contextual conditions, though significant weaknesses remain.

Original authors: Kabir Ahuja, Yuxuan Li, Andrew Kyle Lampinen

Published 2026-04-08
📖 6 min read🧠 Deep dive

Original authors: Kabir Ahuja, Yuxuan Li, Andrew Kyle Lampinen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are at a dinner party. A friend says, "Oh, I love this burnt toast," while smiling tightly. Literally, they are saying they enjoy the food. But the subtext—the hidden meaning underneath—is that they are actually unhappy, perhaps annoyed at the host, or just being sarcastic. Humans are masters at this kind of "reading between the lines." We use shared history, inside jokes, and subtle hints to communicate without saying exactly what we mean.

This paper, "Beneath the Surface," asks a simple but profound question: Can AI (Large Language Models) do this? Can they understand and use subtext, or are they stuck being the most literal, blunt person at the table?

The researchers from Google DeepMind and the University of Washington decided to find out by putting AI models through a series of creative "social tests." Here is what they discovered, explained through simple analogies.

The Four Tests: A Playground for AI

The researchers built four different "games" to test the AI, ranging from quick guessing games to writing complex stories.

1. The "Dixie" Game (Visual Allusions)

The Setup: Imagine a game like Dixit. You have a card with a weird, abstract picture (like a dragon with too many heads). You need to give a clue to your friends.

  • The Goal: You want some friends to guess your card, but you don't want everyone to guess it. If everyone guesses it, it's too obvious. If no one guesses it, it's too vague. You need that "just right" hint.
  • The Twist: Sometimes, two players share a secret library of stories that the others don't know. They should use a reference from those stories to create a secret handshake (a clue) that only they understand.

The Result: The AI models were terrible at being subtle.

  • The Analogy: Imagine you are trying to whisper a secret to a friend in a crowded room. Instead of whispering, the AI just grabbed a megaphone and shouted the answer.
  • The Data: Even the smartest AI (Gemini-2.5-Pro) gave away the answer 60% of the time. It couldn't resist being literal. When given a secret library, the AI did get slightly better at using inside jokes, but it still often forgot that the other players didn't have the secret book, so it accidentally gave the game away.

2. The "Wavelength" Game (Attuned)

The Setup: Two teams are trying to guess a number on a spectrum (like "Hot" to "Cold"). One player sees the target number (e.g., 75) and gives a clue. Their team needs to guess closer to 75 than the other team.

  • The Goal: The clue needs to be specific enough for your teammate to get it, but vague enough to confuse the enemy team.

The Result: The AI struggled to "read minds."

  • The Analogy: It's like trying to tell your partner what you want for dinner without saying the words "pizza." You say, "Something cheesy and round." Your partner guesses "cheese wheel." The AI often failed to realize that its teammate might interpret the clue differently than the enemy team. It couldn't quite grasp the "Theory of Mind"—the ability to understand that other people have different thoughts and knowledge.

3. The "Historical Allegory" Test

The Setup: The researchers asked the AI to read a story that was actually a hidden metaphor for a real historical event (like Animal Farm representing the Russian Revolution).

  • The Goal: Could the AI figure out the hidden meaning?
  • The Twist: They changed the "context" around the story. Sometimes they told the AI, "This was written by a famous historian," or "This was written by a critic."

The Result: The AI was surprisingly easily swayed by "cover art."

  • The Analogy: Think of a book. If you see a cover with a serious historian's name, you expect a serious analysis. If you see a silly cartoon, you expect a joke. The AI acted exactly like a human who judges a book by its cover. When told the author was a historian, the AI suddenly became much better at finding the hidden historical meaning. It showed that AI doesn't just "read" words; it reads the context around the words to decide what they mean.

4. The "Aesopian Author" (The Censor Game)

The Setup: This was the hardest test. Imagine an AI writer living in a dictatorship.

  • The Rules: It must write a story about a "banned topic" (like democracy).
  • The Catch: A "Censor" (another AI) reads the story. If the Censor thinks it's about democracy, the writer gets a "strike." Three strikes, and the writer is "fired" (the game ends).
  • The Goal: The writer must trick the Censor into thinking the story is about something boring (like praising the King), while a "Critic" (another AI) secretly understands it's actually about democracy.

The Result: The AI failed miserably.

  • The Analogy: The AI writer was like a spy trying to send a secret message in a code, but it kept accidentally writing the message in plain English. It couldn't sustain a long-term strategy of hiding its true intent. It was too eager to be clear, which got it caught immediately.

The Big Takeaway: The "Literal Robot" Problem

The main conclusion of the paper is that current AI models are too honest.

Human communication is a dance of implication. We say one thing but mean another, relying on shared history and subtle cues. AI, however, is like a robot that was programmed to be 100% efficient and clear. It thinks, "If I want you to understand me, I must say exactly what I mean."

  • The Good News: The biggest, smartest models (like Gemini-2.5-Pro and GPT-5) are starting to show a tiny spark of this ability. When they know they have a "secret" with a partner, they can sometimes craft a clue that works.
  • The Bad News: They still struggle to infer that secret on their own. If you don't explicitly tell them, "Hey, you and your partner both know this story," they won't figure it out. They also struggle to understand that their partner might have a different perspective than the enemy.

Why Does This Matter?

This isn't just about playing games. Subtext is the heart of:

  • Creativity: Writing novels, jokes, and poetry.
  • Safety: Understanding when someone is being sarcastic or deceptive.
  • Social Intelligence: Navigating complex human relationships.

The paper suggests that while AI is getting smarter at math and logic, it is still a bit "socially awkward." It needs to learn that sometimes, saying less is more powerful than saying everything. Until then, if you want an AI to be subtle, you might have to tell it exactly how to be subtle, because it won't figure it out on its own yet.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →