← Latest papers
💬 NLP

When Models Know More Than They Say: Probing Analogical Reasoning in LLMs

This paper reveals an asymmetry in open-source LLMs where probing internal representations significantly outperforms prompting in detecting rhetorical analogies, suggesting that prompting mechanisms fail to fully access latent information required for complex analogical reasoning.

Original authors: Hope McGovern, Caroline Craig, Thomas Lippincott, Hale Sirin

Published 2026-04-07
📖 4 min read☕ Coffee break read

Original authors: Hope McGovern, Caroline Craig, Thomas Lippincott, Hale Sirin

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a giant, super-smart library robot (a Large Language Model, or LLM) that has read almost everything ever written. You want to know: Does this robot actually understand stories, or is it just really good at guessing the next word based on patterns it's seen before?

To find out, the authors of this paper built a special test called NARB (Narrative Analogical Reasoning Benchmark). They wanted to see if the robot could spot "hidden connections" between different stories, a skill humans use all the time.

Here is the breakdown of their discovery, explained with some everyday analogies.

The Two Types of Tests

The researchers gave the robot two different kinds of puzzles:

  1. The "Story Match" (Narrative Parallelism):

    • The Task: Imagine you tell the robot a story about a baker who fails, learns, and succeeds. Then you ask it to pick the best matching story from a list of three options: one about a baker, one about a soldier, and one about a gardener.
    • The Catch: The "right" answer might be about a soldier. The robot has to ignore the surface details (flour vs. guns) and see the deep structure (struggle \rightarrow growth \rightarrow triumph).
    • The Result: The robot struggled. It was like trying to find a needle in a haystack. It could only get about 35% of these right. It seemed to get lost in the details and missed the big picture.
  2. The "Poem Match" (Rhetorical Parallelism):

    • The Task: This is like spotting a specific rhythm or rhyme scheme in a poem. "The sun rises in the east, the moon sets in the west." The robot has to find another sentence that follows that exact same structural pattern.
    • The Result: When the researchers looked inside the robot's "brain" (using a technique called probing), they found the robot knew the answer almost perfectly (93%). It was like the robot had the answer written on its forehead.

The Big Surprise: The "Silent Genius"

Here is where it gets weird. The researchers compared two ways of asking the robot questions:

  • Method A: Asking it nicely (Prompting). "Hey robot, which story matches this one?"
  • Method B: Looking inside its brain (Probing). "Robot, what are your internal thoughts about these stories?"

The Shocking Discovery:

  • For the Poem Match: When they asked the robot, it gave terrible answers (18% correct). But when they looked inside its brain, the robot knew the answer perfectly (93% correct).
    • Analogy: Imagine a student taking a math test. They get the answer wrong on the paper because they are nervous or can't explain their steps. But if you ask them to whisper the answer to a teacher, they get it right every time. The knowledge is there, but they can't show it when asked.
  • For the Story Match: Both methods failed. The robot didn't know the answer when asked, and it didn't know the answer when they looked inside its brain.
    • Analogy: This student genuinely didn't study for the test. They don't know the material, and they can't fake it.

What Does This Mean?

The paper concludes that LLMs are "Silent Geniuses" for some things and "Shallow Guessers" for others.

  1. They know more than they say: For things like spotting patterns in language (rhetoric), the robot has the knowledge locked inside its layers, but it can't access it when you just ask it a question. It's like having a library in your head but forgetting how to open the door.
  2. They struggle with deep stories: For understanding complex life lessons or story structures, the robot genuinely doesn't have the deep understanding yet. It's not just a communication problem; it's a knowledge problem.

The Takeaway

If you only judge these AI models by how they answer your questions (prompting), you might think they are dumber than they actually are. They might be hiding their true intelligence.

However, for the really hard stuff—like understanding the deep meaning of a story—they are still learning. The paper suggests we need to use both methods (asking them and peeking inside their brains) to get the full picture of what these digital minds can actually do.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →