← Latest papers
💬 NLP

Parallelograms Strike Back: LLMs Generate Better Analogies than People

This study demonstrates that large language models generate higher-quality four-term word analogies than humans by more consistently adhering to the geometric "parallelogram" relational structure, suggesting that the model's perceived failure in human cognition stems from human inconsistency rather than the model's inadequacy.

Original authors: Qiawen Ella Liu, Raja Marjieh, Jian-Qiao Zhu, Adele E. Goldberg, Thomas L. Griffiths

Published 2026-03-20
📖 4 min read☕ Coffee break read

Original authors: Qiawen Ella Liu, Raja Marjieh, Jian-Qiao Zhu, Adele E. Goldberg, Thomas L. Griffiths

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are playing a word game called "Complete the Pattern." The game looks like this:

King is to Queen as Man is to ___?

Most people would instantly say "Woman." This is a classic analogy. For decades, scientists thought our brains solved these puzzles like a geometry teacher: they imagined words as dots on a map, and the relationship between them as a straight line (a vector). If you draw a line from "King" to "Queen," you should be able to draw the exact same line starting from "Man" to land on "Woman." This is called the Parallelogram Model.

However, a few years ago, researchers found that humans are actually terrible at following these perfect geometric lines. When asked to come up with the answers, people often just grab the first word that sounds similar to the third word (e.g., "Man" is similar to "Boy," so maybe the answer is "Boy?"). It seems like our brains are lazy shortcuts artists, not geometry masters.

Enter the "Robot" (LLMs)

The authors of this paper asked a fascinating question: Is the geometry model actually the right way to solve these puzzles, or are we just bad at doing it?

To find out, they pitted Humans against Large Language Models (AI) (like the smart chatbots you might use today). They gave both groups thousands of these word puzzles and asked a third group of humans to judge: "Which answer makes more sense?"

Here is what they discovered, explained through a few simple metaphors:

1. The Robots are Better at the "Geometry"

When the AI generated answers, human judges rated them as better than the answers humans gave.

  • The Metaphor: Imagine a game of darts. Humans are throwing darts while looking at the board with one eye closed; they often hit the wall or the floor. The AI, however, is throwing darts with a laser sight. The AI's answers fit the "perfect geometric shape" (the parallelogram) much more often than human answers do.
  • The Twist: The AI didn't necessarily "think" in geometry. It just happened to produce answers that looked like perfect geometry to us.

2. The "Long Tail" of Bad Human Answers

Why did the AI win? Was it because the AI was a genius and humans were stupid? No.

  • The Metaphor: Think of a bakery. Humans are like a chaotic kitchen where 80% of the time, the baker makes a perfect loaf of bread. But 20% of the time, they burn the bread, drop it on the floor, or serve a rock.
  • The AI, on the other hand, is like a robot baker that always makes a decent loaf. It rarely makes a masterpiece, but it almost never makes a disaster.
  • The Result: When you average all the answers, the AI wins because humans have a "long tail" of terrible, random answers dragging their average down. If you only compare the best answer humans gave against the best answer the AI gave, they are actually tied. The AI's advantage comes from being consistent, not from being a genius.

3. The "Clever Word" Factor

Humans tend to pick words that are very common and easy to think of (like "Boy" or "Child"). The AI, however, often picked slightly more obscure or "fancier" words (like "Immaturity" instead of "Youth").

  • The Metaphor: If you ask a human to describe a "sad feeling," they might say "sad." If you ask an AI, it might say "melancholy."
  • Humans actually prefer the AI's "fancier" words. We think the AI's answers are smarter because they use less common words that still fit the pattern perfectly.

4. The Big Conclusion: We Know What a Good Analogy Looks Like

The most surprising finding is about the "Parallelogram Model" itself.

  • Old Idea: "The parallelogram model is wrong because humans don't use it."
  • New Idea: "The parallelogram model is right about what makes a good analogy, but humans are just bad at making them."

The Takeaway:
We humans are like drivers who know exactly what a "perfectly parked car" looks like, but we are terrible at actually parking the car ourselves. We get distracted, we rush, and we hit the curb. The AI, however, is a self-driving car that parks perfectly every single time.

The study proves that the "geometric" way of thinking about analogies isn't broken; it's just that our messy human brains struggle to follow the rules, while the AI follows them effortlessly. We value the perfect pattern, even if we can't always create it.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →