Structural Ranking of the Cognitive Plausibility of Computational Models of Analogy and Metaphors with the Minimal Cognitive Grid
This paper applies the Minimal Cognitive Grid framework to systematically evaluate and rank the cognitive plausibility of leading computational models of analogy and metaphor, including SME, CogSketch, METCL, and LLMs, using a formalized quantitative approach based on functional/structural ratio, generality, and performance match.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a judge at a talent show, but instead of singers or dancers, the contestants are computer programs trying to act like human thinkers. Specifically, they are trying to solve puzzles involving analogies (like "A is to B as C is to D") and metaphors (like "My job is a jail").
The authors of this paper, Alessio Donvito and Antonio Lieto, created a special scoring system called the Minimal Cognitive Grid (MCG). Think of this grid as a three-legged stool. For a computer to sit comfortably on the stool and be considered a "cognitively plausible" model (meaning it thinks like a human, not just looks like it), it needs to balance on all three legs.
Here is how the three legs work, using simple analogies:
Leg 1: The Blueprint Match (Functional/Structural Ratio)
The Question: Does the computer build its answer the same way a human brain does, or does it just get the right answer by magic?
- The Analogy: Imagine two people trying to build a birdhouse.
- Person A (The Structural Model): Uses wood, nails, and a hammer, following the exact same steps a carpenter would. They might be slow, but their process is human-like.
- Person B (The Functional Model): Uses a 3D printer or a laser cutter. They get the exact same birdhouse, but the machine inside is totally different from a human's hands.
- The Paper's Claim: The authors argue that for a computer to be a good model of how we think, it needs to use the "wood and nails" approach (structural).
- SME and CogSketch: These are like the carpenters. They use specific rules (like "one thing maps to one thing") that mimic how humans make analogies. They get high scores here.
- LLMs (like GPT-4): These are like the 3D printers. They are amazing at making the birdhouse (getting the right answer), but they do it by guessing patterns from a massive library of books, not by following human-like logical steps. They get very low scores here because their "internal machinery" is alien to human thought.
- METCL: This is a middle-ground carpenter. It uses a different set of blueprints (focusing on categorization rather than mapping), so it scores in the middle.
Leg 2: The Swiss Army Knife (Generality)
The Question: Can the computer do many different types of thinking, or is it a one-trick pony?
- The Analogy:
- The Specialist: A master chess player who can beat a grandmaster but can't tie their own shoes or cook an egg.
- The Generalist: Someone who can play chess, cook, drive a car, and solve math problems.
- The Paper's Claim: The authors look at whether these systems can handle math, visual puzzles, language, and even physical movement (sensory/motor skills).
- LLMs: These are the ultimate Generalists. They can read, write, do math, and look at pictures. They score highest here.
- SME and CogSketch: These are the Specialists. They are great at specific logic puzzles (like visual patterns) but can't really do math or understand complex sentences. They score low here.
- METCL: Also a specialist, focused mostly on language metaphors.
Leg 3: The Mirror Test (Performance Match)
The Question: When the computer gets it wrong, does it make the same kind of mistakes a human would?
- The Analogy: Imagine taking a test.
- If you get a question wrong because you misunderstood the concept, that's a "human" mistake.
- If you get it wrong because the computer glitched or guessed randomly, that's a "machine" mistake.
- The goal is to see if the computer's "glitches" look like human confusion.
- The Paper's Claim:
- SME and CogSketch: They are excellent at this. When they fail a visual puzzle (like the Raven's Progressive Matrices), they fail on the exact same hard puzzles that humans struggle with. They even take a similar amount of time to think. They score very high.
- LLMs: They often get the right answer, but when they fail, their mistakes look different from humans. They might get confused by a slight change in wording that wouldn't trip up a person. They score lower here.
The Final Scoreboard
The authors combine these three scores to see who wins the "Cognitive Plausibility" title. They decided that Leg 1 (The Blueprint) is the most important because the main goal is to understand how humans think, not just to get the right answer.
- The Winner (Under strict rules): CogSketch and SME. Even though they are "one-trick ponies" (low generality), they think like humans. They are the best models for studying the human mind.
- The Runner-Up: METCL. It sits in the middle, doing some human-like things but not as comprehensively as the top two.
- The "Magic" Box: LLMs (GPT-4, etc.). They are incredibly smart and versatile (high generality), but they don't think like humans. If you want a tool to solve problems, they are great. But if you want a model to explain why humans think the way they do, the paper says they are not the best fit.
The Big Takeaway
The paper concludes that being "smart" (getting the right answer) is different from being "cognitively plausible" (thinking like a human).
- LLMs are like a super-actor who can memorize a script perfectly and sound exactly like a human, but they don't actually feel the emotions.
- SME/CogSketch are like a method actor who actually feels the emotions and thinks through the character's logic, even if they can't memorize as many lines.
The authors built this grid to help scientists decide which computer program is the best "method actor" for studying the human mind. They found that while the "super-actors" (LLMs) are impressive, the "method actors" (SME/CogSketch) are still the gold standard for understanding the mechanics of human thought.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.