Taming the Centaur(s) with LAPITHS: a framework for a theoretically grounded interpretation of AI performances
The paper introduces the LAPITHS framework to challenge the theoretical validity of claims that AI models like CENTAUR possess human-like cognition, arguing instead that their performance stems from behavioristic patterns rather than genuine cognitive plausibility.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Picture: The "Centaur" vs. The "Lapiths"
Imagine a new, super-smart AI named Centaur. Its creators claim it is a "Unified Model of Cognition." In plain English, they say: "We taught this AI to mimic human behavior so perfectly that it must think like a human. It's not just a calculator; it's a digital mind."
The authors of this paper, a team of researchers from Italy, say: "Hold on a minute."
They introduce a framework called LAPITHS (named after the ancient Greek warriors who fought the Centaurs). Their job is to "tame" the hype around Centaur. They argue that just because an AI acts like a human, it doesn't mean it thinks like a human. They want to separate acting (performance) from thinking (cognition).
The Core Problem: The "Ascription Fallacy"
The paper identifies a logical trap called the Ascription Fallacy.
The Analogy:
Imagine you see a robot that can perfectly mimic a human playing chess. It moves the pieces, pauses to "think," and even makes the same mistakes a human would make when tired.
- The Fallacy: You look at the robot and say, "Wow, it must have a brain like a human's! It understands strategy and feels the pressure of the game."
- The Reality: The robot might just be a super-fast calculator that memorized millions of chess games. It doesn't "understand" anything; it just predicts the next move based on patterns.
The authors argue that Centaur is this chess robot. It was trained on a massive dataset of human decisions (called Psych-101). Because it was trained to predict what humans would do, it predicts them very well. But the authors claim this is just behavioral mimicry, not cognitive alignment.
The Tool: The "Minimal Cognitive Grid" (MCG)
To prove their point, the authors built a scorecard called the Minimal Cognitive Grid (MCG). Think of this as a "Cognitive Truth Detector" that grades AI on three things, not just how well it answers questions.
Structure vs. Function (The "How"):
- The Question: Does the AI use the same internal "gears" as a human brain?
- The Finding: Humans learn incrementally (trial and error in real-time) and have limited working memory (we can only hold a few things in our heads). Centaur, however, was trained in batches (like reading a whole book at once) and has a massive "context window" (it can remember everything in a conversation).
- The Verdict: Centaur fails this test. It acts like a human, but its internal engine is totally different. It's like a plane flying like a bird, but using jet engines instead of flapping wings.
Generality (The "Scope"):
- The Question: Can it do many different types of thinking?
- The Finding: Centaur is good at math, reasoning, and language. But it is unimodal, meaning it only processes text. It has no eyes to see, no hands to touch, and no body to move.
- The Verdict: It gets partial points. It's smart, but it lacks the "embodied" experience of a human.
Performance Match (The "Result"):
- The Question: Does it get the right answer, make the same mistakes, and take the same amount of time?
- The Finding: Centaur is excellent at this. It predicts human choices almost perfectly and even mimics human error patterns.
- The Verdict: High score here. But the authors warn: High score here does not prove it thinks like us.
The Final Score: When you combine these, Centaur gets a low "Cognitive Plausibility" score. It's a great emulator (a perfect actor), but a poor cognitive model (a true thinker).
The Experiment: "Can Anyone Do This?"
To prove that Centaur isn't special, the researchers ran a test. They took five other powerful AI models (like GPT-4, Gemini, and others) that were NOT trained on the human data.
They gave these models a simple instruction: "Here is the rule of the game. Here is the history of what happened. Now, what would a human do?" (This is called RAG or Retrieval-Augmented Generation).
The Result:
- These "untrained" models performed almost as well as Centaur.
- In some cases, the difference in performance was so small it was statistically meaningless.
- The Metaphor: It's like giving a group of actors a script. Centaur is the actor who memorized the script perfectly. But when you give the same script to other actors, they can also deliver the lines almost perfectly without having memorized the whole book beforehand.
Conclusion: You don't need a "special" cognitive model to mimic human behavior; you just need a smart language model and the right instructions.
The Brain Scan Test (fMRI)
Centaur's creators claimed that after training, the AI's "internal thoughts" (neural representations) started to look like human brain scans (fMRI data). They said this proved the AI was becoming "brain-like."
The authors tested this too. They asked other AI models to guess what the brain activity would look like for a specific task, just by reading the task description.
The Result:
- The other models also produced guesses that correlated highly with human brain scans.
- The Metaphor: If you ask a group of people to describe the feeling of "fear," they might all use similar words (heart racing, sweating). That doesn't mean they all have the exact same biological fear mechanism; it just means they are all describing the same concept.
- The authors argue that high correlation in brain scans is easier to get than the creators admit. It doesn't prove the AI has a brain; it just proves it understands the concept of the task.
The Final Takeaway
The paper concludes with a warning for the AI world:
- Don't confuse the map with the territory. Just because an AI draws a perfect map of human behavior (predicts what we do), it doesn't mean the AI is the territory (the human mind).
- Centaur is a great predictor, but a weak explanation. It is a fantastic tool for guessing what humans will do next. But we cannot use it to explain why humans think the way they do, because its internal machinery is fundamentally different from ours.
- We need better tests. We need to stop just asking "Did it get the right answer?" and start asking "Did it get the right answer using the same mental gears as a human?"
In short: Centaur is a masterful actor, but it is not a human. The authors have built a framework (LAPITHS) to ensure we don't get fooled by the performance.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.