← Latest papers
💻 computer science

GPT-4o reads the mind in the eyes

This study reveals that GPT-4o outperforms humans in inferring mental states from upright faces but underperforms on inverted faces and exhibits racial bias, demonstrating a complex mix of human-like cognitive signatures and distinct, systematic processing differences compared to human cognition.

Original authors: James W. A. Strachan, Oriana Pansardi, Eugenio Scaliti, Marco Celotto, Krati Saxena, Chunzhi Yi, Fabio Manzi, Alessandro Rufo, Guido Manzi, Michael S. A. Graziano, Stefano Panzeri, Cristina Becchio

Published 2026-06-24
📖 5 min read🧠 Deep dive

Original authors: James W. A. Strachan, Oriana Pansardi, Eugenio Scaliti, Marco Celotto, Krati Saxena, Chunzhi Yi, Fabio Manzi, Alessandro Rufo, Guido Manzi, Michael S. A. Graziano, Stefano Panzeri, Cristina Becchio

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: Can an AI "Read" a Room?

Imagine you are at a party. You look at someone's eyes and instantly know they are bored, anxious, or secretly amused. Humans are experts at this; we can "read minds" just by looking at the eyes.

This paper asks a simple question: Can a super-smart computer (GPT-4o) do the same thing?

The researchers didn't just ask the AI to guess; they put it through a rigorous test called the "Reading the Mind in the Eyes Test." They showed the AI and a group of humans pictures of eyes and asked, "What is this person feeling?" The answers had to be chosen from a list of four words (e.g., Anxious, Confident, Sarcastic, Bored).

The Experiment: The Upside-Down Trick

To see how the AI and humans were thinking, the researchers used a classic trick: they showed the pictures upside down.

  • The Human Reaction: When humans see a face upside down, it's like trying to read a book written in a foreign language you barely know. It's harder, and we make more mistakes. But we still use the same "mental dictionary" to guess the emotion.
  • The AI Reaction: The AI was actually better than humans at reading the eyes when the faces were right-side up. However, when the faces were flipped upside down, the AI didn't just get a little confused; it completely lost its way. Its accuracy plummeted, dropping below what you'd expect from random guessing.

The Analogy: Imagine a human and a robot trying to solve a puzzle.

  • Right-side up: The robot is a genius, solving the puzzle faster and more accurately than the human.
  • Upside down: The human just turns the puzzle over in their hands and keeps trying. The robot, however, seems to have forgotten how puzzles work entirely and starts throwing pieces at the wall.

The "Bias" Problem: Seeing What You Expect

The researchers also tested the AI with faces of different races (White and Non-White).

  • Humans: In this study, people were equally good at reading the eyes of people from all racial backgrounds.
  • The AI: The AI was much better at reading the eyes of White faces than Non-White faces. It's as if the AI has a "favorite" group it understands better, likely because it was trained on more data featuring those faces.

The "Error" Detective Work: How They Got It Wrong

The most fascinating part of the study wasn't just about who got the right answer, but how they got the wrong answer.

Think of mistakes like footprints in the snow.

  • Human Footprints: When humans make a mistake, their footprints are a bit scattered. If they guess "Angry" instead of "Scared," they might guess "Happy" next time. Their errors are somewhat random and change depending on the situation.
  • AI Footprints: The AI's footprints were incredibly consistent. If it was going to make a mistake, it made the exact same mistake every single time.
    • The Twist: When the faces were upside down, the AI didn't just get random footprints. It developed a completely new, rigid pattern of mistakes that was totally different from how it made mistakes with right-side-up faces.

The Takeaway: Humans change their efficiency when things get hard (upside down), but they use the same logic. The AI, however, seems to switch to a totally different, broken logic when things get hard.

The "Information" in Mistakes

The researchers used a math concept called "Information Theory" to analyze these mistakes. They found that the AI's mistakes actually contained more information than human mistakes.

The Analogy:

  • Human Mistake: Like a student who guesses randomly on a test. Their wrong answers tell you nothing about what they were thinking.
  • AI Mistake: Like a student who consistently confuses "Cat" with "Dog" but never confuses "Cat" with "Car." Even though they are wrong, their pattern tells you exactly how their brain (or code) is connecting ideas. The AI wasn't just guessing; it was consistently applying a specific, albeit incorrect, rulebook.

Summary of Findings

  1. Superhuman (but fragile): GPT-4o is incredibly good at reading emotions from eyes when things look normal, often beating humans.
  2. The Inversion Effect: When faces are upside down, the AI breaks down much more severely than humans do.
  3. Racial Bias: The AI is better at reading White faces than Non-White faces, whereas humans in this study showed no such difference.
  4. Systematic Errors: The AI doesn't make random mistakes. It makes highly consistent, structured mistakes that reveal a specific way it processes information—one that is fundamentally different from how humans process information, especially when the input is tricky (like upside-down faces).

In short, the AI is a brilliant but rigid student. It has memorized the "right" way to read faces from its training data, but when the test gets weird (upside down) or the faces look different (different races), it doesn't adapt like a human; it just follows a broken script.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →