← Latest papers
🤖 AI

On Benchmarking Human-Like Intelligence in Machines

This paper argues that current AI evaluation paradigms fail to accurately assess human-like cognitive capabilities due to issues like unvalidated labels and ecologically invalid tasks, and proposes five concrete recommendations for more rigorous benchmarking based on a human evaluation study of nine existing benchmarks.

Original authors: Lance Ying, Katherine M. Collins, Lionel Wong, Ilia Sucholutsky, Ryan Liu, Adrian Weller, Tianmin Shu, Thomas L. Griffiths, Joshua B. Tenenbaum

Published 2026-08-13
📖 4 min read☕ Coffee break read

Original authors: Lance Ying, Katherine M. Collins, Lionel Wong, Ilia Sucholutsky, Ryan Liu, Adrian Weller, Tianmin Shu, Thomas L. Griffiths, Joshua B. Tenenbaum

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to be a person. You don't just want it to be super smart at math or able to predict the weather; you want it to think, feel, and react the way you do. This is the world of "human-like intelligence." For decades, scientists have been building these digital minds, hoping to create machines that don't just solve problems, but understand the messy, confusing, and sometimes contradictory way humans handle the world. But here's the tricky part: how do you know if the robot is actually acting like a human, or if it's just a really good actor pretending? To find out, researchers usually give the robot a test. But what if the test itself is broken? What if the test expects the robot to give a single, perfect answer, when real humans would actually give a dozen different, slightly unsure answers? This paper dives into that exact problem, arguing that our current way of testing AI is like judging a jazz improvisation by asking the musician to play a single, perfect note.

The authors of this paper, a team of researchers from top universities, argue that we are currently failing to measure "human-like" intelligence correctly. They suggest that many popular AI tests are too simple and rely on "ground truth" labels that don't actually match how real people think. To prove their point, they ran a massive experiment where they asked 117 real humans to take nine different AI tests. These tests covered things like spotting sarcasm, deciding if a joke is funny, or figuring out if a social situation is supportive. Instead of just picking "A, B, or C," the humans used sliders to show how strongly they felt about an answer and how confident they were in that feeling.

The results were eye-opening. The paper suggests that for many of these tests, the "correct" answer the AI was trained on was actually wrong compared to what most humans thought. For instance, on a task about "social support," the test said a specific sentence was "unsupportive," but the humans rated it as moderately supportive on average. Even more interestingly, the humans didn't just disagree with the test; they disagreed with each other in a structured way. Some people were very sure about their answers, while others were unsure, and their opinions varied across a wide spectrum. The paper finds that current AI benchmarks often ignore this "messiness." They treat human judgment like a math problem with one right answer, when it's actually more like a conversation where everyone has a slightly different perspective.

The authors propose five new rules for building better tests. First, we need to stop guessing what humans think and actually ask them, collecting real data from people to use as the "gold standard." Second, instead of looking for one right answer, we should compare AI to the whole distribution of human answers—seeing if the AI can capture the fact that some people think one way and others think another. Third, we need to measure "gradedness." Humans rarely feel 100% certain or 0% certain; they feel "mostly yes" or "a little bit no." The paper suggests we should use sliders and confidence ratings to capture this nuance, rather than forcing a simple yes/no choice. Fourth, the tests need to be based on deep psychological theories, not just random questions. Just because a child passes a "false belief" test doesn't mean a robot passing the same test has a real "Theory of Mind"; the test needs to be designed to actually measure the complex parts of that skill. Finally, the tasks themselves need to be more like real life. Real life is ambiguous, confusing, and full of context. If a test is too clean and simple, it doesn't tell us if the AI can handle the real world.

In short, the paper suggests that to build AI that truly thinks like us, we have to stop treating human intelligence like a multiple-choice quiz. We need to embrace the uncertainty, the disagreement, and the gray areas. By designing tests that reflect the rich, varied, and sometimes confusing nature of human thought, we can finally start to build machines that don't just look smart, but actually feel human.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →