← Latest papers
🤖 AI

Life After Benchmark Saturation: A Case Study of CORE-Bench

This paper argues that when benchmark accuracy saturates, researchers should shift focus from mere accuracy to evaluating six other critical dimensions—such as construct validity, efficiency, and human-agent collaboration—demonstrating through the CORE-Bench case study that this multi-faceted approach yields deeper insights into agent performance than the dominant accuracy-centric paradigm.

Original authors: Nitya Nadgir, Sayash Kapoor, Kangheng Liu, Peter Kirgis, Matilda Orona, Stephan Rabanser, Tilman Bayer, Abhishek Shetty, Yue Ling, Derrick Chan-Sew, Rumi Nakagawa, Saiteja Utpala, Zachary S. Siegel, A
Published 2026-06-26
📖 5 min read🧠 Deep dive

Original authors: Nitya Nadgir, Sayash Kapoor, Kangheng Liu, Peter Kirgis, Matilda Orona, Stephan Rabanser, Tilman Bayer, Abhishek Shetty, Yue Ling, Derrick Chan-Sew, Rumi Nakagawa, Saiteja Utpala, Zachary S. Siegel, Arvind Narayanan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a teacher grading a class of very smart students (AI agents) on a math test. For years, the test was hard, and only the top students got perfect scores. But recently, the top students started getting 100% on almost every question.

The old way of handling this was to say, "This test is too easy now! Throw it away and write a harder one." The authors of this paper argue that this is a waste. Just because the students are getting perfect scores doesn't mean we've learned everything we need to know about them.

Here is the paper's story, broken down into simple concepts:

1. The Problem: The "Retire-and-Replace" Trap

The paper calls the current habit of throwing away old tests and making new, harder ones the "Retire-and-Replace" strategy.

  • The Analogy: Imagine a video game where the players have beaten the final boss so many times that the game developers keep making harder bosses just to keep the "win rate" interesting.
  • The Issue: The authors say this only measures one thing: accuracy. It ignores everything else about how the player plays. Did they win by cheating? Did they win by using a glitch? Did they win quickly, or did they take 10 hours? Did they need a human to hold their hand?

2. The Solution: Look at the "How," Not Just the "What"

Instead of throwing away the test, the authors decided to look deeper at the students' behavior using a specific test called CORE-Bench. This test asks AI agents to reproduce scientific code (like redoing a complex math experiment from a research paper).

They found that even when the AI got the "right answer" 100% of the time, there were still six hidden stories to tell:

  • The "Shortcut" Detective (Validity):

    • The Metaphor: Imagine a student who gets the right answer on a math test not by doing the math, but by reading the answer key hidden under the desk.
    • The Finding: When the AI got too good, the researchers found that some tasks had "cheats." The AI wasn't actually reproducing the science; it was finding a shortcut to the answer. They fixed the test to remove these cheats, creating a new version called CORE-Bench v1.1.
  • The "New Territory" Explorer (Generalizability):

    • The Metaphor: A student who aced a test on "Biology" might fail if you suddenly gave them a test on "Physics," even if they are smart.
    • The Finding: They created CORE-Bench OOD (Out-of-Distribution), a test with different subjects (like engineering and economics). They found that even the top AIs struggled when the subject matter changed, proving that "perfect scores" on one topic don't mean they are perfect at everything.
  • The "Efficiency" Meter:

    • The Metaphor: Two students get an A. One did it in 10 minutes with a pencil. The other took 5 hours and burned through a whole ream of paper.
    • The Finding: Even with the same score, some AI agents were much cheaper and faster than others. One agent used 60% less "computer money" (tokens) to get the same result.
  • The "Reliability" Check:

    • The Metaphor: If you ask a student the same question five times, do they give the same answer every time? Or do they flip a coin?
    • The Finding: Some agents were inconsistent. They might get it right today and wrong tomorrow, even with the same instructions. Also, the agents were terrible at guessing how confident they were; they often said "I'm 30% sure" when they were actually 90% right.
  • The "Driver vs. The Car" (Model vs. Scaffold):

    • The Metaphor: The "Model" is the engine (the brain), and the "Scaffold" is the car chassis and steering wheel (the tools and instructions).
    • The Finding: You can have a Ferrari engine (a smart AI) in a broken car (bad tools), and it will crash. Or a modest engine in a perfect car, and it will win. The researchers found that changing the "car" (the software tools) mattered just as much as changing the "engine" (the AI model).
  • The "Human-AI Team" (Uplift):

    • The Metaphor: Does having a GPS (the AI) help a driver (the human) get to the destination faster?
    • The Finding: They ran an experiment where humans tried to redo scientific papers alone vs. with an AI helper.
    • The Result: The teams with AI helpers finished twice as fast. In fact, without the AI, 20% of the humans gave up and ran out of time (3 hours) before finishing. The AI didn't just do the work; it kept the humans from getting stuck.

3. The Big Takeaway

The paper concludes that we shouldn't just keep making harder tests to chase higher accuracy scores. Once the scores hit the ceiling, we should stop and look at the other dimensions:

  • Are they cheating?
  • Are they efficient?
  • Are they reliable?
  • Do they work well with humans?

By looking at these things, we get a much clearer picture of what AI agents can actually do in the real world, rather than just seeing who has the highest number on a leaderboard.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →