← Latest papers
🤖 machine learning

Benchmarked Yet Not Measured -- Generative AI Should be Evaluated Against Real-World Utility

This paper argues that the disconnect between generative AI's strong benchmark performance and its limited real-world utility stems from flawed evaluation practices, and proposes the SCU-GenEval framework to shift assessment toward measuring sustained improvements in human outcomes within specific deployment contexts.

Original authors: Ishani Mondal, Shweta Bhardwaj

Published 2026-05-11
📖 5 min read🧠 Deep dive

Original authors: Ishani Mondal, Shweta Bhardwaj

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Problem: The "Video Game Score" vs. Real Life

Imagine you are training a robot to drive a car. You test it in a perfect, empty video game simulation where the weather is always sunny, the roads are straight, and there are no other cars. The robot gets a perfect score of 100/100.

You are thrilled! You buy the robot and put it on a real highway. Crash.

This paper argues that Generative AI (like the chatbots and code writers we use today) is exactly like that robot. It gets perfect scores on standard tests (benchmarks), but when we actually use it in real life (schools, hospitals, law offices), it often fails to help people or even causes harm.

The authors looked at 28 real-world stories where AI looked great on paper but failed in practice. They found three main reasons why this happens:

1. The "Fake Proxy" Problem (Proxy Displacement)

The Analogy: Imagine a restaurant judge who only rates food based on how loudly the chef claps when the dish is served. The chef learns to clap very loudly and gets a 10/10 rating. But the food is actually burnt.
The Reality: AI is often graded on easy-to-measure things like "fluency" (how smooth the sentences sound) or "pass rate" (does the code run?). But in the real world, we care about hard-to-measure things like "is this medical advice actually safe?" or "did this student actually learn the concept?"

  • Example: A coding AI might write code that passes all the tests (high score) but contains a security hole that lets hackers in. The test didn't measure safety; it only measured if the code ran.

2. The "Snapshot" Problem (Temporal Collapse)

The Analogy: Imagine a student who uses a calculator to solve a math problem instantly. They get an A on the test. But if you take the calculator away a month later, they can't do basic math at all. The calculator helped them in the moment, but it didn't help them learn.
The Reality: Current AI tests are "snapshots." They ask, "Can the AI do this task right now?" They don't ask, "Does using this AI help the human get better over time, or does it make them lazy and forgetful?"

  • Example: In education, AI might help a student finish homework quickly, but the student might forget how to write or think critically on their own later.

3. The "Average" Problem (Distributional Concealment)

The Analogy: Imagine a doctor says, "This new medicine works great! The average patient feels better." But they don't tell you that it works perfectly for men but makes women sick. The "average" hides the fact that half the people are getting hurt.
The Reality: AI systems often look good when you look at the "average" score. But that average hides the fact that the AI fails miserably for specific groups of people (like minorities, beginners, or people in rural areas).

  • Example: A medical AI might work well for white patients but fail to detect sickness in Black patients, or a legal AI might be great for big law firms but give terrible advice to regular people.

The Solution: A New Way to Measure Success

The authors say we need to stop asking "How good is the AI's output?" and start asking "How much did the AI change the human's ability to achieve their goals?"

They call this "Utility." It's not about the AI's score; it's about the human's progress.

To measure this, they propose a new framework called SCU-GenEval. Think of it as a four-step recipe for testing AI before you let it loose on the world:

  1. Who and What? (Stakeholder-Goal Mapping)
    • Don't just say "The Developer." Say "The Junior Developer," "The Security Expert," and "The End User." What does success look like for each of them?
  2. What Matters? (Construct-Indicator Specification)
    • Don't just measure "speed." Measure "did they learn?" "Is the code secure?" "Did the patient get better?"
  3. How Does it Change? (Mechanism Modeling)
    • Predict the future. Will using this AI make the junior developer lazy? Will it make the doctor trust the machine too much?
  4. Measure Over Time (Longitudinal Utility)
    • Don't just test once. Test today, test next week, and test next month. Did the human get better, or did they get worse?

The Tools to Make It Happen

The authors know this sounds expensive and hard to do. So, they suggest three tools to make it practical:

  • The "Pre-Flight Checklist" (Structured Protocols): Before launching an AI, you must write down exactly what you are testing, who you are testing it on, and what you expect to happen. This stops you from changing the rules after you see the results.
  • The "Digital Twin" (User Simulators): Instead of waiting months to see how real humans react, use computer simulations of specific types of people (e.g., "a tired junior coder") to predict how they will perform over time.
  • The "Specialized Ruler" (Persona-Conditioned Metrics): Instead of using one ruler for everyone, use different rulers for different groups. Measure the "Junior Developer" separately from the "Senior Expert."

The Bottom Line

The paper concludes that a high benchmark score is not enough. Just because an AI wins a video game doesn't mean it's ready for the real world.

We need to shift our focus from "How smart is the machine?" to "How much did the machine help the human become more capable?" If we don't make this shift, we risk deploying AI systems that look impressive on paper but fail to help, or even harm, the people who need them most.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →