← Latest papers
🤖 AI

Is your AI Model Accurate Enough? The Difficult Choices Behind Rigorous AI Development and the EU AI Act

This paper challenges the notion of AI accuracy as a purely objective metric by analyzing the context-dependent normative choices required to define and measure it, using the EU AI Act to demonstrate how these techno-normative decisions shape risk distribution and offer practical guidance for regulators and developers.

Original authors: Lucas G. Uberti-Bona Marin, Bram Rijsbosch, Kristof Meding, Gerasimos Spanakis, Gijs van Dijck, Konrad Kollnig

Published 2026-04-07
📖 5 min read🧠 Deep dive

Original authors: Lucas G. Uberti-Bona Marin, Bram Rijsbosch, Kristof Meding, Gerasimos Spanakis, Gijs van Dijck, Konrad Kollnig

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a judge in a high-stakes cooking competition. The rulebook (the EU AI Act) says: "The dish must be accurate enough to be served to the public."

Now, imagine a developer brings you a soup and says, "Look! My soup is 99.8% accurate!"

You might think, "Wow, that's perfect!" But then you realize: the soup is made of 99.8% water and 0.2% salt. It's technically "accurate" at being water, but it's a terrible soup. If you serve this to someone with a salt deficiency, they could get sick.

This paper argues that asking "Is the AI accurate enough?" is a trick question. Accuracy isn't just a number on a scoreboard; it's a series of hidden choices about what kind of mistakes we are willing to accept.

Here is the breakdown of the paper's four "secret choices" using simple analogies:

1. Choosing the Scorecard (Selecting Metrics)

Imagine you are testing a new Melanoma (skin cancer) detector.

  • The Trap: If you use a simple "Accuracy" score, the AI could just say "No cancer" to everyone. Since most moles are harmless, it would be right 99% of the time. But it would miss every single cancer case. That's a disaster.
  • The Choice: The developer has to decide: What matters more?
    • Option A: Catch every single cancer (even if it means screaming "Cancer!" at a harmless mole). This is called Recall.
    • Option B: Only scream "Cancer!" when you are absolutely sure (even if it means missing a few real cases). This is called Precision.
  • The Paper's Point: There is no "right" scorecard. Choosing one over the other is a moral decision. Are we okay with scaring healthy people (False Positives) or risking the lives of sick people (False Negatives)? The AI Act demands we justify which scorecard we picked.

2. Mixing the Ingredients (Balancing Metrics)

Now, let's say you want both high Precision and high Recall. You can't have 100% of both.

  • The Trap: Developers often mix these two numbers into a single "F-Score" (like blending flour and sugar into a cake). Once blended, it's hard to taste how much sugar or flour is actually in there.
  • The Choice: Do we blend them into one number to make it look simple? Or do we keep them separate and say, "We need 90% Precision AND 80% Recall"?
  • The Paper's Point: Blending them hides the trade-offs. If you blend them, you might accidentally decide that catching cancer is less important than not scaring people, without ever admitting it. The paper suggests we should keep the ingredients separate so we can see exactly what we are prioritizing.

3. Picking the Tasting Panel (Measuring Metrics)

You have your recipe, but how do you test it?

  • The Trap: Imagine you test your soup only on people who love salty food. The soup tastes great to them! But what about people who are sensitive to salt?
  • The Choice: In AI, this is about Data. If you test a skin cancer AI only on photos of fair skin, it will fail miserably on dark skin.
  • The Paper's Point: The developer has to decide: Who is in our test group? Do we test on everyone equally? Do we make sure we have enough photos of elderly people, or people with different skin tones? If the test group doesn't look like the real world, the "accuracy" score is a lie.

4. Setting the Pass/Fail Line (Determining Thresholds)

Finally, you have your scores. Is the soup good enough to serve?

  • The Trap: Is a 90% score good? What about 95%?
  • The Choice: This is the Acceptance Threshold.
    • If the AI is replacing a doctor, it needs to be better than a human.
    • If the AI is just a "second opinion" helper, maybe it just needs to be as good as a human.
  • The Paper's Point: The EU AI Act says the system must be "accurate enough," but it doesn't give a specific number (like "95%"). The developer has to draw the line in the sand. How much risk are we willing to take? If the line is too low, people get hurt. If it's too high, we never deploy life-saving technology.

The Big Picture

The paper concludes that accuracy is not a technical fact; it is a political and ethical choice.

The EU AI Act is trying to fix this by saying: "You can't just say 'it's accurate.' You have to write down exactly how you chose your scorecard, how you balanced the ingredients, who you tested on, and where you drew the pass/fail line."

Why does this matter?
Because if we don't make these choices carefully, we might build AI that is "accurate" on paper but dangerous in real life. The paper urges developers, lawyers, and regulators to stop hiding behind numbers and start having honest conversations about what kind of mistakes we are willing to accept.

In short: Don't just ask, "Is the AI accurate?" Ask, "Accurate at what, for whom, and at what cost?"

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →