Same System, Opposite Verdicts: Metric Discretion in AI Ethics Audits and the Limits of Disclosure
This paper demonstrates that the lack of standardized fairness metrics in AI ethics allows for arbitrary and contradictory audit outcomes, arguing that implementing disclosure requirements similar to financial alternative performance measures—rather than mandating a single correct metric—can significantly reduce this discretion and improve transparency.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a judge in a courtroom, but instead of deciding guilt or innocence, you are judging whether a computer program is being fair. In the world of Artificial Intelligence (AI), "fairness" isn't a single, simple number like a height or a weight. It's more like a recipe. You can make a cake that tastes sweet, or one that tastes sour, or one that is perfectly balanced, depending on which ingredients you choose and how much of each you add. In the same way, to measure if an AI is fair, you have to choose which mathematical recipe to use. You have to decide who counts as a "group" to compare, what score counts as a "pass," and which specific math formula to run.
For a long time, experts argued about which recipe was the "right" one. Some said, "We should look at how many people get hired," while others said, "No, we should look at how many people get rejected." This paper steps into that messy kitchen to see what happens when different chefs use different recipes on the exact same ingredients. The big question is: Does it matter which recipe you pick? The answer turns out to be a loud, resounding "Yes." If you change the recipe, you might flip the verdict from "This AI is fair" to "This AI is biased," even though the computer didn't change a single thing. This matters because governments and companies are now demanding these fairness scores to make big decisions about who gets a loan, a job, or even freedom in court. If the score depends entirely on the recipe you chose, and no one tells you which recipe was used, the score might be meaningless.
The Great AI Fairness Switcheroo
This paper is like a massive experiment where the author, Shay Tsaban, acts as a super-scrupulous auditor. They took four different AI systems—one real-world tool used in a Florida jail to predict if someone will re-offend, and three famous test datasets used by scientists to check for bias in income, credit, and jobs. Then, they didn't just run one test. They ran 91,572 different versions of the test.
Think of it like this: Imagine you have a giant box of Legos representing an AI system. The author built a machine that could snap together every possible combination of "fairness rules" to see what kind of tower it would build. They tried every defensible way to measure fairness: changing the math formula, changing the cutoff score for a "pass," changing which groups of people were compared, and even changing the specific year of data used.
The Shocking Result: The Verdict Flips
The most mind-blowing thing the paper found is that for every single system they tested, the result flipped back and forth. In all 7 of the specific cases they looked at (like "Black vs. White defendants" or "Men vs. Women"), the AI could be declared "Fair" under one set of rules and "Unfair" under another.
It gets even wilder. Sometimes, the group that was labeled "disadvantaged" (the ones being treated unfairly) would switch! Under one recipe, the AI was hurting Group A. Under a different, equally valid recipe, the AI was hurting Group B. The paper found that the numbers reported could swing wildly—6 to 30 times further than the normal "sampling error" (the tiny bit of wiggle room we expect just from random chance). In fact, the choice of which math formula to use was responsible for 45% of the difference in the final score, while random chance only moved the needle by a tiny fraction.
The "Fairwashing" Problem
The paper also looked at what companies and governments are actually doing right now. They read 6,797 public documents where organizations claimed to have audited their AI. The result? Almost no one is playing fair with the information.
- Only 38 of those nearly 7,000 documents actually gave a specific number for fairness.
- Of those 38, none explained why they picked that specific math recipe.
- None said what other recipes they considered and rejected.
- None compared the current score to a score from last year to see if things were getting better or worse.
It's like a chef telling you, "This cake is delicious," but refusing to tell you if they used sugar or salt, or if they baked it for 10 minutes or an hour. You can't tell if the cake is actually good or if they just got lucky with the recipe.
The Solution: The "Recipe Card" Rule
So, how do we fix this? The paper suggests borrowing a trick from the world of finance. In accounting, companies are allowed to report "adjusted earnings" (their own version of profit), but they must show a "reconciliation" table that explains exactly how they got that number from the standard profit.
The author ran a simulation to see what would happen if AI auditors had to do the same thing. They found that if auditors simply had to write down four things:
- Which math family they chose (the recipe).
- The specific rules they used (the ingredients).
- Who they compared (the diners).
- Whether they kept the rules the same over time (consistency).
...this simple disclosure would close 78% of the "wiggle room." It would make the results much more consistent and honest without needing to agree on which single recipe is the "correct" one. Even better, if they also reported how the score changed over time, it would close 82% of the wiggle room.
The Bottom Line
The paper concludes that we don't need to solve the impossible math problem of finding the one "perfect" definition of fairness to fix this mess. We just need to stop hiding the choices. If an AI system can be judged as "fair" or "unfair" just by changing the ruler you use to measure it, then the ruler itself must be part of the report.
Right now, the "residual" problem—the tiny bit of wiggle room that remains even after we force people to write down their choices—is about 0.16 on the scale. This is the limit of what transparency can fix. But the paper shows that the vast majority of the confusion comes from people not telling us which rules they played by. By forcing them to show their work, we can stop the "Same System, Opposite Verdicts" game and start getting real answers.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.