← Latest papers
🤖 AI

Position: Fairness Failure in Generative Models is an Evaluation Problem

This position paper argues that fairness failures in generative models are primarily an evaluation problem caused by non-comparable and non-actionable findings, and proposes "Fairness Cards" as a standardized reporting artifact to enable reproducibility, comparability, and accountability in bias assessment.

Original authors: Mariia Vladimirova, Jean-Yves Franceschi, Thibaut Issenhuth

Published 2026-08-19
📖 7 min read🧠 Deep dive

Original authors: Mariia Vladimirova, Jean-Yves Franceschi, Thibaut Issenhuth

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the last decade, computers have learned to create art, write stories, and hold conversations with a fluency that once seemed impossible. These generative systems, which can produce images, text, and video from simple instructions, are now woven into the fabric of daily life. Yet, as these tools become more powerful, a persistent shadow follows them: the tendency to reproduce and amplify societal inequalities. When a computer generates an image of a doctor, it often defaults to a man; when it writes a story about a leader, it frequently chooses a specific gender or race. This is not merely a glitch in the code but a reflection of the data the machines learned from, a mirror held up to human history's biases. For years, researchers and regulators have tried to fix this, developing methods to detect and reduce these unfair patterns. The goal has always been to ensure that these powerful tools treat all people with equal dignity and accuracy.

However, a new perspective suggests that the problem might not be the lack of better fixes, but rather the way we measure whether the fixes work. A team of researchers argues that the field is stuck in a cycle of confusion because the tests used to judge fairness are too flexible. Just as a ruler made of rubber would give different measurements depending on how you stretch it, the current methods for testing AI fairness change their results based on tiny, often unreported choices made by the researchers. One study might use a specific way of asking a question and conclude the AI is fair, while another study using a slightly different phrasing might conclude the same AI is deeply biased. Both studies could be technically correct within their own narrow rules, yet they tell opposite stories to the public and to policymakers. This inconsistency makes it impossible to know if the technology is actually improving or if we are just seeing different angles of the same problem.

The researchers behind this new work, Mariia Vladimirova and her colleagues, set out to diagnose why fairness in generative models remains such a stubborn issue. They did not simply propose a new algorithm to make AI more polite; instead, they examined the very process of evaluation itself. Through a series of controlled experiments, they demonstrated that the verdict on whether an AI is fair or unfair often depends entirely on the "protocol" used to test it. In one striking demonstration, they took a single, unchanging AI model and tested it against a grid of different scenarios. They varied the way questions were asked, the style of the answers expected, and the settings used to generate the text. They found that the same model could be judged as having a severe fairness problem under one set of conditions and as being perfectly balanced under another.

For instance, when the researchers asked the model to write a story about a person in a specific job, the results were heavily skewed toward stereotypes. But when they asked the same model to write a professional memo about that same person, the bias vanished. In another test, they changed the way the computer selected its words, a setting known as the decoding regime. A slight shift in this setting, which is often left to chance or default settings in other studies, was enough to flip the conclusion from "biased" to "unbiased." The researchers found that these variations were not random noise but systematic shifts that could hide or reveal harm depending on the choices made by the evaluator. This means that a company could release a model with a report claiming it is fair, simply because they chose a specific testing method that happened to show good results, while a different method would have revealed significant problems.

The paper argues that this lack of standardization is the true bottleneck. It is not that we do not know how to build fairer models, but that we cannot reliably tell if we have succeeded. The current landscape is filled with "ad-hoc" checks, where researchers pick a few examples to test and report the results. This approach allows for what the authors call "cherry-picking," where a system looks good because the test was designed to make it look good, or where a genuine improvement is missed because the test was too narrow. Furthermore, the researchers highlighted that modern AI systems often include safety layers that refuse to answer certain questions. If a model refuses to answer a question for one group of people but answers it for another, that is a form of unfairness called an access disparity. Yet, many current tests simply ignore these refusals or treat them as errors to be discarded, effectively hiding the fact that some people are being silenced while others are heard.

To solve this, the team proposes a new standard called a "Fairness Card." Imagine a nutrition label on a food package, but for AI systems. Just as a nutrition label tells you exactly what is in the food, how much sugar it contains, and how it was processed, a Fairness Card would list every single detail about how the AI was tested. It would specify exactly what questions were asked, how many times the test was run, what settings were used to generate the answers, and how the researchers handled refusals or safety blocks. This document would not dictate what "fair" means, but it would make the process of measuring fairness transparent and reproducible. If two different researchers want to compare their results, they would know exactly how to replicate the test, ensuring that they are comparing apples to apples rather than apples to oranges.

The researchers tested this idea by creating a detailed Fairness Card for their own experiment. They documented every variable, from the specific words used in the prompts to the random seeds that influenced the computer's choices. When they published these details, they showed that the same model could produce a wide range of fairness scores depending on which part of the card you looked at. Under one prompt style, the model showed a high rate of stereotypical language; under another, it showed almost none. By laying all these numbers out in a single, standardized format, they proved that the "truth" about the model's fairness is not a single number, but a complex picture that only becomes clear when all the variables are visible.

This approach shifts the focus from trying to find a perfect, one-size-fits-all definition of fairness to building a system where the evaluation itself is honest and complete. The authors acknowledge that this will not solve every problem. They note that some companies may still hide their data, and that different cultures may have different ideas about what is fair. However, they argue that without a shared, detailed record of how a system was tested, no progress can be made. It is impossible to improve what you cannot measure consistently, and it is impossible to compare solutions if the rules of the game keep changing.

The paper concludes with a call to action for researchers, developers, and regulators. They urge that any time a new AI system is released or a claim is made about its fairness, a Fairness Card must be provided. This document should be treated as a mandatory part of the release, just like a safety manual or a user guide. It should detail the specific groups of people tested, the exact methods used to ask questions, and the full range of results, including the worst-case scenarios. By making these choices explicit, the field can move away from vague promises and toward a future where fairness is a measurable, trackable, and accountable part of how these powerful tools are built and used. The goal is not to eliminate all bias, which may be impossible, but to ensure that we are all looking at the same evidence when we decide whether a system is safe to use.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →