Evidence for Whom? Contestability and Socio-Technical Assurance in High-Risk AI Governance
This paper argues that current high-risk AI governance frameworks (NIST, EU AI Act, ISO/IEC 42001) create a critical evidentiary gap by prioritizing artifact-centered evidence over process-based evidence needed for contestation, a structural flaw highlighted by the inherent ambiguity in interpreting these requirements and which ultimately undermines the procedural and epistemic rights of governed stakeholders.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where a computer program decides who gets a job, who receives a loan, or whether someone is granted parole. These systems are not just code running in a vacuum; they are part of a larger human ecosystem involving the companies that build them, the organizations that use them, and the people whose lives are altered by their output. When these systems make a mistake, the question of who is responsible and how to prove it becomes urgent. This is the realm of artificial intelligence governance, a field dedicated to ensuring these powerful tools are safe, fair, and accountable. A central idea in this field is "contestability," the ability for a person affected by a decision to question it, understand why it happened, and seek a correction. For this to work, there must be evidence—proof that the system was checked, that it works as intended, and that a human could step in if things went wrong. But a critical question remains: who gets to see this evidence, and is the evidence actually useful for the person who needs it most?
A team of researchers from universities in the United States and India set out to investigate this exact problem. They examined three major sets of rules and guidelines that govern high-risk artificial intelligence: a voluntary framework from the United States, a binding law from the European Union, and an international standard for managing AI systems. Their goal was to create a detailed map, or "crosswalk," that connects specific requirements in these documents to the actual technical methods used to test AI. They wanted to see if the evidence required by the rules was the kind of evidence that could actually help a person contest a decision, or if it was merely a collection of technical reports that only the companies and regulators could read.
The researchers began by gathering a vast collection of technical evaluation methods. They identified twelve distinct families of tests, ranging from checking how accurate a model is on new data to seeing if it can be tricked by malicious inputs, and from analyzing the quality of the data used to train it to studying how well humans can supervise the system. They then took these twelve families and compared them against twenty-two specific requirements found in the three governing documents. They asked a simple but profound question for each pairing: does this specific test provide the main proof needed to satisfy this specific rule?
What they found was a striking imbalance. The rules that asked for evidence about the AI model itself—such as its accuracy, its robustness against attacks, or its fairness in statistical terms—were well-supported. There were plenty of technical tests available to generate the necessary reports. However, the rules that asked for evidence about how the system works in the real world, after it has been deployed, were largely unsupported. These rules concerned things like whether a human overseer could actually catch an error, whether the system's logs were detailed enough to reconstruct a specific decision, and whether the organization could prove they were monitoring the system's impact on people's lives. For these crucial questions, the researchers found that the available technical methods were either missing entirely or provided only weak, secondary support.
The study revealed that the evidence generated to prove an AI system is safe is primarily designed for intermediaries: the companies that build the software, the auditors who check it, and the government regulators who enforce the law. The people who are actually subject to these decisions—the job applicant who was rejected, the borrower who was denied—have no direct route to this evidence. They might receive an explanation of what the system did, but they cannot access the underlying records that would allow them to verify if that explanation is true or to understand exactly how the decision was reached. The researchers argue that this is not just a technical gap or a missing piece of a puzzle; it is a fundamental flaw in the system of justice. By denying affected people access to the evidence needed to challenge a decision, the current framework treats them as passive subjects rather than active participants in a process that determines their futures.
The researchers were careful to note that their findings are based on a specific analysis of the available technical literature and the text of the laws. They did not claim that companies are not doing these checks internally, but rather that the methods to prove these checks are effective are not yet standardized or widely available in a form that satisfies the rules. They also highlighted that their own work involved a degree of interpretation. When they asked independent experts to review their mapping, the experts agreed on the broad strokes but sometimes disagreed on the finer details, suggesting that the rules themselves are open to different reasonable readings. This disagreement, the authors suggest, is a feature of the problem, not a bug: it shows that the current requirements are not clear enough to guarantee that the right kind of evidence is being produced.
Ultimately, the paper concludes that the current approach to AI governance is incomplete. It relies heavily on evidence that can be produced inside a company's development pipeline, such as test results and technical documentation, while neglecting the evidence that can only be gathered by observing how the system interacts with people in the real world. The authors propose that for AI to be truly accountable, the rules must evolve to require evidence that is not just about the machine, but about the entire socio-technical system. This includes proof that human oversight is effective, that logs are sufficient to reconstruct events, and that the system's impact on society is being monitored. Without these changes, the promise of AI accountability remains a promise made only to the powerful, leaving those most affected by the technology without the tools to defend themselves.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.