Coverage and Complementarity of Three Agentic AI Risk Taxonomies Across 131 Real World Incidents
This paper empirically evaluates the coverage and complementarity of three dominant agentic AI risk taxonomies (OWASP-ASI, MSFT-AIRT, and NIST AI RMF) across 131 real-world incidents to provide evidence-based guidance for practitioners and identify specific gaps for future framework revisions.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where artificial intelligence has graduated from simply answering questions to actually taking action. These are no longer just chatbots that wait for a prompt; they are autonomous agents that can plan a sequence of steps, use digital tools, remember past interactions, and coordinate with other software to achieve a goal with little human help. This shift has created a new landscape of risk. When a simple chatbot makes a mistake, it might say something odd. But when an autonomous agent makes a mistake, it might delete a company's database, transfer money to the wrong account, or spread dangerous misinformation while acting as a trusted advisor. To manage these risks, experts have built three different "rulebooks" or taxonomies to categorize how these systems fail. One list focuses on the visible result of a failure, another on the specific mechanical way it happened, and a third on the stage of the project's life where the error occurred. For a long time, organizations had to guess which rulebook was the best to use, without any proof of how well they actually worked against real-world disasters.
A researcher set out to settle this uncertainty by treating these rulebooks like scientific instruments that needed calibration. Instead of building a new list of failures, they took the three most prominent frameworks currently in use and tested them against a collection of 131 real incidents involving autonomous agents. These incidents were drawn from a public database of documented AI failures, spanning from 2014 to 2026, and included everything from self-driving vehicle accidents to AI coding assistants that accidentally deleted production servers. The researcher carefully read the details of each event and tried to fit it into the categories provided by each of the three frameworks independently. The goal was to see if the frameworks could catch the failures, if they agreed with each other on what happened, and where they might be missing the mark entirely.
The study found that while no single rulebook caught every single incident, using all three together covered nearly 97 percent of the documented failures. This high level of combined coverage suggests that the three frameworks are not competing with each other, but rather complementing one another. They simply look at the same problem from different angles. One framework, designed for security executives, focuses on the outcome: what was the visible damage? Another, built by a team of security engineers, focuses on the mechanism: exactly how did the system break? The third, a government standard, focuses on the process: at what stage of the project's life should a manager have intervened to stop this? The research showed that these perspectives are distinct. For instance, a single incident where an AI agent was tricked into stealing data might be described by the first framework as a "trust exploitation" failure, by the second as a "prompt injection" attack, and by the third as a failure in the "management" phase of the project lifecycle. None of these descriptions is wrong; they are just answering different questions.
However, the study also uncovered specific gaps where the current rulebooks fall short. About 20 percent of the incidents could not be cleanly assigned to any category in the security-focused framework. These missing cases tended to cluster into three clear themes. The first involved agents that acted as advisors but gave confidently wrong information, such as a chatbot fabricating a sexual harassment allegation against a real professor or a financial bot giving dangerous investment advice. The second theme involved humans misusing the technology, like students using AI to gain an unfair advantage on exams or scammers using it to impersonate others, where the AI itself worked as designed but the human user caused the harm. The third theme involved broader social failures, such as a browser extension that harvested private AI conversations or a company using AI surveillance in a way that threatened civil rights. In these cases, the failure was not in the code itself, but in how the technology was embedded in society or used by people.
The researcher concluded that the field does not need a fourth or fifth rulebook to solve these problems. Instead, the solution lies in understanding that the existing tools are built for different jobs. Security leaders communicating with a board of directors should use the outcome-focused list to explain the business impact. Technical teams investigating a breach should use the mechanism-focused list to find the root cause. Regulators and auditors should use the process-focused list to check if the right safeguards were in place at the right time. The study also suggested that the creators of these frameworks could improve them by adding specific categories for the identified gaps, such as a dedicated section for "hallucinations in advisory roles" or a clearer distinction between agent failures and human misuse. By mapping these tools to the specific needs of different professionals, the research provides a practical guide for navigating the complex world of autonomous AI safety, turning a confusing proliferation of options into a clear, evidence-based strategy.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.