← Latest papers
🤖 AI

The Eticas AI Risk Taxonomy: Open Infrastructure for Operationalizing AI Audits

This paper introduces the Eticas AI Risk Taxonomy v2.0.0, an open infrastructure that bridges the gap between abstract risk catalogs and practical AI audits by providing a standardized, operational framework to measure, calibrate, and grade specific risks like PII leakage across multiple external compliance standards.

Original authors: Gemma Galdon Clavell, Pablo Accuosto, Usman Gohar

Published 2026-07-03
📖 5 min read🧠 Deep dive

Original authors: Gemma Galdon Clavell, Pablo Accuosto, Usman Gohar

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to build a house. You have a list of 74 different "Risk Checklists" from various architects, engineers, and government agencies. One says, "Check for leaks." Another says, "Check for structural cracks." A third says, "Check for fire hazards."

The problem, according to this paper, is that all these lists stop at naming the problem. They tell you what to look for, but they don't tell you how to look for it, what tool to use, or how to decide if the house is "safe enough" to live in. They are like dictionaries that define words but don't teach you how to speak.

This paper introduces a new tool called the Eticas AI Risk Taxonomy. Think of it not just as a dictionary, but as a complete construction manual that bridges the gap between "naming a risk" and "actually testing for it."

Here is how the paper breaks it down, using simple analogies:

1. The Core Problem: The "Glossary" Trap

Most AI risk lists are like a glossary. They say, "Privacy Risk: The danger of personal data leaking." That's useful for talking about the problem, but it's useless for an auditor trying to test a real AI system.

  • The Paper's Claim: The hard part of auditing isn't naming the risk; it's operationalizing it. It's turning the word "Privacy Risk" into a specific test, a number, and a final grade (like an A, B, C, D, or F).

2. The Solution: A "Bridge" from Concept to Grade

The authors built a bridge. They show exactly how to take a risk concept and walk it all the way to a graded result.

  • The Analogy: Imagine a doctor diagnosing a patient.
    • Old Way: The doctor says, "The patient has a 'fever risk'." (End of report).
    • Eticas Way: The doctor says, "The patient has a 'fever risk.' We measured their temperature at 104°F using a specific thermometer (the test). This falls into the 'Severe' band (the scale). Therefore, the patient gets a Grade 'E' (the diagnosis)."
  • The Example: They tested a famous AI model (GPT-4) on a specific risk: PII Leakage (accidentally revealing private info like email addresses).
    • They didn't just say "It might leak."
    • They ran specific tests (probes) where they tried to trick the AI.
    • Result: When the AI was tricked once, it leaked 51% of the time. When tricked three times, it leaked 84% of the time.
    • The Grade: Because it leaked so much under pressure, they gave it a severe grade of E (the worst rating) and flagged it as a "Systemic" problem (meaning it's a deep flaw, not a one-time glitch).

3. The Secret Sauce: Separating the "What" from the "How"

This is the most important technical part, explained simply:

  • The Risk (The "What"): This is the abstract danger, like "Data Leakage."
  • The Mechanism (The "How"): This is how the danger happens. For data leakage, it could happen because the AI reveals data you typed in, or because it memorized data from its training and spits it back out.
  • The Innovation: The paper separates these two.
    • Analogy: Think of a car. The "Risk" is a "Flat Tire." The "Mechanisms" are "Punctured by a nail" vs. "Worn out tread."
    • You need different tools to fix a nail than you do to fix worn tread.
    • By separating them, the Eticas taxonomy allows different auditors to use different tests (tools) on the same risk, as long as they agree on the "Mechanism." This makes the system flexible and open.

4. The "Open Core" Model

The authors are offering a "Open Core" model, which is like a Lego set.

  • The Open Part (The Baseplate): The categories, definitions, and the "Mechanism" labels are free for everyone to use (under a Creative Commons license). This is the shared language.
  • The Proprietary Part (The Bricks): The specific tests, the exact numbers, and the detailed scoring rules are kept by the authors (Eticas) as their "practitioner layer."
  • Why this matters: Anyone can build on the open baseplate. A regulator, a university, or a company can take the "Data Leakage" concept and build their own specific tests on top of it, knowing they are speaking the same language as everyone else.

5. The "Agentic AI" Surprise

The paper highlights a new category called Agentic AI.

  • The Analogy: Most AI today is like a calculator: You press a button, it gives an answer.
  • Agentic AI is like a robot butler: It can plan, use tools, and act on its own.
  • The Gap: The paper points out that all the major laws and safety rules (like the EU AI Act) were written for the "calculator" era. They don't really know how to handle the "robot butler."
  • The Finding: By creating a specific category for Agentic AI, the paper exposes a "Governance Gap." It shows that current rules are outdated for this new type of AI, and we need new tools to audit them.

Summary

The paper argues that the field of AI auditing is currently stuck with too many dictionaries and not enough tools. The Eticas AI Risk Taxonomy is a new infrastructure that:

  1. Standardizes the language (so everyone agrees on what "Risk" means).
  2. Operationalizes the testing (showing exactly how to measure it).
  3. Grades the results (giving a clear pass/fail score).
  4. Exposes Gaps (showing where current laws fail to cover new AI technologies like Agentic AI).

It's not just a list of problems; it's a blueprint for solving them.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →