← Latest papers
🤖 AI

TrustX Agent Risk Classification Framework (ARC): Risk-Tiering Internally Created Agentic AI Systems

This paper introduces the TrustX Agent Risk Classification Framework (ARC), a structured instrument that utilizes a twelve-dimension scoring rubric and existing autonomy models to classify seven types of agentic AI systems into three risk tiers with corresponding governance controls, specifically designed for practitioners, developers, and regulators to manage the unique risks of internally created agentic AI.

Original authors: Hannah M. Liu, Rhea Saxena, Shiv Asthana

Published 2026-07-13
📖 6 min read🧠 Deep dive

Original authors: Hannah M. Liu, Rhea Saxena, Shiv Asthana

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the world of Artificial Intelligence as a bustling city. For a long time, the "AI citizens" were mostly like helpful librarians or calculators: they sat on a shelf, waited for a question, and gave an answer. But recently, a new type of citizen has arrived: Agentic AI. These aren't just librarians; they are like tiny, super-smart interns who can not only find information but also do things—book flights, write code, move robots, or even manage bank accounts. They can plan, act, and keep going without a human holding their hand every second.

The problem? The city's old rulebooks (like the NIST framework or the EU AI Act) were written for the librarians. They don't know how to handle these new, energetic interns who might accidentally delete a file, steal a secret, or start a chain reaction of mistakes.

Enter Hannah Liu, Rhea Saxena, and Shiv Asthana from the Responsible AI Institute. They've built a new tool called the TrustX Agent Risk Classification Framework (ARC). Think of ARC as a high-tech "Risk Radar" designed specifically to scan these new AI interns and sort them into three safety zones: Low, Medium, and High.

The 12-Point Safety Scan

To figure out how dangerous an AI intern might be, ARC doesn't just guess. It runs a 12-dimension scoring rubric. Imagine a checklist with 12 questions, like:

  • Autonomy: Does the AI need a human to say "go" for every single move, or does it just run wild on its own?
  • Blast Radius: If this AI messes up, does it just annoy one person, or does it crash the whole company's network?
  • Reversibility: If the AI breaks something, can you hit "undo," or is the damage permanent?
  • Data Sensitivity: Is the AI reading public weather reports, or is it digging through top-secret "crown jewel" files?

Each question gets a score of 1 (Low), 2 (Medium), or 3 (High).

The "Critical Dimension" Rule: No Hiding the Danger

Here is the most important part of their discovery, and it's a bit like a fire alarm. In many risk systems, if you have one huge problem but ten tiny ones, the system might average them out and say, "Eh, it's mostly fine."

The authors of this paper explicitly reject that idea. They argue that you cannot average away a disaster. Their framework uses a "critical dimension" approach. This means if any single one of those 12 questions gets a "3" (High Risk), the whole system gets bumped up to the highest danger zone, no matter how safe the other 11 questions are.

For example, in their illustrative examples, they looked at a "Decision Support System" (like a medical AI that helps doctors). Even though most of its scores were low (1s), it had one "3" because it handled sensitive, regulated medical data. Because of that single "3," the framework suggests it must be treated as a Tier 3: High Risk system. A simple average would have hidden that danger, but ARC catches it.

The Seven Types of AI Interns

The framework sorts these agents into seven categories, and the results are surprising:

  1. Autonomous Agents: These are the "bosses." They set goals and do everything themselves. In the paper's simulation, these scored an average of 2.33 and were classified as Tier 3 (High Risk). They are like a self-driving car that never stops to ask for directions.
  2. Coding Assistants: These are the "programmers." The paper notes a special problem here: because they write code that runs on computers, they need a special "extension" to the rules. Even a simple "autocomplete" tool (which just suggests code) was scored as Tier 2 (Medium Risk) because it handles internal, confidential code.
  3. Decision Support Systems: As mentioned, even if they mostly just talk, if they touch private data, they jump to Tier 3.
  4. AI Embedded/Physical Agents: These are robots and self-driving cars. Because they can hurt people or break physical things, and you can't always "undo" a car crash, they are Tier 3.
  5. Knowledge Assistants: These are the "researchers" (like chatbots). Usually, they are Tier 2, but if they remember too much or touch secret data, they can become Tier 3.
  6. Tool-Using Agents: These are the "connectors" that talk to other apps. Because they can accidentally leak data to the outside world, they are Tier 3.
  7. Transaction/Commerce Agents: These handle money. Surprisingly, in the paper's example, they landed on Tier 2 (Medium Risk) because they usually still need human approval. However, the authors warn that if these agents get more freedom, they could easily become High Risk.

What the Paper Rules Out

The authors are very clear about what their tool is not.

  • It is not a magic wand that solves all AI problems.
  • It is not a final, unchangeable law. The paper admits the framework is a "structured, repeatable instrument" that will need to be updated as AI gets smarter.
  • It is not a replacement for human judgment. The scores are based on how the user answers the questions, and the authors acknowledge that "scoring subjectivity" is a limitation.
  • It does not claim that high-risk AI should never be built. Instead, it argues that if you build high-risk AI, you need "rigorous" controls (like third-party validation and board-level approval) to keep it safe.

How Sure Are They?

The paper presents this as a proposed framework based on existing rules and real-world incidents (like the "IDEsaster" of 2025 where coding tools had vulnerabilities). They use illustrative examples and simulated scoring to show how it works. They don't claim to have tested this on every AI in the world yet; rather, they suggest that this method is a necessary step forward because current rules are "outpaced" by how fast AI is growing.

The Takeaway

The authors believe that by using this "Risk Radar," companies and regulators can stop guessing. Instead of hoping their AI interns are safe, they can run them through the 12-point scan. If the AI gets a "3" in any category, the lights go red, and strict rules kick in. It's a way to make sure that as AI gets more powerful, our safety nets get stronger, too.

As the paper concludes, this framework is meant to be a "practical stepping stone." It's not the final destination, but it's a much better map than the ones we had before.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →