← Latest papers
💻 computer science

Toward Comprehensive Risk Assessments and Assurance of AI-Based Systems

This paper critiques the insufficient adaptation of traditional safety and security methodologies for AI-based systems and proposes a novel end-to-end risk framework that integrates Operational Design Domains (ODD) to establish a consistent assurance terminology and a concrete operational envelope for more effective risk assessment and mitigation.

Original authors: Heidy Khlaaf

Published 2026-06-30
📖 5 min read🧠 Deep dive

Original authors: Heidy Khlaaf

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Picture: Why We Need a New Rulebook

Imagine the world of Artificial Intelligence (AI) is like a sudden explosion of new, super-powerful cars hitting the road. Everyone is excited, but these cars are driving in ways we don't fully understand. Some are giving weird directions, others are saying rude things, and nobody has a clear map of where they might crash.

The author, Heidy Khlaaf, argues that we are trying to test these new "AI cars" using old rulebooks designed for regular cars, computers, and even hardware parts. The problem? AI isn't like those things. It's too complex, too unpredictable, and the old tests don't catch the real dangers.

This paper proposes a new, better way to check if AI is safe before we let it loose on the public.


1. The Confusion: "Alignment" vs. "Safety"

The Analogy: Imagine you hire a very obedient robot butler.

  • Value Alignment: You tell the robot, "Be nice to everyone." The robot follows this rule perfectly. It is "aligned" with your values.
  • Safety: However, the robot decides the best way to be "nice" is to lock everyone in the house so they can't get hurt by the outside world. It followed your instruction (alignment), but it caused a disaster (unsafe).

The Paper's Point:
The AI community often confuses these two. They think if an AI does what it's told (is aligned), it must be safe. Khlaaf says no. Safety isn't just about following orders; it's about making sure the system doesn't hurt people, even if it's trying to do exactly what you asked. We need to check for harm, not just check if the robot is "obedient."

2. The Mistake: Using the Wrong Tools

The paper says people are trying to fix AI problems using tools designed for other industries. Here is why that doesn't work:

  • Hardware Safety (The "Random Breakage" Test):
    • Old Way: Engineers test toaster parts. If a toaster breaks, it's usually because a wire randomly snapped due to wear and tear. You can predict this by counting how many toasters break over time.
    • The AI Problem: AI doesn't break randomly. It breaks because of bad design or confusing instructions. It's like a toaster that decides to burn bread because it misunderstood the word "toast." You can't predict this by counting broken wires; you have to understand the recipe.
  • Cybersecurity (The "Hacker" Test):
    • Old Way: Security experts ask, "Can a bad guy break in and steal our data?" They focus on protecting the system from outside enemies.
    • The AI Problem: The danger isn't always a hacker. The danger is the AI itself doing something harmful by accident. Asking "Can a hacker break this?" doesn't answer "Will this AI accidentally fire a gun at a crowd?" We need to test the AI's behavior, not just its locks.
  • Software Safety (The "Code Check" Test):
    • Old Way: Programmers check code line-by-line to make sure it follows strict rules.
    • The AI Problem: AI learns on its own. You can check the code that teaches the AI, but you can't check the code that is the AI, because the AI changes its own "brain" based on what it learns. It's like trying to write a rulebook for a student who invents new math problems every day.

3. The Solution: The "Operational Design Domain" (ODD)

Since we can't test AI for everything (because there are too many things it could do), the paper suggests we define exactly where and how the AI is allowed to work.

The Analogy: The Driving License
Imagine a driver's license. You don't get a license to drive anywhere, anytime.

  • You might have a license to drive a car on highways in good weather.
  • You do not have a license to drive a tank in a war zone or a boat in a storm.

The paper calls this the Operational Design Domain (ODD). It's a "safety envelope" or a "fence" around the AI.

How the New Framework Works:
Instead of trying to test the AI for every possible scenario in the universe, we define the fence first. The paper suggests a checklist (a taxonomy) to draw this fence:

  1. Where is it used? (Is it in a hospital, a newsroom, or a factory?)
  2. Who is touching it? (Is it a doctor, a child, or a data entry clerk?)
  3. How does it connect? (Is it talking to a human, a database, or a robot arm?)
  4. Who could get hurt? (Are we protecting specific groups of people based on race, age, or gender?)
  5. What are we protecting? (Is it money, private data, or physical safety?)

4. Putting It All Together

The paper proposes a new process for developers and auditors:

  1. Draw the Fence: Clearly define the ODD. "This AI is only for writing marketing emails for small businesses."
  2. Test Inside the Fence: Check if the AI is safe only within that specific context.
  3. Check the Edges: See what happens if the AI is pushed against the fence (e.g., What if it tries to write a medical diagnosis instead of an email?).
  4. Fix the Gaps: If the AI acts dangerously near the fence, you either fix the AI or make the fence smaller (restrict its use).

The Bottom Line

We cannot treat AI like a toaster, a computer virus, or a standard software program. It is a new kind of system that learns and changes.

To keep people safe, we must stop trying to test AI for "everything" and start defining exactly where it is allowed to operate. By clearly drawing the boundaries (the ODD) and testing the AI strictly within those boundaries, we can finally know if an AI system is truly ready for the real world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →