The AI Resilience Gap: Bringing Artificial Intelligence Inside the Operational Resilience Perimeter
This paper argues that current AI governance frameworks focused on trustworthiness fail to address operational resilience, and proposes the "AI Resilience Framework" to integrate AI dependencies into operational continuity planning through dependency mapping, substitutability tiering, and concentration management.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a financial company as a busy, high-stakes restaurant. For years, the regulators (the health inspectors) have been very focused on Trustworthiness. They check: Is the food safe? Is the chef fair? Are the recipes documented? Is the kitchen clean? This is the world of "Trustworthy AI."
But this paper argues that being "safe and fair" isn't enough. There is a second, equally important set of rules about Operational Resilience. This asks: If the power goes out, or the main supplier runs out of flour, can the restaurant still serve customers?
The author, Jonathan Shelby, says that while companies are getting very good at making their AI "safe and fair," they are failing to make sure their AI can survive a disaster. They have built a "Trustworthy" kitchen, but they haven't checked if the restaurant can keep cooking if the stove breaks.
Here is the breakdown of the paper's argument using simple analogies:
1. The Two Different Checklists
The paper says there are two separate rulebooks that companies are trying to follow, but they aren't talking to each other.
- The "Trustworthy" Checklist (The Health Inspector): This looks at the AI itself. Is it biased? Is it lying? Is it dangerous? If the AI is perfect, this checklist says "Pass."
- The "Resilience" Checklist (The Fire Marshal): This looks at the service. If the AI stops working, does the business stop? Can you switch to a backup plan? If the AI is perfect but you have no backup, this checklist says "Fail."
The Gap: A company can have an AI that is 100% "Trustworthy" (safe, fair, documented) but 0% "Resilient" (if it breaks, the whole business collapses). The paper calls this the AI Resilience Gap.
2. Why AI is a Special Kind of Breakage
The paper explains that AI breaks in weird ways that old safety rules didn't expect.
- The "Silent Drift" (The Grey Failure): Imagine a GPS app. Usually, if it breaks, the screen goes black (a clear failure). But AI is different. It might keep giving you directions, but slowly, the directions get worse and worse. It's still "on," but it's leading you into a ditch. The old rules only check if the screen is "on," so they miss this slow, silent disaster.
- The "Monoculture" Problem: Imagine every restaurant in a city buys their flour from the exact same giant mill. If that mill has a fire, every restaurant closes at the same time. The paper warns that everyone is using the same few "Frontier AI" models. If one of those big models fails, the whole financial system could stumble together.
3. The Solution: The "AI Resilience Framework"
The paper proposes a new 5-step method to fix this. Think of it as a way to audit your restaurant's backup plans.
- Step 1: Map the Ingredients. You need to know exactly which AI tools are running your "Important Business Services" (like taking orders or checking credit). You can't fix what you can't see.
- Step 2: The "Can You Swap It?" Test. The paper introduces a Criticality-Substitutability Matrix.
- High Criticality + No Swap: Danger Zone. (e.g., The only chef who knows the secret recipe, and if they quit, the restaurant closes).
- High Criticality + Swap Available: Managed. (e.g., The chef quits, but you have a trained sous-chef ready to step in).
- Low Criticality: Light Touch. (e.g., The AI just picks the playlist; if it breaks, no big deal).
- Step 3: Redefine "Broken." You can't just say "The AI is down." You must also say, "The AI is giving wrong answers." You need to set a limit: "If the AI is wrong more than 5% of the time, we treat it as broken and switch to the backup."
- Step 4: The "Real" Backup Doctrine. This is the most important part. Many companies say, "If the AI fails, a human will take over." But the paper says: If you haven't practiced that, it's not a backup; it's a fantasy. If the human process was deleted years ago to save money, you have no backup. You must keep the "human path" alive and practice switching to it.
- Step 5: Watch the Big Suppliers. You need to check if you are relying too heavily on one giant AI provider. If they are the "Mill" that supplies everyone, you need a plan to switch to a different mill if they fail.
4. What This Means for Leaders
The paper tells security bosses and company boards:
- Stop waiting for new rules. The regulators (like the Bank of England) aren't writing new "AI Safety" laws. They are saying, "You already have to be resilient. Now apply those rules to AI."
- Don't just trust the AI; trust your backup. Being "safe" isn't enough. You need to prove you can survive if the AI goes silent, drifts, or disappears.
- Connect the dots. The people checking for "fairness" (Model Risk) and the people checking for "survival" (Resilience) need to talk. They are looking at the same AI but asking different questions.
In short: The paper argues that we are currently building AI that is "good" but fragile. The goal is to build AI that is not only "good" but also "tough," with real, practiced plans for when things go wrong.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.