← Latest papers
💰 quantitative finance

Agentic AI Systems and Financial Stability, From Model Risk to Systemic Risk

This paper argues that the shift from predictive to agentic AI systems transforms financial risk from diversifiable idiosyncratic errors into non-diversifiable systemic threats, demonstrating through six mathematical frameworks that such risks cannot be mitigated by detection or post-hoc controls and require ex-ante structural prevention.

Original authors: Sriram Nagaraj, Seung Jung Lee

Published 2026-10-08
📖 7 min read🧠 Deep dive

Original authors: Sriram Nagaraj, Seung Jung Lee

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Financial stability relies on a comforting assumption: that when one part of the system breaks, the rest can absorb the shock. If a single bank fails, the others survive; if one trader makes a mistake, the market corrects itself. This works because failures are usually unique to that specific institution, scattered and unrelated. But a new kind of risk is emerging that shatters this safety net. It comes from "agentic" artificial intelligence—systems that do not just predict the future but take actions in the real world, such as moving money, writing code, or managing patient care. When many of these agents are built on the same underlying software foundation, a single error or a single hack does not stay local. Instead, it strikes every agent at once, turning a small glitch into a system-wide collapse.

Researchers at the Federal Reserve have spent years studying how to manage risk in traditional banking, where they measure the chance of a loan defaulting and set aside money to cover the loss. They have now applied this same rigorous thinking to these new AI systems, but they found that the old rules do not work. The core problem is that these AI agents are different from human workers or traditional software. They act with speed and autonomy, and when they make a mistake, the damage is often irreversible. You cannot "un-send" a fraudulent wire transfer or "un-delete" a critical database. Because the harm is permanent and happens instantly, the usual safety nets—like having a backup plan or waiting to see what happens before reacting—are useless. The researchers discovered that the only way to stop a catastrophe is to prevent it from being possible in the first place, by removing the dangerous capabilities before the system ever runs.

The study, which spans six distinct mathematical investigations, begins by looking at a single AI agent. In traditional finance, risk is calculated by multiplying the probability of an event by the size of the loss. The researchers adapted this for AI, breaking down harm into three parts: how likely the agent is to take a harmful action, how severe that action is, and how far its reach extends. They found that when an agent is under attack, these three factors do not act independently. A single malicious event can simultaneously make the agent more likely to act, make the damage worse, and expand the scope of the damage. This creates a "wrong-way" risk where the danger compounds rapidly. Most importantly, they proved that for irreversible harms, no amount of monitoring or ability to reverse mistakes can fix the problem. If an action cannot be undone, the only solution is to ensure the agent never has the ability to take that action at all.

The researchers then expanded their view from a single agent to a fleet of thousands, all running on the same foundation model. This is where the risk becomes truly systemic. In a normal portfolio, if you have many different investments, a failure in one does not hurt the others. But if every agent in a fleet shares the same brain, a flaw in that brain affects them all simultaneously. The researchers showed that no matter how many agents you add to the fleet, the risk does not go down. Instead, it hits a "floor" that cannot be diversified away. Even if the agents are in different companies or different countries, if they rely on the same underlying software, a single failure can bring down the entire group. This is what they call a "monoculture": a system where everyone is vulnerable to the exact same threat.

The study also examined how these agents interact with one another. In the real world, agents often talk to each other, sharing information or coordinating tasks. The researchers found that this interaction creates a new, hidden danger. Even if the agents are built on different software, if they are programmed to listen to each other, they can start to move in lockstep during a crisis. If one agent panics and changes its behavior, the others follow, creating a chain reaction that spreads the risk. This happens even without a shared software flaw. Furthermore, a compromised agent can use language to persuade its peers to do something harmful, creating a contagion that spreads through conversation rather than code. This means that simply separating the technical networks is not enough; the way agents communicate must also be controlled.

To understand how these risks play out over time, the researchers modeled the system as a continuous flow of events, similar to how earthquakes or financial crashes are studied. They found that the risk behaves like a rare, massive jump rather than a slow drift. The system can appear stable for a long time and then suddenly collapse. They also looked at how to control these systems while they are running, using "guardrails" to stop bad actions. They proved that for irreversible events, these runtime controls are fundamentally limited. If an action is already irreversible, waiting to detect it and stop it is too late. The probability of a disaster depends entirely on the initial design of the system, not on how well you watch it while it runs. No amount of monitoring can fix a flaw that was built in from the start.

The final chapters of the study introduced an adversary—a hacker or attacker who is also using advanced AI. The researchers treated this as a game where the attacker tries to find the weakest point and the defender tries to block it. They found that when attackers use AI, the system becomes a much more attractive target. A shared software foundation is like a single key that opens every door in a building; if a thief finds that key, they can enter everywhere at once. The study showed that in a decentralized market, where each company tries to protect itself individually, everyone will underinvest in the most important safety measure: structural prevention. They will spend money on detection and reaction because those are easier to see and measure, but they will ignore the hard work of removing the dangerous capabilities entirely. This leaves the whole system vulnerable.

The researchers concluded that the only effective way to manage this risk is through "structural prevention." This means designing the system so that catastrophic actions are physically impossible to perform, rather than hoping to catch them after they happen. It involves removing dangerous tools, limiting permissions, and ensuring that agents cannot access the parts of the system that could cause irreversible harm. The study proved that this approach is the only one that works against a growing, intelligent adversary. As AI technology improves and attackers become more capable, the cost of trying to detect and react to every new threat becomes infinite. In contrast, the cost of removing a dangerous capability is a one-time investment that protects the system forever, regardless of how smart the attacker becomes.

The paper ends with a clear message for regulators and system designers. Financial stability in the age of AI cannot be achieved by watching each institution more closely or by hoping that a single failure won't spread. The risk is too deep and too interconnected. The only way to secure the system is to make structural choices before the system ever starts running. This means limiting how much any single software foundation is relied upon, ensuring that agents cannot take irreversible actions, and designing the system so that a single point of failure cannot bring down the whole network. The mathematics is clear: you cannot diversify away a shared flaw, you cannot detect your way out of an irreversible mistake, and you cannot react fast enough to stop a coordinated attack. The only solution is to build the system so that the disaster never has a chance to happen.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →