Measuring agentic artificial intelligence intensity in banking: a validated disclosure-based index and identification strategy for estimating effects on bank performance and systemic risk
This paper introduces and validates the Agentic AI Intensity Index (AAII), a novel disclosure-based measurement instrument designed to distinguish autonomous AI execution from advisory systems in banking, while outlining a comprehensive identification strategy to estimate its effects on bank performance and systemic risk.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of banking, artificial intelligence has long been a helpful assistant. For years, these systems have acted as digital scribes, drafting reports, summarizing data, or suggesting the next best step for a human manager to consider. The human remains firmly in the driver's seat, holding the keys to every final decision. However, a significant shift is now underway. Banks are beginning to deploy a new kind of intelligence, one that does not just suggest but acts. These are agentic systems, capable of planning complex financial workflows and executing transactions with delegated authority. Once such a system is given the green light to act, it can move money, approve loans, or settle debts without waiting for a human to sign off on every single move. This transition from advice to action changes the fundamental nature of risk and efficiency in finance, yet for researchers and regulators, a critical problem has remained unsolved: how do you measure it?
Until now, there has been no reliable way to distinguish a bank that is merely talking about artificial intelligence from one that has actually handed over the reins of its operations to autonomous machines. Existing methods often treat all mentions of technology as the same, failing to see the difference between a chatbot that answers customer questions and a software agent that independently releases a payment. Without a clear measurement tool, it is impossible to know if these new systems are making banks more efficient, more risky, or both. A team of researchers at Dayananda Sagar University in India has addressed this gap by creating a new, rigorously tested tool called the Agentic AI Intensity Index. Rather than simply counting how often a bank mentions technology, this index measures the specific extent to which a bank has delegated the power to make and execute consequential financial decisions to autonomous systems.
The researchers did not just invent a new score; they built a complete scientific protocol to ensure the score is real and reliable. They started by gathering thousands of pages of public documents from banks, including annual reports, regulatory filings, and transcripts of earnings calls where executives speak to investors. From these texts, they developed a method to identify sentences that describe actual delegation of authority. They distinguished between language that describes a system waiting for human approval and language that describes a system acting on its own within a set mandate. To ensure this distinction was accurate, they brought together a panel of experts, including academic researchers, bank technology officers, and central bank supervisors. These experts refined a list of keywords and phrases over two rounds of discussion, agreeing on which terms truly signaled autonomous action and which were merely promotional fluff.
Once the list of keywords was finalized, the team tested their method to ensure it worked consistently. They had multiple trained coders read the same set of sentences to see if they agreed on whether a system was truly autonomous. They also used advanced computer models to help sort through the massive amount of text, but with a crucial safeguard: every time the computer identified a high-level autonomous action, a human expert reviewed it to confirm the classification. This human-in-the-loop approach prevented the system from being fooled by vague or misleading language. The researchers then checked if their new index actually measured what it claimed to measure by comparing it against independent evidence, such as job postings for specific types of engineers, patent filings for autonomous systems, and known instances where banks had deployed these technologies. The results showed that the index successfully separated banks with true autonomous capabilities from those that were simply discussing general digital transformation.
The core of this work is the index itself, which assigns a score to each bank based on the intensity of its agentic AI adoption. The score is not just a raw count of words; it is weighted by the level of authority the bank has granted to its machines. A system that executes a payment within a pre-set limit receives a higher score than one that merely drafts a recommendation for a human to review. The researchers also normalized the data to account for the fact that some banks write much more about technology than others, ensuring that a high score reflects genuine delegation of power rather than just a lot of talking. This allows for a fair comparison between a massive global bank and a smaller regional institution. The index is designed to be used in future studies to answer pressing questions: Does handing over decision-making power to machines make banks more efficient, or does it introduce new, unpredictable risks? Does it help banks detect fraud faster, or does it create new vulnerabilities?
The paper does not provide the final answers to these performance questions, as it is a methodological study focused on building the measuring stick rather than using it to measure a specific outcome. Instead, it lays out a clear, repeatable strategy for other researchers and regulators to use. It specifies exactly how to construct the index, how to validate it, and how to use it in statistical models to isolate the effects of autonomous AI from other factors like bank size or economic conditions. The researchers acknowledge that the transition to agentic AI is complex and that the benefits and risks are likely to depend heavily on how well a bank governs these systems. They propose that the impact of these technologies will vary based on the quality of a bank's internal oversight, the maturity of its technology, and the strictness of the regulations it operates under.
By providing a validated tool, this work removes a major barrier to understanding the future of banking. It allows the financial community to move beyond vague claims about "digital transformation" and start measuring the specific, tangible shift toward machine autonomy. The researchers have shown that it is possible to separate the signal of true agentic action from the noise of general technology discussion. This clarity is essential for supervisors who need to understand the risks of irreversible automated payments and for investors trying to gauge the true capabilities of the institutions they fund. The study concludes that while the technology is advancing rapidly, the ability to measure its adoption with precision is now within reach, provided that researchers and regulators adhere to the rigorous validation steps outlined in the new protocol. This foundation sets the stage for a new era of evidence-based analysis in banking, where the effects of autonomous systems can be studied with the same rigor as any other financial innovation.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.