AI foundation model choice and false positive variation in Anti Money Laundering transaction screening
This study demonstrates that selecting different large language models for anti-money laundering transaction screening leads to drastic variations in false positive rates and operational workloads, establishing model choice as a critical operational risk parameter that requires explicit governance and monitoring.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Every year, vast sums of money move through the global financial system, some of it intended to hide the origins of illegal profits. To stop this, banks and other financial institutions are required to watch their customers' transactions closely, looking for patterns that suggest money laundering. For decades, they have relied on computer systems to do this initial watching. These systems act like a sieve, catching anything that looks even slightly suspicious and flagging it for a human investigator to review. The problem is that these sieves are often too coarse. They catch a tremendous amount of harmless activity along with the real threats, creating a mountain of false alarms. This forces analysts to spend hours investigating innocent people, wasting resources and causing unnecessary stress for customers.
Recently, a new type of computer intelligence has entered the conversation. These are large language models, powerful tools trained on vast amounts of text that can understand and reason about complex situations without needing to be taught specific rules for every scenario. Because they can read transaction histories and understand the story behind the numbers, many experts hoped they could solve the problem of false alarms. The idea was that these models could be smarter than the old rule-based systems, distinguishing between a legitimate business payment and a suspicious transfer with greater accuracy. However, a critical question remained unanswered: if a bank chooses one of these new models over another, does it actually matter? Or are they all just different versions of the same tool, interchangeable and predictable?
A team of independent researchers set out to find the answer by treating the choice of model not as a simple software decision, but as a fundamental risk to the bank's operations. They gathered five different, commercially available large language models from two major technology providers. They did not train these models on bank data or tweak their settings; instead, they asked each one to act as a fresh, zero-shot screener, meaning they had to make a judgment based solely on their existing knowledge. To ensure a fair test, the researchers created a controlled set of six hundred synthetic transaction cases. Three hundred of these cases were designed to look exactly like known money laundering schemes, while the other three hundred were perfectly legitimate transactions that were crafted to look suspicious enough to fool a less careful system. These were the "hard negatives," the innocent transactions that often get flagged by mistake.
The researchers then fed every single one of these six hundred cases to all five models, asking each one to decide if the transaction was suspicious or not. The results were startling. When the models looked at the three hundred legitimate transactions, their behavior varied wildly. One model flagged only a tiny fraction of them as suspicious, while another flagged the vast majority. In fact, the difference in how often they raised false alarms between the best and worst performers was enormous. For the same set of innocent transactions, one model made a mistake less than once in a hundred times, while another made a mistake more than eight times out of ten. The models did not just disagree on a few tricky cases; they disagreed on nearly ninety percent of the legitimate transactions. Whether a specific innocent customer got flagged for investigation depended almost entirely on which specific computer model was doing the screening.
The study also revealed a clear pattern within the families of models provided by each company. In both cases, the smaller, less expensive models were far more aggressive, flagging innocent activity at a much higher rate than their more powerful, expensive counterparts. The researchers found that these models were not simply better or worse at the same task; they were operating at completely different points on the spectrum of caution. One model was so eager to catch criminals that it accused almost everyone, while another was so cautious that it let some suspicious activity slip through to avoid bothering innocent people. This trade-off meant that choosing a model was not just a technical choice; it was a choice about how many false alarms the bank's staff would have to deal with every day.
To understand what this meant for a real-world bank, the researchers imagined a scenario where an institution processes one million transactions a day, with a very small percentage of them actually being suspicious. Under these conditions, the difference in model choice translated into a massive difference in workload. The most cautious model would generate a manageable number of alerts for investigators to review. The most aggressive model, however, would generate a queue of alerts nearly two hundred times larger. In this scenario, the most aggressive model would produce a flood of false alarms, where almost every single alert turned out to be a mistake, drowning out the few genuine threats. This showed that the choice of model could fundamentally alter the daily reality of the compliance team, shifting them from a state of focused investigation to a state of overwhelmed noise.
The researchers also checked whether these models would fail in the same way, a concern known as algorithmic monoculture, where everyone relies on the same tool and thus makes the same mistakes. They found that this was not the case. The models did not miss the same suspicious transactions; instead, they missed different ones. The real danger was not that they would all agree to let a criminal slip through, but that they would all disagree on who was innocent. This divergence meant that two banks using different models could have completely different views of the same customer's activity. One bank might see a normal business pattern, while the other sees a crime.
The study concluded that treating these large language models as interchangeable components is a dangerous mistake. Unlike traditional software, where the code is fixed and controlled by the user, these commercial models can be updated, replaced, or changed by the provider without the bank's direct knowledge. A model that works well today might be swapped for a different version tomorrow that behaves completely differently, shifting the bank's entire operation from a state of high precision to a state of chaos without any change in the bank's own systems. The researchers argued that banks must treat the specific identity of the model as a critical part of their risk management. They need to know exactly which version they are using, monitor how it behaves on a regular basis, and understand that changing the model is a major event that could drastically alter their ability to catch crime without harassing their customers. The choice of model is not just a procurement detail; it is the very setting that determines how the system sees the world.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.