Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows
This paper identifies a critical evidence gap in deploying LLMs for fraud detection and trust-and-safety workflows, noting that while fraud-focused studies lack operational metrics like latency and cost, moderation research provides more robust data, and it proposes the FORTE framework and a minimum evidence checklist to guide future deployment claims.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are the head of security for a massive, bustling digital city. This city has two main jobs: keeping the streets safe from bad actors (like scammers and fraudsters) and making sure everyone follows the rules of the community (like stopping hate speech or spam). In the past, you hired a team of human detectives to read every single note, check every transaction, and decide who was telling the truth. But the city has grown too big, and the human team is drowning.
Enter the new recruits: Large Language Models (LLMs). Think of these not as human detectives, but as incredibly fast, super-smart robots that can read a million pages in a second and understand the meaning behind the words. They are being proposed to help the human team by flagging suspicious notes, summarizing complex cases, or even acting as the first line of defense. But here is the tricky part: just because a robot is smart doesn't mean it's ready for the real world. In a video game, a robot might be perfect. But in a real city, if the robot is too slow, it causes traffic jams. If it costs too much to run, the city goes bankrupt. And if it gets tricked by a clever liar, the whole system could collapse. The big question isn't just "Is the robot smart?" but "Can we trust this robot to work inside our busy, high-pressure city without breaking anything?"
This paper, titled "Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows," is like a rigorous inspection report for these new robot recruits. The author, Keyur Gabani, looked at 49 different studies and reports that talk about using these AI robots to catch fraud, stop scams, and moderate content. The goal was to see if there is enough proof that these robots can actually do the job in a real, live system, or if they are just playing in a sandbox.
The main finding of the paper is a bit of a shock: there is a huge gap between what we think these robots can do and what we actually know they can do in the real world. The author calls this an "evidence imbalance."
Here is the story the paper tells:
The Fraud Team is Missing the Scorecard
In the world of catching fraud (like credit card scams or phishing emails), the researchers found 18 different studies. These studies are great at showing that the robots can spot a scam in a test tube. They can say, "Hey, this looks like a scam!" with high accuracy. But when the author looked for the "scorecard" of real-world performance, it was almost empty.
- The Missing Numbers: None of the 18 fraud studies reported how long it takes the robot to make a single decision (latency). None of them told us how much money it costs to make that single decision (cost). None of them explained how sure the robot is when it makes a call (calibration).
- The Metaphor: It's like having a race car that is proven to be fast on a test track, but nobody has ever told you how much gas it burns per mile or if it can handle a rainy day. You know it can go fast, but you don't know if it's safe to drive on the highway.
The Moderation Team is Showing More Cards
In the world of content moderation (stopping bad comments or spam), the author found 14 studies. Surprisingly, this group was more honest about the real-world stuff. They were more likely to talk about how much time the robot takes, how much it costs, and how to handle fairness (making sure the robot doesn't treat different groups of people unfairly).
- The Metaphor: The moderation team is like a construction crew that not only builds the house but also gives you a detailed bill of materials and a timeline. They are still working out the kinks, but they are showing their work.
The "FORTE" Framework: A New Way to Look at the Robots
To make sense of all this, the author created a new way to organize the robots, called FORTE. Instead of just asking "What model is this?", FORTE asks "What job is this robot doing?"
- The Classifier: The robot acts as a judge, saying "Guilty" or "Not Guilty."
- The Assistant: The robot helps a human detective by summarizing a long case file.
- The Investigator: The robot goes out, gathers clues, and writes a report.
- The Escalator: The robot handles easy cases and only sends the hard, confusing ones to a human.
The paper argues that we need to stop looking at these robots as magic black boxes and start looking at them as specific tools with specific jobs.
What the Paper Rules Out
The paper is very clear about what it doesn't say. It does not say that these robots are bad or that we shouldn't use them. It also does not say that the robots are ready to replace human detectives entirely.
- The Reality Check: The author explicitly states that we do not have enough proof to say an LLM can replace a real-time fraud engine right now. We don't know if they are fast enough, cheap enough, or safe enough against clever hackers who try to trick them.
- The "Sandbox" Problem: Many of the studies the author looked at were done "offline." This means they tested the robots on old data in a quiet room. The paper argues that this is not the same as testing them in a noisy, chaotic, live city where hackers are actively trying to break in.
The Missing Pieces and the Future
The paper concludes that before we can fully trust these robots in our digital city, we need to fill in the missing scorecards. The author suggests a "minimum checklist" for anyone wanting to deploy these robots:
- Latency Budget: How fast does it need to be?
- Cost Per Decision: How much does each check cost?
- Decision Threshold: When does it say "I'm not sure, let a human handle this"?
- Explanation Integrity: If the robot gives a reason for its decision, is that reason true, or did it just make something up?
- Adversarial Pressure: What happens when a bad actor tries to trick the robot?
The paper suggests that the best path forward right now is a "hybrid" approach. Let the fast, simple, and cheap tools handle the easy stuff. Save the super-smart, expensive robots for the hard cases, for explaining why a decision was made, or for helping human detectives.
In short, the paper is a call to action. It tells us that while the technology is exciting and promising, we are currently flying blind in many areas. We have the robots, but we don't have the manual. To make these systems safe and useful, researchers and companies need to stop just showing off how smart the robots are and start showing us how they behave when the pressure is on, the clock is ticking, and the cost is real.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.