← Latest papers
💻 computer science

Multi-LLM Consensus Framework for Evaluating Banking-Sector NIDS Dataset Coverage of MITRE ATT&CK Techniques

This paper proposes a multi-LLM consensus framework to evaluate how well existing NIDS benchmark datasets cover banking-specific MITRE ATT&CK techniques, revealing significant gaps such as the 89.9% blind spot in CIC-DDoS2019 and establishing a need for sector-native datasets to ensure real-world operational effectiveness.

Original authors: Sanjida Khanom, Sadia Afrin Khan, Adrita Rahman Tory, Md. Ahsan Habib, Khondokar Fida Hasan

Published 2026-08-04
📖 4 min read☕ Coffee break read

Original authors: Sanjida Khanom, Sadia Afrin Khan, Adrita Rahman Tory, Md. Ahsan Habib, Khondokar Fida Hasan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the internet as a massive, bustling city where banks are the most important vaults. To keep these vaults safe, security guards called Network Intrusion Detection Systems (NIDS) patrol the digital streets, watching for anyone trying to break in. These guards are trained using "practice drills"—huge collections of fake attack data called datasets. The problem is, most of these drills are generic. They teach the guards how to stop a common pickpocket or a car thief, but they rarely show them how to stop a mastermind who knows exactly how to hack a bank's special money-transfer system or trick an ATM. It's like training a firefighter only on kitchen fires and then sending them to put out a chemical plant explosion; they might know how to use a hose, but they won't know what to do when the chemicals start reacting. This paper asks a simple but critical question: Are our current training drills actually preparing these digital guards for the specific, sneaky attacks that target banks, or are they just making the guards feel confident while leaving them vulnerable?

The authors of this paper decided to investigate this gap by building a new, smarter way to test these training datasets. Instead of just counting how many "bad guys" are in a dataset, they used a method that mimics how real-world security sensors actually work. They started with a massive list of 210 different ways hackers might attack a bank, taken from a famous guide called MITRE ATT&CK. However, they knew that real security sensors have limits: they can't see inside encrypted messages (like a locked letter), they can't see what's happening inside a computer's brain (like a file being deleted), and they can only watch traffic passively, like a camera on a pole, without being able to stop it.

To figure out which of those 210 attacks a sensor could actually see, the researchers didn't just ask one expert. They used a "consensus engine" made of four powerful artificial intelligence models. Think of it as a panel of four super-smart detectives who all have to agree on whether a specific attack leaves a visible trail. If three out of four detectives said, "Yes, we can see that," it counted as detectable. If they disagreed, the team was careful and assumed the attack was invisible to the sensor. This strict rule filtered the list down to just 68 attacks that a real bank sensor could realistically spot.

Next, they took these 68 "detectable" bank attacks and checked five popular training datasets to see how well they covered these specific threats. The results were eye-opening. They found that the dataset called UNSW-NB15 was the best of the bunch, covering about 82.2% of the important banking attacks. However, there was a catch: only 18.4% of that coverage was a direct match. The rest was just a vague hint, like seeing smoke but not knowing if it's from a fire or a barbecue. On the other hand, a dataset called CIC-DDoS2019, which is often used for testing, was a disaster for banking security. It missed 89.9% of the core banking behaviors, leaving a massive blind spot. It turns out this dataset is great for spotting big, loud denial-of-service attacks, but it's completely useless for the subtle, multi-step tricks hackers use to steal money.

The paper suggests that relying on these old, generic datasets gives banks a "false sense of security." Just because a security system scores high on a generic test doesn't mean it will work when a real hacker tries to manipulate a SWIFT money transfer or jack an ATM. The authors conclude that we urgently need new training data that is built specifically for banks, including the unique languages and protocols they use. Until then, the digital guards might be very good at spotting common thieves, but they are likely to be caught off guard by the sophisticated heists that actually threaten the global financial system.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →