ICS Cybersecurity Datasets: A Systematic Meta-Review of Coverage, Evaluation Practice, and Structural Gaps
This paper presents a systematic meta-review of 83 ICS cybersecurity datasets identified from 18 studies, revealing critical structural gaps such as architectural shallowness and progression compression that undermine current evaluation practices, and proposes a coordinated research agenda to address these deficiencies through improved corpus construction, benchmarking, and data governance.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where the invisible digital commands that keep our power grids humming, our water treatment plants clean, and our factories running are under constant threat. These systems, known as Industrial Control Systems, are the nervous system of modern civilization. For years, researchers have tried to build digital immune systems to spot and stop cyberattacks before they cause physical damage. To train these digital defenders, scientists rely on vast libraries of data—recordings of normal operations and simulated attacks. The hope has been that by feeding these recordings into computer programs, the programs would learn to recognize the subtle signs of a hacker's approach. But a new, comprehensive look at the entire collection of these data libraries suggests a troubling reality: the training ground is fundamentally broken.
A team of researchers from universities in Norway, Greece, and Spain recently decided to take a step back and examine the entire landscape of these data collections. Instead of testing a new algorithm or building a new simulator, they performed a massive audit of the research itself. They gathered eighteen major studies published between 2019 and 2026, which collectively described eighty-three different datasets used to train and test security systems. By organizing this information into a single, unified framework, they discovered that the data available to researchers is not just incomplete; it is structurally skewed in ways that make it difficult to fully validate if our defenses are working.
The researchers found that the vast majority of these data collections focus on a very specific, late stage of an attack. About eighty-five percent of the datasets only show what happens when an attacker has already broken in and is actively trying to disrupt the system, such as shutting down a valve or stopping a motor. What is missing is the story of how the attacker got there in the first place. The early stages of a cyber intrusion—where a hacker might quietly scan a network, steal passwords, or move laterally from a corporate office to an industrial control room—are almost entirely absent from the data. In fact, only about eight percent of the datasets capture the full journey of an attack as it crosses from the business side of a network into the physical control side. This means that the digital defenders being trained are like firefighters who have only ever practiced putting out a blaze after the house has burned down, but have never learned how to spot the arsonist lighting the match in the hallway.
The problem goes deeper than just the timing of the attacks. The data also fails to show the most critical parts of the physical machinery. Industrial systems are often described as a hierarchy, with the top levels managing business operations and the bottom levels directly controlling the physical sensors and motors. The researchers found that the data is heavily concentrated in the middle layers, where supervisors monitor the systems. The very bottom layer, where the actual physical work happens, is effectively invisible in the data collections. Without recordings from these field devices, researchers cannot build security systems that understand the true physical consequences of an attack. Furthermore, most of the data comes from simulated environments or laboratory testbeds rather than real, working industrial plants. While these simulations are useful, they often lack the messy, unpredictable noise of real-world operations, creating a gap between what the security systems learn in the lab and what they might face in reality.
Perhaps the most surprising finding concerns how these data collections are used to test security systems. The researchers discovered that the methods used to evaluate success are often flawed. In many cases, the data is split randomly between training and testing, which can accidentally leak information. Because industrial data is a continuous stream where one moment is closely related to the next, shuffling it randomly allows the testing program to utilize patterns from the future that it shouldn't know yet. This leads to inflated scores that make security systems look much better than they actually are. Additionally, the labels used to mark what is an attack and what is normal are often too vague. Instead of pinpointing exactly when an attack started and stopped, many datasets use broad labels that cover long periods, making it impossible to tell if a system detected a threat early or late.
The team concluded that the current state of research is trapped in a cycle of narrowness. The data is too shallow to teach systems about the full scope of an attack, too simulated to prepare them for the real world, and too poorly labeled to measure their true performance. They argue that simply making better algorithms will not solve the problem if the training data remains so limited. Instead, they propose a coordinated effort to build new datasets that cover the entire timeline of an attack, include recordings from the lowest levels of physical machinery, and come from real operational environments. They also call for stricter rules on how these datasets are tested, ensuring that security systems are evaluated on their ability to detect threats in real-time and under realistic conditions. Until these structural gaps are filled, the confidence we place in our industrial cybersecurity defenses remains, at best, an illusion, as benchmark performance claims are often bounded by the structural properties of the datasets on which they are obtained.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.