A Multi Method Importance and Performance Efficiency Analysis of Topological Metrics for Natural Visibility Graph Based Cyber Attack Detection
This study demonstrates that applying a consensus-based importance analysis to Natural Visibility Graph metrics allows for the selection of a compact subset of three topological features that significantly improves cyber-attack detection accuracy and reduces computational runtime by over 96% compared to using the full set of 21 metrics.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the digital world, network traffic flows like a constant river of data, carrying everything from harmless web browsing to sophisticated cyber-attacks. To protect systems, security experts use computer programs to watch this flow, looking for patterns that signal danger. For years, these programs have relied on counting simple things, like how many packets of data pass through a door or how long a connection lasts. However, researchers have begun to realize that the shape of the data itself might hold more clues than the numbers alone. Imagine taking a line of numbers representing a data stream and turning it into a map where each point is a dot and lines connect dots that can "see" each other without being blocked by anything in between. This creates a unique web, or graph, that preserves the rhythm and structure of the original data. By studying the shape of these webs—how tightly the dots cluster together, how many paths exist between them, and how central certain points are—scientists can uncover hidden signatures of malicious activity that standard counting might miss.
A team of researchers at Karabuk University in Turkey set out to test just how useful these shape-based clues are for catching cyber-attacks. They started with a massive collection of network records containing both normal traffic and various types of attacks. Instead of feeding the raw data directly into a computer program, they first transformed small chunks of the data into these visual webs. From each web, they extracted twenty-one different measurements, or topological metrics, to describe its structure. Some of these measurements looked at how connected the dots were, others examined the average distance between points, and some analyzed how the dots grouped into communities. The researchers then used a powerful type of artificial intelligence, a convolutional neural network, to learn how to distinguish between the webs created by normal traffic and those created by attacks.
The central question the team wanted to answer was whether all twenty-one of these shape measurements were actually necessary. In many scientific fields, researchers often assume that gathering more data leads to better results, but in this case, extracting every single measurement takes a significant amount of time and computing power. The team suspected that some of these measurements might be redundant or less helpful than others. To find out, they applied four different methods to rank the importance of each measurement. One method looked at how much the computer's confidence dropped when a specific measurement was scrambled; another compared the real measurements against random noise to see which ones stood out; a third method systematically removed the least useful ones to see what remained; and a fourth method used a mathematical approach to explain exactly how much each measurement contributed to the final decision. By combining the results of these four distinct approaches, they created a single, agreed-upon ranking of the twenty-one metrics.
The results of this analysis revealed a clear hierarchy. The measurements that described how tightly the points in the web clustered together proved to be the most valuable. Specifically, the median, standard deviation, and mean of these clustering values were ranked as the top three most important indicators. In contrast, other measurements, such as those describing the longest paths or the overall size of the web, were ranked much lower. The researchers then tested this finding by training their computer program using only the top-ranked measurements, gradually reducing the number of inputs from all twenty-one down to just the top three. They compared the performance of these smaller sets against the full set of twenty-one.
The outcome was surprising. When the computer program was trained using only the top three clustering measurements, it performed better than when it was trained with the full set of twenty-one. The simplified model correctly identified attacks with an accuracy of 97.148 percent, compared to 95.999 percent for the full model. It also achieved a higher score in balancing the detection of different attack types. More importantly, this improvement came with a massive gain in speed. The full set of measurements took nearly fifteen thousand seconds to process through the entire experiment, while the top three measurements required only about five hundred and eighty-nine seconds. This represents a reduction in computing time of over ninety-six percent. The study suggests that by focusing on the most informative structural clues and discarding the rest, security systems can become both faster and more accurate.
The researchers noted that this result does not mean that all other measurements are useless, but rather that for this specific type of data and this specific computer model, the extra information was not helping and may have even been distracting. The top three measurements, which describe the local clustering of the data points, seemed to contain the essential signature of the attacks. The study concludes that a compact, carefully selected set of features can outperform a larger, unfiltered collection. While the findings are specific to the dataset and methods used in this experiment, they offer a clear path forward for building more efficient cyber-defense systems. By understanding which structural details matter most, future tools can be designed to detect threats with greater speed and precision, without the burden of processing unnecessary data.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.