On fair and realistic performance evaluations for graph-based lateral movement detectors
This paper critiques the inconsistent preprocessing and labeling practices in existing lateral movement detection benchmarks, proposes standardized methodologies to ensure fair and realistic evaluations, and demonstrates that re-evaluating established detection methods under these new policies yields significantly different performance results.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the internet as a giant, bustling city where every computer is a building and every data packet is a delivery truck. In this city, security guards (cybersecurity experts) are constantly watching for trouble. Usually, they look for someone trying to break into a building from the outside. But the sneakiest thieves don't just break in; once they get into one building, they start hopping from room to room, stealing keys, and opening doors to the most valuable vaults. This sneaky hopping around inside the city is called "lateral movement." To catch these digital burglars, researchers build "detectives"—computer programs that look at the city's traffic logs and try to spot the weird patterns of a thief moving between buildings.
For years, these researchers have been racing to build the best detectives. To see who is winning, they use "training grounds": giant, pre-made datasets that simulate a city with a known thief running around. The idea is simple: if your detective finds the thief in the training ground, it's good. But here's the catch: the training grounds are messy. Some researchers clean up the logs by throwing away "boring" traffic, while others label every single step the thief takes as a crime. It's like judging a detective's skill by giving them a test where you've already hidden all the difficult clues and told them exactly where the bad guy is. If the rules of the test change from one researcher to another, you can't really tell who is actually the best detective.
This is exactly what the paper "On Fair and Realistic Performance Evaluations for Graph-Based Lateral Movement Detectors" investigates. The author, Corentin Larroche, acts like a referee who decides to check the scorecards of the last few big competitions. He looks at two famous training grounds (datasets called LANL and OpTC) and discovers that researchers have been using wildly different rules to prepare the data and label the "crimes." Some were filtering out innocent people to make the test easier, while others were counting harmless traffic as malicious. The paper argues that these inconsistent rules are making the results look much better than they really are.
Larroche then proposes a new, stricter set of rules for preparing these datasets. He suggests that we should keep all the traffic in the test, not just the "interesting" parts, and that we should be very careful about what we call a "lateral movement." When he re-runs the tests for three famous detection methods using his new, fairer rules, the results change dramatically. The detectives that looked like superheroes in the old tests suddenly look much less impressive. For example, one method that previously claimed a near-perfect score drops significantly when forced to deal with the messy, realistic data. The paper concludes that while lateral movement detection is still a very hard problem, we need to stop using biased adjustments in our tests if we ever want to build a system that works in the real world.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.