Fully Automated End-to-End Adversary Emulation from MITRE ATT\&CK Based Cyber Threat Intelligence Using LLMs
This paper introduces a fully automated end-to-end framework that leverages LLMs to generate, execute, and iteratively revise Caldera playbooks from MITRE ATT&CK-aligned CTI reports, significantly outperforming prior state-of-the-art systems like AURORA by achieving high execution success rates through an effective failure recovery mechanism.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the digital world as a giant, bustling city where bad guys are constantly trying to break into banks, steal secrets, or cause chaos. To keep this city safe, security guards (cyber defenders) need to know exactly how these bad guys think and act. But here's the catch: the bad guys are getting smarter, faster, and more creative every day. To stay ahead, defenders have to practice fighting these specific, real-world attacks. This is called "adversary emulation." It's like a fire drill, but instead of smoke and hoses, they use digital simulations to see if their security systems can spot and stop a hacker.
The problem is that these "fire drills" are incredibly hard to organize. Usually, a human expert has to read a long, complicated report about a recent cyberattack, figure out exactly what the hacker did, and then manually write a script to make their computer do the same thing. It's like trying to build a complex Lego castle by reading a manual written in a foreign language, then building it brick-by-brick with your eyes closed. It takes forever, costs a lot of money, and if you make a tiny mistake, the whole castle falls apart. This paper asks a big question: Can we teach a super-smart computer (an AI) to read these reports, build the Lego castle for us, and even fix it if it falls over, all without us needing to lift a finger?
The researchers in this paper say, "Yes, we can!" They built a fully automated robot system that acts like a digital master builder. Here is how it works: First, the system reads a real cyber-threat report (which is like a story about a hacker's heist). Then, it uses a powerful AI brain to translate that story into a set of instructions called a "playbook." Think of this playbook as a recipe for a cyber-attack. The system then cooks up this recipe on a test kitchen (a safe, isolated computer network) to see if it works.
But here is the magic part: in the past, if the recipe failed because the oven was too hot or the ingredients were missing, the human had to step in, figure out why, and rewrite the recipe. This new system doesn't wait for a human. If the attack fails, the system instantly looks at the error, guesses what went wrong (like "Oops, I used the wrong password" or "The file path is missing"), and tries to fix the recipe on its own. It keeps trying and tweaking the instructions until the attack runs smoothly.
The team tested this system on 11 different real-world attack stories. They found that the system was surprisingly good at its job. When using the best AI model they tried (called Claude Sonnet 4.5), the system successfully built and ran the attack recipes about 84% of the time after it fixed its own mistakes. Before the system started fixing things, it only succeeded about 70% of the time. This means the "self-repair" feature was a huge help, boosting the success rate by about 15 percentage points.
The researchers also compared their robot builder to the current best human-led system (called AURORA). Their robot was not only fully automatic but also managed to create more detailed attack plans and succeeded more often than the previous champion. It took the system less than 200 seconds and only about 35 cents of computer processing power to create and fix a playbook for each report. In contrast, humans might take months to do the same job manually.
The paper suggests that this approach is a major step forward because it removes the need for humans to constantly babysit the process. However, the researchers are careful to note that while the system is impressive, it isn't perfect. Sometimes, the environment is just too different from what the AI expected, and the robot can't fix the problem no matter how hard it tries. But for the vast majority of cases, this "self-healing" robot proves that we can automate the entire process of learning from past attacks and testing our defenses against them, making our digital cities much safer and our security teams much more efficient.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.