← Latest papers
🤖 AI

To Defend Against Cyber Attacks, We Must Teach AI Agents to Hack

This position paper argues that to counter inevitable AI-driven cyber attacks, defenders must fundamentally shift their strategy by responsibly developing and deploying their own offensive AI capabilities within controlled, governed environments to master threats before adversaries do.

Original authors: Terry Yue Zhuo, Yangruibo Ding, Wenbo Guo, Ruijie Meng

Published 2026-02-04
📖 5 min read🧠 Deep dive

Original authors: Terry Yue Zhuo, Yangruibo Ding, Wenbo Guo, Ruijie Meng

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: The Old Rules Don't Work Anymore

For the last decade, cybersecurity has worked like a game of "Whac-A-Mole."

  • The Old Way: Hackers had to be highly skilled humans. Building a complex attack was like building a custom house; it took months of hard work and deep expertise. Because it was so expensive and slow, hackers only attacked the "big houses" (major companies) or used simple, generic tools to knock on thousands of doors hoping one was open. Defenders assumed hackers couldn't afford to knock on every door.
  • The New Threat: AI Agents are changing the game. Imagine a hacker who isn't a person, but a swarm of thousands of tiny, tireless robots. These robots don't need to be experts. They can knock on millions of doors in a second. Even if they fail 99% of the time, they only need to succeed 1% of the time to make a profit. They can attack the "small houses" (niche systems) that humans used to ignore because they weren't worth the effort.

The paper argues that defenders are currently trying to stop these robots by putting up "Do Not Enter" signs on the robots themselves. But the paper says this won't work because bad actors can just build their own robots without the signs.

Why Current Defenses Are Failing

The authors say we are trying to stop AI hackers by using three main methods, but they are all like trying to stop a flood with a sieve:

  1. Filtering the Training Data (The "Clean Library" approach):

    • The Idea: We remove all bad instructions (like "how to build a bomb") from the AI's training books so it never learns to be bad.
    • Why it Fails: Bad hackers don't need to read the "bad books." They can teach the AI to figure out how to hack using basic logic and common sense, just like a human can figure out how to pick a lock without a manual. Also, bad hackers can just download the AI and teach it themselves.
  2. Safety Alignment (The "Good Behavior" approach):

    • The Idea: We train the AI to say "No" when asked to do something bad.
    • Why it Fails: Bad hackers are clever. They can trick the AI into thinking a bad request is actually a good one (like asking a guard to "test the security" instead of "break in"). Also, if the AI is running on a hacker's own computer, they can just retrain it to ignore the "No" button.
  3. Output Guardrails (The "Bouncer" approach):

    • The Idea: We put a bouncer at the door to check every message the AI sends and block anything that looks dangerous.
    • Why it Fails: A hacker can break a big attack into a thousand tiny, harmless-looking steps. The bouncer sees each step as safe, but when you put them all together, they build a bomb.

The Proposed Solution: Build Your Own "Red Team" Robots

The paper suggests a radical shift: We must teach our own AI agents how to hack, so we can use them to defend us.

Think of it like a fire department.

  • Current Defense: We wait for a fire to start, then try to put it out.
  • The Paper's Idea: We hire a team of professional arsonists (our "Offensive AI") to work inside a secure, walled-off training facility (a "Cyber Range").
    • These "Good Arsonists" try to burn down the building using every trick in the book.
    • Because they are AI, they can try millions of ways to start a fire in seconds.
    • We watch where they succeed, find the weak spots in our building, and fix them before a real bad guy shows up.

The Three Steps to Make This Safe

The authors know this sounds dangerous, so they propose three strict rules to keep the "Good Arsonists" from causing real damage:

  1. Build Better Test Tracks: We need better "driving courses" (benchmarks) to measure exactly how good these hacking robots are. We need to test them on real-world scenarios, not just simple puzzles.
  2. Train, Don't Just Script: Instead of giving the robots a fixed list of steps to follow, we need to train them to learn and adapt, just like a human hacker would. This makes them better at finding new, unknown weaknesses.
  3. The "Sandbox" Rule: This is the most important part. The "Good Arsonists" must never leave the secure training facility.
    • They stay inside a digital cage where they can't touch the real internet.
    • They find the holes and write a report.
    • Then, we take that report and give it to a different, "Safe AI" that only knows how to fix holes, not make them. This "Safe AI" is what we release to the public to protect our systems.

The Bottom Line

The paper concludes that we cannot wait for bad hackers to figure out how to use AI to attack us. By the time they do, it will be too late.

Instead, we must master the weapon first. We need to build our own offensive AI, keep it locked in a safe cage, and use it to find every possible way our systems can be broken. Only by understanding how the enemy thinks and acts at a massive scale can we hope to defend ourselves.

In short: To stop the bad robots, we have to build our own good robots that are even better at breaking things, but we must keep them in a cage so they only help us fix our house.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →