← Latest papers
💻 computer science

Cybersecurity AI (CAI) Dataset

The paper introduces the CAI Dataset, a massive corpus of real-world cybersecurity LLM interactions that highlights the critical bottleneck of expert operator trajectories while underscoring the urgent need for on-premise, privately-hosted models to mitigate the systemic risks of centralizing sensitive offensive and defensive data within frontier-model API providers.

Original authors: Víctor Mayoral-Vilches

Published 2026-05-28
📖 5 min read🧠 Deep dive

Original authors: Víctor Mayoral-Vilches

Original paper dedicated to the public domain under CC0 1.0 (http://creativecommons.org/publicdomain/zero/1.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the world of cybersecurity as a massive, high-stakes game of chess played by millions of people every day. Some players are the "Blue Team" (defenders) trying to protect the castle, and others are the "Red Team" (attackers) trying to break in. For a long time, these players used their own brains and manual tools to play.

Recently, they started using a new, super-smart assistant: Artificial Intelligence (AI). But there was a problem. The AI was smart, but it didn't know how the experts actually played the game. It knew the rules, but not the tricks, the mistakes, or the messy, real-life strategies humans used in the heat of battle.

This paper introduces CAI Dataset, a massive library designed to teach AI how to think like a real cybersecurity expert. Here is the breakdown in simple terms:

1. The Problem: The "Expert" Gap

Think of the AI models (like the ones you chat with) as brilliant students who have read every textbook in the world. However, when it comes to cybersecurity, they were failing because they hadn't watched the masters at work.

  • The Discovery: Researchers found that the AI wasn't failing because it wasn't smart enough; it was failing because it lacked the "muscle memory" of real experts. It needed to see the actual step-by-step logs of how humans hack or defend systems, including the dead ends, the typos, and the "aha!" moments.

2. The Solution: The "Black Box" Recorder

To fix this, the authors built a tool called CAI (Cybersecurity AI). Imagine this tool as a transparent glass box that sits between the human expert and the AI.

  • How it works: When a human expert uses the CAI tool to do their job (hacking a test system, fixing a bug, or defending a network), the tool records everything. It writes down every command the human types, every tool they use, every mistake they make, and how they fix it.
  • The Collection: Over 14 months, this tool recorded 230,000 sessions from 16,000 different people in 123 countries. It's a library of 26 million prompts (instructions) and 18 terabytes of data. That's like recording every move in a billion games of chess.

3. What's Inside the Library?

This isn't just a list of "good" answers. It's a raw, unfiltered look at real work:

  • The Good and the Bad: About 36% of the data is offensive (attackers trying to break in), 20% is clearly malicious intent, 27% is business work (integrating systems), and 4% is defensive.
  • Real Life, Not Theory: Unlike other datasets that are made up by computers (synthetic), this data is real. It includes people pasting real passwords, server names, and secret keys into the chat because they need the AI to work right now.
  • The "Leak" Reality: The paper notes a surprising and slightly scary fact: Because the AI is so helpful, experts often paste their live secrets (like passwords and secret keys) into the chat to get the job done faster. They know the data is being logged, but the speed and convenience are worth the risk to them. This creates a massive concentration of the world's most sensitive secrets inside a few AI companies.

4. Why This Matters (The "Recipe" for a Better AI)

The authors say this dataset is the "secret sauce" needed to train the next generation of cybersecurity AIs.

  • Current State: Most AI training uses clean, perfect examples.
  • CAI Dataset: This uses messy, real-world examples. It teaches the AI how to handle a broken tool, how to recover from an error, and how to chain together a complex attack or defense over many steps.
  • The Goal: To build an AI that doesn't just know what a vulnerability is, but knows how to find and fix it in a real environment, just like a human expert.

5. The Big Warning and The Fix

The paper ends with a critical observation about safety and privacy:

  • The Risk: Right now, all these real-world secrets and attack strategies are flowing into the servers of a few big AI companies. If one of those companies gets hacked, or if the data is misused, it could cause chaos on a global scale.
  • The Solution: The only way to keep the speed and power of AI without leaking secrets is to run a specialized AI inside your own building (on-premise), where you control the data.
  • The Contribution: The CAI Dataset is being released to partners and customers specifically to help them build these private, on-site AI experts. This way, organizations can have a super-smart cybersecurity assistant that knows the real tricks of the trade but keeps all their secrets safe inside their own walls.

In a Nutshell

The paper says: "We recorded 14 months of real cybersecurity experts doing their jobs, mistakes and all, to create the ultimate training manual for AI. This data proves that experts are already trusting AI with their deepest secrets to get work done. To keep those secrets safe while staying efficient, we need to build private, specialized AI experts trained on this real-world data, rather than relying on public AI giants."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →