← Latest papers
💻 computer science

Open Security Benchmark: Towards Autonomous Enterprise Cyber Defense

This paper introduces the Open Security Benchmark (OSB), a novel framework designed to evaluate and build trust in autonomous AI agents for enterprise cyber defense by providing a frozen, holistic, and queryable security environment with ground-truth answers to bridge the critical gap in public, cross-vendor evaluation data.

Original authors: Gal Engelberg, Michael Arenzon, Leon Goldberg

Published 2026-07-31
📖 3 min read☕ Coffee break read

Original authors: Gal Engelberg, Michael Arenzon, Leon Goldberg

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a world where computers don't just follow orders but act like junior detectives, hunting down security holes in massive, messy office buildings before the bad guys can break in. This is the dream of "autonomous cyber defense," where AI agents constantly scan an organization's digital state—checking who has keys to the server room, which software is outdated, and if anyone is sneaking around. But here's the catch: for a detective to be trusted, you need to know if they are actually solving the case or just guessing. Right now, there's a huge problem: real company data is secret, so researchers can't test these AI detectives on the same messy, real-world scenarios. It's like trying to teach a pilot to fly a plane by only showing them pictures of clouds, never letting them touch the controls in a real storm. Without a shared, fair testing ground, we can't know if these AI agents are ready to save us or if they're just making things up.

This paper introduces a solution called the Open Security Benchmark (OSB), which is essentially a giant, frozen, fake "digital city" built specifically to test these AI detectives. Think of it as a massive, unchangeable video game level that looks and feels exactly like a real corporation, complete with fake employees, fake cloud servers, and fake passwords. The authors created this "city" to close the gap between theory and reality. They built a framework where AI agents are given a mystery to solve—like "Find all the employees who were fired last month but still have access to the company's secret files"—and the agents have to find the answer using two different methods: either by writing a query in a special language (SQL) to search a giant database, or by pretending to be a human operator clicking through real vendor tools like AWS or Google Workspace.

The paper doesn't just say "look, we made a game." It lays out a rigorous framework to test these agents by freezing the environment so that every time an agent runs, the "ground truth" (the correct answer) stays exactly the same. This allows researchers to grade the AI not just on whether it got the right answer, but on how it got there. Did it ask the right questions? Did it connect the dots between different systems? The study instantiates this framework with synthetic, fake organizations, allowing for testing across different sizes of companies (from small teams of 75 people to large ones with 2,000 employees). The framework is designed to reveal whether AI agents can handle the complexity of real-world tasks, such as connecting a human resources record to a cloud account across different vendors, a capability that is notoriously difficult for single-vendor tools.

Crucially, the paper argues that we cannot trust an AI agent just because it gives a confident-sounding answer. The authors explicitly rule out the idea that we can evaluate these agents in isolation or on small, clean datasets; they insist that the environment must be "frozen" and cross-vendor to be fair. They are not claiming that AI has solved cyber defense yet. Instead, they suggest that this benchmark is a necessary tool to measure progress, turning the vague question of "Is this AI safe?" into a concrete scorecard. By providing a shared, reproducible target, OSB aims to move the field from guessing to measuring, ensuring that when we eventually let AI agents take the wheel in our real digital cities, we know exactly how they will perform.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →