Evaluating Frontier AI Agents as Autonomous Clinical Security Auditors
This paper introduces an open evaluation task based on the METR Task Standard that tests whether frontier AI agents can autonomously conduct structured clinical AI security audits by implementing attacks, calculating security scores, and generating reports, demonstrating that models like Claude Sonnet 4.6 and GPT-4.1 can achieve perfect scores across multiple datasets and architectures.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where computers are learning to be doctors, helping to diagnose diseases and decide on treatments. These "AI doctors" are incredibly smart, but they have a secret weakness: they can be tricked. Just like a human might be fooled by a clever optical illusion, an AI can be confused by tiny, almost invisible changes to the data it receives. If a hacker knows how to make these tiny changes, they could trick an AI into saying a patient is healthy when they are actually sick, or vice versa. This is called an "adversarial attack," and it's a huge problem because if we don't find these tricks before we let the AI help real patients, people could get hurt.
To stop this, we need "security auditors"—specialists who try to break the AI on purpose to see how strong it is. But here's the catch: being a security auditor is hard. It requires deep math skills, special computer programs, and a lot of time. Most hospital teams don't have experts who can do this, and the tools to do it are often missing or too complicated. So, the big question is: Can a new kind of super-smart computer program, called an "AI agent," do this hard job all by itself? These agents are like digital interns that can read instructions, write their own computer code, run tests, and write a report without a human holding their hand at every step.
This paper puts that idea to the test. The researchers created a challenging "exam" for three of the world's most advanced AI agents. The exam asked them to act as autonomous security auditors for a clinical AI model. The agents were given a medical prediction model (like one that predicts breast cancer or ICU mortality), a dataset of patient records, and a written set of instructions on how to perform four different types of security attacks. The agents had to write their own code from scratch to execute these attacks, crunch the numbers, and write a final security report in a specific format, all inside a secure digital container with a strict time limit of 40 "turns" (or steps).
The results were a mix of amazing success and interesting failures. Two of the three AI agents, Claude Sonnet 4.6 and GPT-4.1, aced the test. They completed every single version of the exam perfectly, writing correct code and generating accurate security reports every time they tried. They did this efficiently, using relatively few "tokens" (the units of text the AI processes). However, the third agent, GPT-4o, struggled. It only succeeded in about 61% of its attempts. When it failed, it didn't fail because it couldn't understand the math; it failed because it got lost in the process. Sometimes it stopped working before it could save the final report, sometimes it made a simple math error when adding up the final scores, and sometimes it submitted an empty file. The study suggests that while top-tier AI agents are now capable of doing complex, multi-step security auditing on their own, they aren't all equally reliable. The difference between success and failure often came down to "discipline"—the ability to stay focused, track progress, and finish the job without getting distracted or making careless mistakes. This means that for hospitals wanting to use AI to check their own medical AI, choosing the right "digital auditor" is just as important as the medical model itself.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.