← Latest papers
💻 computer science

Testing and Evaluation of Agentic AI Systems In Military Command and Control

This paper argues that the inherent properties of agentic AI systems undermine eight foundational assumptions of current military testing and evaluation methods, thereby breaking the logical link between test evidence and fielded behavior, which necessitates a shift toward narrower, recoverable assurance claims and a governance model that treats deployment as a continuing act of risk management.

Original authors: Ulysse Richard, Heather Frase, Sarah Cao, Di Cooke, Sebastian Kwon, Adrianna Tan

Published 2026-08-24
📖 6 min read🧠 Deep dive

Original authors: Ulysse Richard, Heather Frase, Sarah Cao, Di Cooke, Sebastian Kwon, Adrianna Tan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the high-stakes world of military command, leaders rely on a system to gather information, make sense of it, and suggest courses of action. This system, known as command and control, is the nervous system of an operation, connecting sensors on the ground and in the sky to the commanders who must decide where to send troops or ships. For decades, the tools used in these systems have been predictable machines: they follow strict rules, and if you feed them the same information twice, they give you the same answer. But a new kind of technology is now entering these rooms. These are not just tools that follow orders; they are agents. Unlike a simple calculator, an agent can look at a situation, decide what information it needs to find, go out and get it, and then use that new information to change its own plan. It can remember past events, talk to other similar systems, and adapt to surprises without a human telling it exactly what to do next. This flexibility makes them incredibly powerful, but it also makes them difficult to understand. The central question for the military is no longer just whether these agents work, but how we can be sure they will work safely when the stakes are highest.

A team of researchers set out to answer this question by examining how we currently test these new, flexible systems. They looked at 240 different methods that military and industry experts use to test software and autonomous systems. Their goal was to see if these old methods could handle the new reality of agents that learn, remember, and change on the fly. They found a fundamental problem: the way we currently test these systems relies on assumptions that simply do not hold true for agents. Traditional testing assumes that a system is a fixed object, like a car, where the parts do not change and the way it behaves is predictable. It assumes that if you test the car today, the results will still be valid tomorrow, and that you can test each part of the car separately and know how the whole car will behave.

The researchers discovered that agentic systems break every one of these assumptions. First, these systems are not fixed. An agent builds its own path by choosing which tools to use and which information to gather. Because it makes these choices while it is working, the system it is testing is different from the system it was designed to be. It is like trying to test a car while the driver is simultaneously changing the engine and the tires. Second, the system is not stable. An agent remembers what happened in previous tasks, and that memory changes how it acts in the next task. Even if the software code never changes, the agent's behavior changes because its internal state has shifted. This means that a test result from last week might not apply today. Third, these systems are often made of many smaller agents working together. When they do, new behaviors can appear that were not present in any of the individual parts. A group of agents might solve a problem in a way that none of them could have predicted alone, creating a result that is impossible to foresee by testing the parts in isolation.

Because of these changes, the researchers found that current testing methods can produce a report that looks perfect on paper, satisfying all the official rules, while still failing to prove that the system will be safe in real life. The tests might show that the system works under specific, controlled conditions, but they cannot guarantee that the system will behave the same way when it is out in the field, facing a chaotic and changing environment. The link between the test results and the real-world performance is broken. The evidence exists, but the argument that connects that evidence to a promise of safety is no longer strong enough.

The paper does not say that we should stop using these systems or that they are impossible to test. Instead, it suggests that we must change how we build our case for safety. We cannot rely on broad promises that the system is safe in all situations. Instead, we must make much narrower, specific claims. For example, we can test whether the system stays within a specific, limited set of actions, or whether it follows a correct path of steps even if the final answer varies. We can also accept that some evidence can only be gathered while the system is actually in use. This means that the decision to put an agent into service is not a one-time event. It is a continuous process where we constantly monitor the system, check its memory, and verify that it is still behaving as expected. If the system drifts or changes in a way we did not predict, we must be ready to stop it or adjust it.

The researchers also looked at how humans can supervise these agents. They found that simply having a human in the room is not enough. If the agent is too fast, too complex, or too opaque, the human cannot understand what is happening fast enough to intervene. The system must be designed so that the human knows what the agent is doing, why it is doing it, and has enough time to stop it if things go wrong. Without this, the human becomes a bystander rather than a supervisor.

Ultimately, the study concludes that while we cannot yet prove that these complex, changing systems are perfectly safe in every possible scenario, we can still manage the risk. We do this by being honest about what we do not know. We define the limits of what the system is allowed to do, we monitor it closely while it works, and we accept that our confidence in it is always provisional. The decision to use these agents is no longer a final stamp of approval; it is an ongoing act of governance, where we constantly weigh the evidence, manage the uncertainty, and ensure that human judgment remains in control. The path forward is not to find a perfect test that proves safety once and for all, but to build a system of checks and balances that allows us to use these powerful tools while keeping the risks within our ability to manage.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →