Runtime Continuation after Faults in Partially Executed Agent Workflows
This paper proposes a runtime continuation mechanism for partially executed language-model agent workflows that ensures safe recovery from faults by reconstructing a trusted frontier, deriving residual obligations from a canonical contract, and advancing state only through authorized effects supported by valid, fresh, and uniquely matching evidence.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the modern digital landscape, artificial intelligence has evolved from a system that simply writes text to one that actively performs tasks. These intelligent agents can search the web, edit files, send messages, and update databases on a user's behalf. They operate by following a sequence of steps, often called a workflow, where each action changes the state of the real world. However, this ability to act creates a unique vulnerability. If an agent makes a mistake, gets confused by a trick, or encounters a system error after it has already changed a file or sent a message, simply trying again or starting over can be dangerous. Replaying a failed sequence might repeat a harmful action, while continuing from a confused state might build on a false assumption. The core challenge is determining exactly which part of the history is trustworthy, what work remains unfinished, and what proof is needed to safely move forward without repeating errors or falling for deception.
Researchers Li Zeng and his team at the Institute of Information Technology of the National Immigration Administration and the Hong Kong University of Science and Technology have developed a new method to solve this problem. They call their system AgentMirror. Instead of treating a broken workflow as a simple error to be fixed by a new guess, AgentMirror treats the recovery process as a strict, step-by-step verification of reality. The system operates on the principle that once an agent has performed an action, the history of that action becomes a fixed point of truth. If the agent crashes or is tricked after that point, the system does not ask the artificial intelligence to decide what happened. Instead, it looks at a verified record of what actually occurred, identifies exactly which tasks are still left to do, and demands concrete proof before allowing any new action to change the state of the world again.
The researchers built this system around a concept they call a "trusted frontier." Imagine a line drawn in the sand that marks the exact moment where the agent's history is confirmed to be safe and correct. Everything before this line is accepted as true; everything after is unverified. When a fault occurs, the system does not let the artificial intelligence rewrite this line or pretend that a mistake didn't happen. It reconstructs the state of the agent based only on the verified history up to that line. From this point, the system calculates a list of "residual obligations"—the specific, mandatory tasks that the agent still needs to complete to finish the job. It then presents this list to the artificial intelligence as a guide, telling it what work is left, but without giving the intelligence the power to just do it.
The critical innovation lies in how the system handles the next step. When the artificial intelligence proposes a new action, that proposal is treated as untrusted. The system forces the action through a standard security check, known as "ordinary admission," to ensure it is allowed to run. If the action executes, the system does not immediately accept that the task is done. It waits for a "runtime receipt," which is a digital certificate issued by a trusted authority confirming that the action actually happened and was observed correctly. This receipt must be fresh, meaning it was just created, and it must be bound to the current trusted frontier, meaning it cannot be a replay of an old action or a fake from a different session. Only when this receipt is verified and matches exactly one of the remaining tasks does the system allow the official record of the agent's progress to move forward.
To test their idea, the researchers used a benchmark called AgentDojo, which simulates various tasks and deliberately introduces faults and attacks to see how agents behave. They compared their AgentMirror system against other recovery methods, such as simply retrying the last step or letting the agent guess its way out of trouble. The results showed that AgentMirror was significantly better at completing tasks safely. It successfully recovered from faults without allowing unauthorized actions to slip through, and it prevented the agent from falling for trick instructions that tried to make it repeat harmful operations. The study demonstrated that by separating the guidance given to the artificial intelligence from the authority to execute actions, and by requiring fresh evidence for every step, the system could maintain a secure path forward even when the agent itself was compromised.
The researchers also explored what would happen if they removed specific parts of their safety rules. They found that if they skipped the step of verifying the receipt, the system would accept fake or old evidence, leading to incorrect progress. If they allowed the recovery guide to act as a permission slip, the system would let untrusted actions run. If they did not check that the action matched exactly one specific task, the system might close the wrong task or leave work unfinished. These tests confirmed that every part of their process is necessary; no single check can replace the others. The system works because it forces a chain of trust where the artificial intelligence can suggest, but the system decides, and only verified reality can change the record.
This work does not claim to make artificial intelligence perfect or to solve every possible failure. The system relies on the initial contract or set of rules being correct; if the rules themselves are wrong, the system will faithfully follow the wrong path. It also cannot undo actions that have already been performed in the real world if the proof of those actions is lost. However, it provides a robust framework for managing the moment when things go wrong. By treating the recovery process as a strict transition from a known, safe state to a new, verified state, AgentMirror offers a way to keep intelligent agents working safely even when they encounter errors or malicious interference. The study concludes that the key to safe continuation is not better guessing, but better verification.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.