An Empirical Investigation of Multi-Trial Consistency, Trajectory Pathologies, and Reliability Rankings in Software Engineering Agents
This paper proposes a comprehensive multi-trial evaluation framework to assess the consistency, reliability, and trajectory pathologies of autonomous software engineering agents, demonstrating its operational feasibility through a pilot study that reveals significant infrastructure censoring rates and establishes a foundation for moving beyond single-trial success metrics.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the modern world of software development, vast libraries of code are maintained by teams of engineers who rely on automated tools to find and fix bugs. Recently, a new generation of artificial intelligence has emerged, capable of acting as autonomous assistants that can read these codebases, understand problems, and attempt to write the necessary fixes on their own. These systems, often powered by large language models, have shown they can solve complex coding tasks in controlled tests. However, a critical question remains unanswered: can these digital workers be trusted to do the same job reliably every single time? In the real world, where software updates are deployed automatically, an assistant that succeeds once but fails the next time it is asked to do the exact same thing creates chaos. It introduces unpredictability, forcing human developers to constantly check the work, which defeats the purpose of automation. The core challenge is not just whether an AI can solve a problem, but whether it can solve it consistently without getting confused, making random errors, or changing its approach in ways that break the system.
A researcher at the University of Engineering and Technology in Lahore, Pakistan, has proposed a new way to measure this reliability, moving beyond the current standard of simply counting how often an AI gets a single attempt right. The current method, known as a single-trial pass rate, treats an AI like a student taking a one-time exam: if the answer is correct, the student passes, regardless of whether they could have solved it again. This new study argues that for software engineering, this approach is insufficient. Instead, the researcher designed a rigorous protocol to test these agents multiple times on the same problem, looking for consistency. The goal was to see if the AI could navigate a code repository, find a bug, and fix it in exactly the same way every time it was asked, or if its performance would crumble under repeated scrutiny.
To test this, the researcher set up a large-scale experiment involving a specific set of real-world coding tasks drawn from popular open-source projects. The study planned to run these tasks across a grid of different AI models and different software frameworks, repeating each task five times to gather a complete picture of performance. Before running the full experiment, the researcher conducted a smaller "feasibility pilot" to ensure the testing machinery worked correctly. This pilot involved planning 360 attempts across 30 different coding problems using two specific AI models and two different software scaffolding systems. However, only 312 of these attempts were successfully completed and analyzed, while 48 were interrupted by external factors. The researchers carefully tracked every step the AI took, recording not just whether it succeeded or failed, but also how many times it had to change its mind, how often it made mistakes while trying to use computer tools, and how often the testing infrastructure itself broke down.
The pilot study revealed that the testing environment itself is fragile. Out of the 360 planned attempts, 48 were interrupted by external factors, such as the cloud service provider timing out or the computer container resetting unexpectedly. This resulted in a cancellation rate of 13.3 percent, a finding that highlights the difficulty of running these complex tests reliably. More importantly, the pilot showed that the current testing budget was too short to produce successful repairs. Because the AI was limited to just five turns of conversation or action per attempt, none of the 312 completed episodes resulted in a successful fix. The agents spent their entire limited time simply trying to navigate the code and understand the problem, never reaching the stage where they could apply a solution.
Despite the lack of successful repairs in this short pilot, the study successfully demonstrated that it is possible to collect detailed data on how these agents behave. The researchers identified specific types of errors that occur when the AI tries to interact with the computer system, such as generating commands that the computer cannot understand or trying to access files that do not exist. They also developed new ways to measure "code churn," which tracks how much the AI changes its own work before settling on a final answer. A high amount of churn suggests the agent is indecisive or unstable, constantly rewriting its own code without a clear plan. The study also introduced a metric for "tool failure," distinguishing between errors caused by the AI's poor reasoning and errors caused by the computer system crashing.
The paper concludes that while these AI agents show promise, the current method of evaluating them is incomplete. By focusing only on single attempts, the industry misses the "flakiness" that makes these tools unreliable for real-world use. The researcher argues that a truly reliable agent must be able to solve a problem consistently across multiple trials, not just get lucky once. The pilot study proved that the necessary tools to measure this consistency exist and that the infrastructure can handle the data collection, even if the current AI models are not yet ready to pass the full test. The findings suggest that future evaluations must move beyond simple pass-or-fail scores to include detailed logs of the AI's journey, measuring how often it stumbles, how much it hesitates, and how often it fails to complete a task due to external interruptions.
The ultimate goal of this work is to establish a new standard for trust in automated software engineering. Just as a human employee would be fired for being inconsistent and unreliable, an AI assistant must prove it can perform its duties with the same stability. This study does not claim that current AI models have solved the problem of automated repair; in fact, the pilot data showed zero successful repairs under strict conditions. Instead, it provides the blueprint for how to properly test these systems in the future. By defining new metrics for consistency and tracking the specific pathologies that lead to failure, the researcher has laid the groundwork for a more honest and rigorous assessment of artificial intelligence in software development. The path forward involves running the full-scale experiment with longer time limits and more powerful models to see if the consistency improves, or if the agents simply become more sophisticated in their errors. Until then, the industry must recognize that a single successful fix is not enough to guarantee safety or reliability in the complex world of software maintenance.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.