← Latest papers
💻 computer science

Agent-MIRUPD: Operationalizing LLM Agents for Microservices Incident Response Under Performance Degradation

The paper proposes Agent-MIRUPD, an LLM-based agent framework that unifies observability data with structured multi-step reasoning to automate the full microservices incident-response lifecycle, achieving significant improvements in diagnosis speed, mitigation accuracy, and token efficiency compared to existing baselines.

Original authors: Nayereh Rasouli, Matthijs Jansen, Alexandru Iosup, Cristian Klein, Erik Elmroth

Published 2026-09-23
📖 5 min read🧠 Deep dive

Original authors: Nayereh Rasouli, Matthijs Jansen, Alexandru Iosup, Cristian Klein, Erik Elmroth

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Modern digital services, from banking apps to healthcare platforms, increasingly rely on a complex architecture where a single application is broken into dozens of small, independent programs called microservices. These pieces run in containers, managed by an automated system known as Kubernetes, which acts like a traffic controller, ensuring that if one piece crashes, it is restarted and the service continues. While this setup offers incredible flexibility, it creates a new kind of problem: sometimes, a service does not crash completely but simply slows down or behaves erratically. These subtle issues, often called gray failures or performance degradation, are difficult to spot because the system is still technically "alive," yet it is failing to do its job properly. For human engineers, diagnosing why a specific service is lagging requires sifting through mountains of data logs, performance metrics, and network traces, a task that is slow, error-prone, and exhausting.

Researchers at Umeå University and Vrije Universiteit Amsterdam have developed a new approach to help solve this problem, creating a system that uses an artificial intelligence agent to manage the entire process of finding and fixing these issues. Instead of just looking for obvious crashes, this system is designed to detect when a service is degrading, figure out exactly why it is happening, and then propose a fix. The core of their work is a framework called Agent-MIRUPD, which acts as a digital assistant that can reason through complex situations, much like a senior engineer would, but with the speed of a computer. The researchers tested this system in a controlled environment that mimics real-world microservice failures, including scenarios where services age and slow down over time, and found that it could handle the full cycle of incident response with high accuracy.

The system operates by breaking the problem-solving process into four distinct stages: detection, localization, analysis, and mitigation. First, the agent constantly monitors the system to see if anything is wrong. If it spots a problem, it moves to the second stage, where it tries to pinpoint exactly which service is the culprit. In the third stage, it digs deeper to understand the root cause, such as a memory leak or a configuration error. Finally, in the mitigation stage, it decides what action to take to restore the system. A crucial feature of this design is how it handles uncertainty. For clear-cut problems, like a misconfigured setting, the agent can automatically apply the fix. However, for more subtle issues where the right solution is not obvious, the system pauses and asks a human operator for approval before making any changes. This "human-in-the-loop" approach ensures that the AI does not accidentally make a situation worse while trying to help.

To make this work, the researchers equipped the AI with a set of specialized tools that allow it to query the system for specific information, such as checking the status of running programs or reading error logs. Rather than feeding the AI raw data, these tools provide summarized, structured answers, which helps the AI reason more effectively without getting overwhelmed. The system also uses a technique called stage-context propagation, which means that the conclusion reached in one stage is passed directly to the next. For example, once the agent identifies a faulty service, that information is immediately used as the starting point for the analysis phase, preventing the AI from having to restart its investigation from scratch. This continuity allows the system to move through the diagnostic process much faster than previous methods.

When the researchers tested their system against existing tools and other AI agents, the results were significant. In scenarios involving functional faults, where services fail completely, the system achieved an accuracy of up to 90 percent. For the more difficult performance degradation scenarios, where services slow down but do not stop, it achieved 85 percent accuracy. The system also proved to be much faster and more efficient than its competitors. It reduced the time needed to find the root cause of a problem from 244 seconds down to 52 seconds. Furthermore, it used 51.3 percent fewer computational resources, measured in tokens, which are the basic units of data the AI processes to think. In terms of fixing the problems, the system improved the accuracy of mitigation actions from 40 percent to 70 percent compared to a standard baseline agent.

The study also highlighted the importance of human oversight. When the system was allowed to act completely on its own for performance-related issues, it sometimes struggled to choose the best fix because these problems often have multiple possible solutions, each with different trade-offs. By involving a human expert to review and approve the proposed actions, the system could select more reliable recovery strategies. The researchers found that when the AI provided a clear explanation of its reasoning and the human operator validated the choice, the overall effectiveness of the recovery increased. This suggests that the most powerful approach is not to replace human engineers but to give them a smart assistant that handles the heavy lifting of data analysis and initial diagnosis, leaving the final decision on complex fixes to human judgment.

The researchers acknowledge that their work is still a simulation and that real-world production environments are even more complex. They plan to test the system with actual industrial workloads and explore ways to make the AI more efficient by using smaller, less expensive models. However, the current findings demonstrate that it is possible to operationalize AI agents for the full lifecycle of incident response in microservices. By combining automated detection with human supervision, the system offers a practical path toward more resilient digital infrastructure, capable of handling the subtle, slow-burning failures that often go unnoticed until they cause major disruptions. The work shows that with the right design, artificial intelligence can become a trusted partner in keeping the critical systems of modern life running smoothly.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →