MARCA: Multi-Agent Root Cause Analysis with Multi-Modal Data
The paper proposes MARCA, a multi-agent framework that addresses the challenges of root cause analysis in distributed systems by using a controller-executor-voter architecture to achieve high accuracy and efficiency while reducing token usage and ensuring data privacy.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Modern software systems are no longer single, monolithic machines; they are vast, intricate cities of tiny, independent programs called microservices that talk to one another constantly. When one part of this digital city stumbles, the entire network can grind to a halt, causing everything from delayed payments to crashed websites. Finding out exactly which tiny program caused the collapse is a task known as Root Cause Analysis. For decades, engineers have tried to automate this search, but the sheer volume of data—logs, performance numbers, and error codes—makes it difficult for traditional tools to keep up. Recently, powerful computer programs known as Large Language Models have shown promise in reading and understanding these messy data streams, much like a human expert would. However, these models face their own hurdles: they can get overwhelmed by too much information at once, they are expensive to run, and sending sensitive company data to outside servers raises serious privacy concerns.
A researcher at the University of Luxembourg has developed a new approach to solve these problems, called MARCA. Instead of relying on a single, massive computer program to read everything and guess the answer, MARCA breaks the job down into a small team of specialized digital workers. Imagine a detective team where one person directs the investigation, another gathers specific clues, and a third weighs the evidence to make a final call. This team works together in a loop, asking questions, gathering data, and refining their theory until they are confident they have found the true source of the failure. By dividing the work, the system avoids getting bogged down by massive amounts of data, keeps the sensitive information on local servers for privacy, and uses less computing power than previous methods.
The researcher tested this multi-agent system on a simulated payment platform they built themselves, as well as on public benchmarks used by other scientists. They injected various types of faults into the system, such as overloading the processor, filling up the memory, or cutting off network connections, to see if MARCA could correctly identify the culprit. The results were clear: the new system found the root cause correctly about 77 percent of the time, a significant improvement over the best existing methods, which hovered around 62 percent. It also proved much better at distinguishing between different types of failures, such as telling the difference between a network slowdown and a memory shortage, achieving a high score in accuracy that suggests it can handle the messy, overlapping signals found in real-world systems.
What makes this approach distinct is how it handles the flow of information. Traditional methods often try to feed all the available data into a single model at once, which can confuse the system or force it to ignore important details to fit everything in. MARCA, by contrast, acts more like a focused investigator. It starts at the point where the problem was noticed and then systematically checks the connections leading back to the source. At each step, a "controller" agent decides what to look at next, an "executor" agent uses specialized tools to fetch the relevant logs or performance numbers, and a "voter" agent combines these findings to update the list of suspects. This process repeats until the evidence points strongly to one specific service. This method allows the system to ignore irrelevant data, keeping the task manageable and efficient.
The study also highlighted that simply using a more powerful language model is not enough to solve the problem. When the researcher compared their team-based approach to a single, powerful model trying to do the whole job alone, the team-based system still performed significantly better. This suggests that the structure of the investigation—the way the agents collaborate, filter out noise, and weigh different types of evidence—is more important than the raw intelligence of the individual model. The system was also designed to be adaptable; it can learn from past mistakes to adjust how much it trusts different types of data, such as whether to rely more on error codes or performance metrics, depending on what has worked best in previous scenarios.
While the results are promising, the researcher was careful to note the boundaries of their work. The system relies on having an accurate map of how the services are connected; if that map is missing or outdated, the investigation can go off track. Additionally, in cases where multiple problems happen at the exact same time, the system sometimes struggles to separate the overlapping symptoms. However, the core finding stands: by organizing the analysis into a collaborative, iterative process, it is possible to diagnose complex software failures more accurately and efficiently than before. This approach offers a practical path forward for keeping large-scale digital systems running smoothly, ensuring that when things go wrong, the right answer can be found quickly without compromising data security or overwhelming computing resources.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.