← Latest papers
💻 computer science

Occam’s Razor in AI-assisted complex diagnosis: a comparative effectiveness study of single large language models versus multi-agent systems in resource-constrained primary care settings

Contrary to the prevailing assumption that multi-agent systems outperform single models, this study demonstrates that in resource-constrained primary care settings, a single high-performance local LLM (GPT-oss-20b) achieves superior diagnostic accuracy, safety, and latency compared to complex multi-agent architectures, thereby advocating for "Lean AI" strategies over computationally expensive ensemble workflows.

Original authors: Tengfei Cai, Naiguang Zhang, Yansheng Li, Yanmin Li, Xiaoyan Li, Bo Liu

Published 2026-09-25
📖 5 min read🧠 Deep dive

Original authors: Tengfei Cai, Naiguang Zhang, Yansheng Li, Yanmin Li, Xiaoyan Li, Bo Liu

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the quiet corners of the world's healthcare systems, from rural clinics in low-income nations to understaffed hospitals in developing regions, a silent struggle often plays out. A patient arrives with a confusing mix of symptoms that do not point clearly to a single disease. The local doctor, who is the first and often only line of defense, faces a difficult choice: guess based on limited experience or wait for a specialist who may be hundreds of miles away. This gap between the symptoms a patient presents and the correct diagnosis is a major driver of global health inequality. For decades, the medical community has hoped that artificial intelligence could bridge this divide, acting as a tireless, knowledgeable assistant that can reason through complex cases. The prevailing belief in the technology world has been that to solve such difficult problems, you need a team of digital minds working together. The idea is that if you combine the reasoning of several different computer programs, they can correct each other's mistakes and arrive at a truth that a single program might miss. This approach, known as a multi-agent system, is seen as the future of complex decision-making, mimicking a board of human experts debating a case until they reach a consensus.

However, a new study challenges this assumption, suggesting that in the high-stakes world of medical diagnosis, more agents do not necessarily mean better answers. Researchers from Weifang People's Hospital and iMEDWAY Technology conducted a rigorous test to see if these complex teams of artificial intelligence actually outperform a single, powerful computer program when working in a resource-limited environment. They set up a simulation that mirrored a real-world scenario where internet access is unreliable and data must stay on local servers for privacy. They pitted five different single artificial intelligence models against two different team-based systems. The single models were large, sophisticated programs capable of reading and understanding medical text, while the team systems were designed to route cases between different models and vote on the final diagnosis. The goal was to see which approach was more accurate, safer, and faster when dealing with 915 complex medical cases that had already been solved by human experts.

The results were surprising and counter to the popular trend in computer science. The study found that the single, most capable artificial intelligence model was significantly better at diagnosing these complex cases than the adaptive team-based system, which dynamically routed cases between models. However, the standard team system, which simply aggregated votes from multiple models, performed on par with the single model, showing no statistically significant difference in accuracy. The researchers observed a phenomenon they called "ensemble degradation," where the inclusion of weaker models in the adaptive team diluted the correct reasoning of the strongest model. Instead of correcting errors, the weaker models introduced confusion and incorrect guesses that pulled the final answer away from the truth. It was as if adding a few less-experienced voices to a room of experts only served to muddy the conversation rather than clarify it.

Beyond accuracy, the study highlighted the practical advantages of the simpler approach. The single model was dramatically faster, taking an average of 30 seconds to analyze a case, whereas the team systems took much longer, with one taking up to 200 seconds per case. This speed difference is critical in a busy clinic where time is a scarce resource. Furthermore, the single model required far less computing power. While the team systems needed a massive server with four high-end graphics cards to run simultaneously, the single model could theoretically run on a much smaller, cheaper setup that a local hospital could actually afford. This finding suggests that for global health, the path forward is not to build more complex, expensive networks of artificial intelligence, but to focus on perfecting a single, robust tool that can work offline and reliably.

The researchers also looked at safety, which is the most important factor in medical care. They found that the single model was less likely to make dangerous errors or "hallucinate," which is when an artificial intelligence invents facts that sound real but are false. The adaptive team system, by mixing in outputs from less capable models, produced more of these dangerous errors. The study concluded that in the specific context of diagnosing complex diseases, simplicity is superior. A single, highly capable model provides a safer, more accurate, and more cost-effective solution than orchestrating a swarm of agents, particularly when that swarm includes weaker models that degrade performance. This approach, which the authors describe as "Lean AI," offers a realistic path to bringing high-quality diagnostic support to the parts of the world that need it most, without the burden of expensive infrastructure or unstable internet connections. The study does not claim that artificial intelligence has solved all medical mysteries, but it does provide strong evidence that for now, the best way to help a doctor in a remote clinic is to give them one very smart tool, rather than a complicated team of digital assistants.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →