← Latest papers
💻 computer science

RheumBench: A Specialist-Validated Rubric-Based Benchmark for Evaluating Large Language Models in Rheumatology

RheumBench is the first specialist-validated, rubric-based benchmark for rheumatology that evaluates ten LLMs and RAG systems, revealing that current models exhibit substantial performance gaps, systematic overconfidence, and potentially severe omission errors, thereby underscoring the critical need for human oversight in clinical decision support.

Original authors: Fabian Lechner, Jonathan Bamberger, Phillip Kremer, Jutta Richter, Matthias Schneider, Uta Kiltz, Sebastian Kuhn, Lucie Flek, Philipp Klemm, Axel J Hueber, Martin Krusche, Hannah Labinsky, Johannes Kn
Published 2026-09-01
📖 5 min read🧠 Deep dive

Original authors: Fabian Lechner, Jonathan Bamberger, Phillip Kremer, Jutta Richter, Matthias Schneider, Uta Kiltz, Sebastian Kuhn, Lucie Flek, Philipp Klemm, Axel J Hueber, Martin Krusche, Hannah Labinsky, Johannes Knitza

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the modern clinic, doctors increasingly turn to artificial intelligence to help sort through complex symptoms, weigh treatment options, and make difficult decisions. These tools, known as large language models, are powerful computer programs trained on vast amounts of text that allow them to understand and generate human-like language. They can read medical records, answer questions about diseases, and suggest next steps for patient care. While these systems offer the promise of faster, more accessible support, they also carry significant risks. If a computer program gives a confident but wrong answer, it could lead to a patient receiving the wrong treatment or missing a critical diagnosis. The field of rheumatology, which deals with diseases of the joints, muscles, and immune system, is particularly complex, involving hundreds of different conditions that often look alike. Because the stakes are so high, experts need a reliable way to test whether these digital assistants are actually safe and accurate before trusting them with real patients.

To address this need, a team of researchers from several German universities and medical centers created a new testing ground called RheumBench. This is the first specialized evaluation system designed specifically to measure how well artificial intelligence performs in rheumatology. The researchers did not simply ask the computers to answer multiple-choice questions; instead, they built a rigorous framework based on real-world clinical scenarios. They gathered one hundred detailed patient cases, ranging from seventy diagnostic puzzles where the goal was to identify the correct disease, to thirty management situations that required advice on treatment, counseling, and follow-up care. For each of these cases, a panel of board-certified rheumatologists created a detailed scoring guide, or rubric. This guide listed specific actions that a good doctor should take, such as ordering a specific test or prescribing a certain medication, as well as actions that should be avoided. Each point in the guide was weighted by importance, with some actions earning high points for being correct and others losing points for being dangerous or inappropriate.

The researchers then tested ten different artificial intelligence systems using these one hundred cases. The group included three of the most advanced general-purpose computer models available, one open-source model, and six systems specifically designed for medical use, including two that are officially certified as medical devices and one built specifically for rheumatology. The computers were asked to provide answers to the patient cases, and their responses were evaluated against the expert-created rubrics. To handle the massive amount of grading required, the researchers used a panel of three other advanced computer models to act as judges, checking each answer against the scoring guide. This process generated over fifty thousand individual evaluations, which were then compared against the judgments of human rheumatologists to ensure the computer judges were accurate. The results showed a very high level of agreement between the human experts and the computer judges, validating the method.

The findings revealed a landscape of significant variation in performance. The most advanced general-purpose computer models, which are not specialized for medicine, actually outperformed the systems designed specifically for clinical use. The best-performing system achieved a score of 54.0%, meaning it correctly followed the expert guidelines just over half the time. In contrast, the system specifically built for rheumatology scored the lowest of all, at 17.8%. Even the certified medical devices, which are regulated as medical products, did not consistently outperform the general models. This suggests that having a specialized focus or official medical certification does not automatically guarantee that an artificial intelligence system will provide better or safer advice. The study also found that the way a question was asked made a big difference. When the researchers used more detailed and thorough prompts, the performance of most systems improved, with one model reaching a score of 72.9%. However, even with these improvements, no system reached a level of perfection that would allow it to operate without human supervision.

Safety was a major concern throughout the testing. The researchers looked for errors that could cause serious harm to a patient. They found that potentially severe mistakes occurred in every single system tested, including the certified medical products. The most common type of error was an omission, where the computer failed to suggest a necessary action, such as ordering a crucial test or referring a patient to a specialist. While the systems were generally good at avoiding actions that would directly cause harm, their failure to act when needed was a significant risk. Furthermore, the computers displayed a troubling pattern of overconfidence. They frequently reported high levels of certainty in their answers, often claiming to be more than ninety percent sure, even when their actual performance was much lower. This gap between how sure the computer felt and how right it actually was was present across all systems, indicating that users cannot rely on the computer's own confidence rating to judge the quality of its advice.

The study concludes that while artificial intelligence holds promise for supporting doctors in rheumatology, current systems are not yet ready to replace human judgment. The fact that general-purpose models often performed better than specialized medical tools challenges the assumption that medical certification or niche training automatically leads to superior safety or accuracy. The researchers emphasize that human oversight remains essential, particularly because these systems can make serious errors and often do not realize when they are wrong. The RheumBench framework will continue to be updated with more cases and scenarios to help track improvements and ensure that future versions of these tools are safer for patients. Until then, the role of the computer should be viewed as a supportive tool that requires careful checking by a trained specialist, rather than an autonomous decision-maker.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →