← Latest papers
📄 nephrology

Robustness Gap of Large Language Models in Nephrology

This study reveals a significant robustness gap in state-of-the-art large language models when evaluated on nephrology board questions using "None of the other answers" substitution, demonstrating that high multiple-choice accuracy does not necessarily reflect true clinical reasoning capabilities.

Original authors: Soejima, A., Kitano, F., Ichikawa, D., Shibagaki, Y., Noda, R.

Published 2026-08-18
📖 4 min read☕ Coffee break read

Original authors: Soejima, A., Kitano, F., Ichikawa, D., Shibagaki, Y., Noda, R.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Artificial intelligence has begun to offer support in medicine, acting as a tool that can read vast amounts of text and suggest answers to complex questions. In fields like kidney care, or nephrology, doctors must constantly weigh laboratory results, patient history, and physical signs to decide on treatments. For years, researchers have tested whether these computer programs can pass medical exams, a standard way to see if they know enough to be useful. However, a high score on a multiple-choice test does not necessarily mean the system truly understands the logic behind the answer. It might simply be recognizing patterns in the way questions are written or guessing based on which options look most familiar. The critical question for the future of safe medical care is whether these models can reason through a problem when the usual shortcuts are removed, or if they crumble when forced to think without a familiar crutch.

A team of researchers from St. Marianna University School of Medicine in Japan set out to test this specific weakness using a method that acts like a stress test for the mind. They took a collection of 145 real questions used to renew the board certification of nephrologists in Japan. These questions cover the deep, complex reasoning required to manage kidney disease. The researchers then created a modified version of every single question. In the original test, one specific answer was correct. In the new version, they replaced that correct answer with a phrase meaning "none of the other answers." This forced the computer models to do something difficult: instead of picking the right option from a list, they had to realize that every other option was wrong and that the only valid choice was the one stating that none of them fit. If a model was truly reasoning through the medical facts, it should have been able to identify that the original correct answer was missing and select the "none of the above" option. If it was just guessing based on patterns, it would likely pick one of the remaining wrong answers.

The team tested four of the most advanced language models available, including systems from OpenAI and Google. They fed the models the original questions and the modified versions separately, asking them to choose the best answer each time. The results showed a clear and significant gap between how well the models performed when the correct answer was present versus when it was hidden. When the models faced the original questions, they performed quite well, with the best system answering nearly 88 percent correctly. However, once the correct answer was swapped out for "none of the other answers," their performance dropped sharply. The system that started with the highest score still fell to about 73 percent, a statistically significant decline. Other models saw even steeper drops, with one falling from roughly 66 percent correct down to just 19 percent. This means that when the familiar pattern of a correct answer was removed, the models struggled to rely on their own medical reasoning and often chose a wrong option that looked plausible, rather than admitting that the correct choice was missing.

The study suggests that while newer models are becoming more robust and better at handling these tricky changes, they are not yet fully reliable for independent clinical decision-making. The researchers found that even the most advanced systems tend to anchor on the idea that one of the visible choices must be right, even when the logic of the case proves otherwise. In a real-world hospital setting, this behavior could be dangerous. If a doctor is missing a piece of information, such as a specific lab value or a subtle physical sign, a safe artificial intelligence should recognize that the case cannot be judged and ask for more data. Instead, these models often fill in the gaps with a confident but incorrect guess. The authors conclude that passing a multiple-choice exam is not enough to prove a model can reason like a doctor. To be truly safe for use in kidney care, these systems need to demonstrate that they can handle uncertainty and recognize when an answer is simply not available, rather than just matching patterns to find the highest score on a test.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →