← Latest papers
📄 allergy and immunology

Can Large Language Models Diagnose Primary Immunodeficiency from Patient-Described Symptoms?

This study reveals that while large language models can effectively identify primary immunodeficiency from physician-written histories, their diagnostic accuracy drops significantly when relying on patients' own descriptions of symptoms, highlighting a critical gap in performance based on the language and framing of medical input.

Original authors: Reteig, L. C., Woloshin, S., Maglione, P. J., Farmer, J. R., Ong, M.-S.

Published 2026-10-04
📖 6 min read🧠 Deep dive

Original authors: Reteig, L. C., Woloshin, S., Maglione, P. J., Farmer, J. R., Ong, M.-S.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

For millions of people, the path to a medical diagnosis is not a straight line but a long, confusing journey. This is especially true for those with rare conditions, where symptoms might be vague, common, or easily mistaken for something else. In recent years, many of these individuals have turned to a new kind of helper: large language models. These are advanced computer programs trained on vast amounts of text, capable of holding conversations and answering questions about health. They have become a go-to resource for people trying to make sense of their own bodies, often before they ever see a specialist. But a critical question remains: can these digital tools truly understand a patient when the patient is the one describing the problem? The answer depends heavily on who is speaking. When a doctor summarizes a case using precise medical terms, the computer understands perfectly. When a regular person describes their pain, fatigue, or recurring infections in everyday language, the computer often gets lost.

A team of researchers set out to test this gap between expert language and patient language, focusing on a group of rare disorders known as primary immunodeficiencies. These are conditions where the body's natural defense system, which usually fights off germs, is broken or missing. People with these conditions get sick very often, sometimes with severe infections that require hospitalization, and they face a high risk of other complications like autoimmune diseases. Because these conditions are rare and complex, patients often wait years for a correct diagnosis, a delay that can be dangerous and emotionally draining. During this waiting period, many turn to chatbots for answers. The researchers wanted to know if these chatbots could spot the signs of a broken immune system when the description came directly from the patient, rather than from a doctor.

To find out, the researchers gathered stories from twenty-one adults who had already been diagnosed with primary immunodeficiency. These individuals had all experienced long delays in getting their diagnosis. The team asked them to describe the very first symptoms that made them feel something was wrong. The researchers took these raw, unedited descriptions—exactly as the patients said them, with all their hesitation, repetition, and everyday words—and fed them into a powerful large language model. They stripped away any mention of the final diagnosis so the computer had to rely on the symptoms alone. The goal was simple: would the computer suggest that the patient might have a primary immunodeficiency, or at least recognize that their immune system was the problem?

The results revealed a stark difference between how the computer performs with expert input versus patient input. In a previous study using the same computer model, when doctors wrote the patient histories using professional medical language, the computer correctly identified the condition in ninety-six percent of cases. It was nearly flawless. However, when the researchers used the patients' own words, the computer's performance dropped dramatically. It correctly identified primary immunodeficiency in only seven of the twenty-one cases, which is about one-third. The computer failed to name the specific condition in the majority of cases, even though the patients were describing the exact same illnesses.

Despite missing the specific diagnosis, the computer was not entirely unhelpful. In seventeen of the twenty-one cases, it suggested that the patient might have general issues with their immune system, even if it did not name the specific rare disease. This is a significant finding because it suggests the tool can still act as a useful signal. For a patient who has been told by multiple doctors that their symptoms are just stress or allergies, a computer suggesting "immune system issues" might encourage them to seek a specialist. However, the study also showed that the computer often guessed other, more common conditions. It frequently suggested allergies, environmental factors like poor air quality, or endocrine problems like thyroid issues. These are the kinds of things doctors consider first, but they can also be dead ends for someone who actually has a rare immune disorder.

The researchers found that the computer was most likely to get the diagnosis right when the patient's story included three specific clues: infections that started in childhood, very serious infections that required hospitalization, or symptoms that were not infections at all, such as failure to grow or digestive problems. When patients described their struggles in these clear, classic ways, the computer had a better chance of connecting the dots. But when the descriptions were more subtle or focused on the general feeling of being unwell, the computer struggled. This highlights a fundamental weakness: the tool is highly sensitive to how the story is told. It thrives on the structured, precise language of medicine but falters when faced with the messy, emotional, and varied way people actually speak about their health.

The study authors are careful to note that this is not a final verdict on the safety or utility of these tools, but rather a warning about how they are currently used. They point out that their test was a best-case scenario. The patients they interviewed knew their diagnosis and might have remembered their symptoms more clearly than they would have at the very beginning of their illness. If the computer performs this poorly with clear, retrospective accounts, it might perform even worse with confused, pre-diagnosis patients who are unsure what is wrong. Furthermore, the study did not test how often the computer would give false alarms to healthy people, which is a major concern. If the tool flags too many healthy people as having immune problems, it could overwhelm doctors and delay care for those who truly need it.

Ultimately, this research underscores a growing challenge in modern medicine. As more people turn to artificial intelligence for health guidance, the gap between what the technology can do and how people actually use it becomes a safety issue. The computer is not broken; it is simply trained to understand the language of experts, not the language of patients. For the millions of people navigating a diagnostic odyssey, the promise of these tools remains unfulfilled until they can truly listen to the person in the room, not just the medical chart. The path forward requires building systems that can bridge this divide, ensuring that when a patient describes their suffering in their own words, the answer they get is not just a guess, but a reliable guide toward the care they need.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →