Assessing the Reliability of LLM-Generated Phenotype-Genotype Associations Through External Validation
This study benchmarks four large language models on generating phenotype-genotype associations, finding that while they can produce a substantial volume of candidates with moderate to strong external validation support, their accuracy varies significantly by association type and model, underscoring the critical need for rigorous validation pipelines.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
In the world of modern medicine, a patient's health story is often written in two different languages. One language describes what the doctor sees: a high fever, a specific heart rhythm, or a particular shape of a bone. Scientists call these observable traits phenotypes. The other language is hidden inside the body's cells, written in the code of DNA. This code contains genes and tiny variations in the genetic script that can explain why a person has a certain trait or why they might develop a specific disease. Connecting these two languages—matching a visible symptom to its invisible genetic cause—is the foundation of precision medicine. It allows doctors to diagnose rare conditions, predict risks, and choose treatments that fit a person's unique biology. However, the library of known connections between symptoms and genes is growing so fast that human experts cannot read every new scientific paper to update their records. The gap between what is published and what is officially recorded is widening, leaving doctors without the latest answers they need.
To bridge this gap, researchers are turning to artificial intelligence, specifically a type of computer program known as a large language model. These programs have read vast amounts of text and can generate new sentences that look and sound like human writing. The idea is that if these programs have read enough medical literature, they might be able to instantly suggest which genes or genetic variations are linked to a patient's symptoms, acting as a super-fast research assistant. But there is a serious concern: these programs are known to sometimes invent facts. They can produce a gene name that sounds real but does not exist, or link a real gene to a disease it has nothing to do with. If a doctor relies on such a mistake, it could lead to a wrong diagnosis. The critical question is whether these artificial intelligence tools can actually find real, verified connections between symptoms and genes, or if they are mostly guessing.
A team of researchers at Vanderbilt University set out to test this question with a rigorous experiment. They did not just ask the computers to guess; they built a system to check every single answer against the most trusted databases in the world. They selected four of the most advanced language models available and asked them to perform six different types of tasks. In some tasks, the computer was given a list of symptoms and asked to name the genes or genetic variations that cause them. In others, the computer was given a gene or a genetic variation and asked to describe the symptoms it causes. They tested these models on both common conditions, like high blood pressure or diabetes, and very rare diseases that affect only a handful of people. The researchers then took every single association the models generated and ran it through a verification pipeline. They checked if the gene names were real, if the genetic variations existed in standard records, and most importantly, if the link between the symptom and the gene was supported by evidence in major scientific catalogs.
The results revealed a complex picture of what these tools can and cannot do. Overall, the artificial intelligence models were surprisingly good at generating names that were real. Nearly all the genes and genetic variations they suggested actually existed in the biological world; the computers rarely made up fake names. However, the real test was whether the connection between the symptom and the gene was true. When the researchers checked the evidence, they found that about 74 percent of the associations the models generated could be matched to at least one established scientific database. This means that for nearly three out of every four suggestions, there was some proof in the literature to back it up. Of those, about 9 percent had very strong support, while another 54 percent had moderate support. This suggests that these tools can indeed act as powerful generators of candidate ideas for doctors and scientists to investigate.
Yet, the performance was not uniform across all types of questions. The models performed much better when asked about genes than when asked about specific genetic variations, known as single nucleotide polymorphisms or SNPs. When the task was to link a symptom to a gene, the models were able to find strong or moderate evidence for about two-thirds of their answers. But when the task shifted to linking a symptom to a specific genetic variation, the success rate dropped significantly. The models struggled most with rare diseases. While they could often find connections for common conditions like heart disease or diabetes, they frequently failed to find verified links for ultra-rare disorders. This highlights a limitation not just in the artificial intelligence, but in the databases themselves. The scientific records for rare diseases are often incomplete or scattered, making it harder for the computer to find a match even if one exists.
The study also compared the four different models to see if one was clearly superior. One model, Claude Sonnet, performed the best overall, successfully generating associations that were supported by evidence about 69 percent of the time. Another model, GPT-5.5, came in second with about 65 percent. The other two models trailed behind, with success rates around 61 percent and 57 percent. Interestingly, the models that were best at finding gene connections were not always the same ones that were best at finding genetic variation connections. This suggests that each model has learned different patterns from the vast amounts of text it was trained on. The researchers also looked at how often the different models agreed with each other. They found that the models rarely produced the exact same list of answers, even when given the same question. Sometimes they agreed on well-known facts, but for many other questions, they offered completely different suggestions. This lack of agreement means that relying on a single model is risky; if one model suggests a link, it does not guarantee it is correct, and if two models disagree, it does not mean both are wrong.
Perhaps the most important finding was the nature of the errors the models made. When the researchers manually reviewed the suggestions that could not be found in any database, they discovered that the majority were not just random guesses. About 60 percent of the failed matches were cases where the model linked a real gene or variation to the wrong disease. The other failures were often due to the fact that the scientific databases simply did not have the information yet. This distinction is crucial. It means that while the models are not usually inventing fake names, they are frequently confident about the wrong connections. This creates a specific danger for clinical use: a doctor cannot tell the difference between a correct, verified link and a confident hallucination just by looking at the output.
The researchers concluded that these artificial intelligence tools are useful, but they must be treated with caution. They are excellent at generating a long list of possibilities, acting as a starting point for human experts. However, they cannot be trusted to provide the final answer. Every suggestion they make needs to be checked against established records before it can be used in a medical setting. The study showed that the reliability of the output depends heavily on the type of question being asked. For common diseases and gene-level questions, the tools are quite reliable. For rare diseases and specific genetic variations, they are much less certain, reflecting the current gaps in our collective scientific knowledge. The path forward involves using these models to speed up the discovery process, but always pairing their output with a strict verification step. In this way, the speed of artificial intelligence can be combined with the accuracy of human-curated science, ensuring that the connections made between a patient's symptoms and their genes are both fast and true.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.