Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions
This study evaluates three LLM chatbots on medical question answering, finding that retrieval performance varies significantly by model and user role, with ChatGPT outperforming others and all models showing a bias toward citing larger clinical trials.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you're trying to find the best recipe for a specific dish, but instead of asking a chef, you ask a super-smart robot that has read almost every cookbook ever written. This robot is a "Large Language Model" (LLM), a type of AI chatbot that can chat, write, and answer questions by predicting the next word in a sentence. In the world of medicine, these bots are becoming popular because they can quickly summarize complex health issues and point you to the studies that back up their advice. But here's the catch: sometimes these robots make things up, or "hallucinate," citing studies that don't exist. Even worse, they might miss the most important studies entirely. This paper asks a simple but crucial question: If a human expert (like a medical researcher) has already done the hard work of finding the perfect list of studies for a specific medical question, how well can these AI robots find that same list? Do they act like a curious patient, a busy doctor, or a serious researcher? And does the size of the study matter to the robot?
The authors of this study decided to put three of the newest, most powerful AI chatbots to the test. They treated the chatbots like students taking a very specific exam. The "test questions" were 20 real medical topics taken from the 2026 Cochrane Database of Systematic Reviews, which is like the gold-standard library where expert teams have already gathered and vetted the best evidence for specific health problems. For each of these 20 topics, the researchers asked the chatbots to find the original research studies (the "primary sources") that support the answers. To see if the robot's personality mattered, they asked the same questions three different ways: once as if a patient was asking, once as if a clinician (doctor) was asking, and once as if an evidence-synthesis researcher (a scientist who specializes in finding data) was asking. They ran this entire experiment four times for every combination to make sure the results weren't just a fluke, ending up with 720 total answers to analyze.
The results were a mix of "not bad" and "surprisingly biased." On average, a single chatbot response managed to find about 39.2% of the studies that the human experts had already identified as the "correct" list. That means if the experts found 10 perfect studies, the robot usually found only about 4 of them. However, the type of robot mattered a lot. ChatGPT was the clear winner, finding 63.1% of the expert studies. Claude was in the middle with 37.0%, and Gemini struggled the most, finding only 17.3%. Interestingly, the "personality" of the user also played a small role: when the prompt was written from the perspective of a researcher, the bots found more studies (42.8%) than when it was written as a patient (36.1%) or a clinician (38.6%).
But the most fascinating discovery was about which studies the bots chose to find. The researchers looked at the characteristics of the studies the bots successfully retrieved versus the ones they missed. They found that the bots had a massive preference for studies with larger sample sizes (meaning studies that tested the treatment on more people). Even when they controlled for how new the study was, how often it was cited by others, or whether it was free to read, the size of the study was the only thing that reliably predicted whether the bot would find it. For every unit increase in the log of the sample size, the odds of the bot finding the study jumped by 1.80 times. It's as if the chatbots have a built-in bias that thinks, "If a study has a huge crowd of participants, it must be the most important one," ignoring smaller but equally valid studies.
The study also checked if the bots were making up fake studies or getting details wrong. While they didn't invent many fake studies (a good sign thanks to their advanced reasoning skills), they did make "metadata errors" about 3% of the time. This means they might cite a real study but get the author's name, the year, or the journal title wrong. Claude made these errors the most often (5.9% of the time), while Gemini was the most accurate (0.2%).
In the end, the paper suggests that while these AI chatbots are getting better at finding medical evidence, they aren't perfect replacements for human experts yet. They are incomplete, they change their answers depending on who you pretend to be, and they have a strong bias toward big, famous studies. If you use a chatbot to find medical research, you might get a list of the "big hits," but you could easily miss the smaller, quieter studies that experts know are just as important. The authors conclude that these tools should be seen as helpful assistants that can point you in the right direction, but not as the final authority on what the medical evidence actually says.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.