← Latest papers
📊 statistics

Errors in AI-Assisted Retrieval of Medical Literature: A Comparative Study

This study quantitatively reveals that widely used free-version large language models exhibit significant variability and a high rate of complete failure (47.8%) when retrieving accurate medical literature references, underscoring the critical need for human verification of bibliographic data generated by AI.

Original authors: Jenny Gao (College of Arts and Science, New York University, New York, NY), Yongfeng Zhang (Department of Computer Sciences, School of Arts & Sciences, Rutgers University, Piscataway, NJ), Mary L Disi
Published 2026-03-25
📖 4 min read☕ Coffee break read

Original authors: Jenny Gao (College of Arts and Science, New York University, New York, NY), Yongfeng Zhang (Department of Computer Sciences, School of Arts & Sciences, Rutgers University, Piscataway, NJ), Mary L Disis (UW Medicine Cancer Vaccine Institute University of Washington, Seattle, WA), Lanjing Zhang (Department of Chemical Biology, Ernest Mario School of Pharmacy, Rutgers University, Piscataway, NJ, Department of Pathology, Princeton Medical Center, Plainsboro, NJ, Rutgers Cancer Institute, New Brunswick, NJ)

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a student trying to write a research paper. You ask a super-smart, all-knowing robot librarian for a list of 10 books to help you. The robot hands you a list with titles, authors, and library call numbers. You look happy and start reading. But then, you realize: half the books don't actually exist. Some have the right title but the wrong author. Others have the right author but the wrong year. Some call numbers lead to empty shelves.

This is exactly what a new study by Jenny Gao and Lanjing Zhang discovered about AI "librarians" (called Large Language Models or LLMs) when they try to find medical research papers.

Here is the breakdown of their findings, translated into everyday language:

The Big Experiment: The "Robot Librarian" Test

The researchers decided to put five popular free AI tools to the test: Grok, ChatGPT, Google Gemini, Perplexity, and DeepSeek.

They picked 40 real, high-quality medical articles from four famous journals (like the New England Journal of Medicine and JAMA). They asked each AI: "Here is a summary of this article. Please find me 10 other important papers related to this topic, and give me the exact details (like the DOI, PubMed ID, and links) so I can find them."

The AI tools generated 2,000 references in total. The researchers then acted like strict librarians, checking every single one to see if it was real and correct.

The Shocking Results: The "Hallucination" Problem

The results were a bit scary for anyone trusting AI to do their homework:

  • The "Complete Fail" Rate: Nearly 48% of the time, the AI gave a reference that was a total disaster. It was either a fake paper that never existed, or the details were so wrong you couldn't find the paper at all. It's like the robot handing you a map to a treasure chest that is buried in a field that doesn't exist.
  • The "Average" Score: If you gave the AI a score from 0 to 1 (where 1 is perfect), the average score was only 0.29. That's a failing grade.
  • The Winners and Losers:
    • Grok was the "honor student," getting the most correct details (though it still made mistakes).
    • Google Gemini was the "struggling student," getting the details wrong most of the time.
    • ChatGPT and Perplexity were in the middle, but still far from perfect.

Why Did This Happen?

The study found two main reasons why the robots got confused:

  1. The "Journal" Factor: It was harder for the AI to find correct details for papers from the New England Journal of Medicine (NEJM) compared to the British Medical Journal (BMJ). The researchers suspect this is because NEJM abstracts are shorter and have less information for the AI to "grab onto" to find the right paper. It's like trying to find a specific person in a crowd when you only have their first name versus having their full name and address.
  2. The "Confidence" Trap: The AI tools were very confident. They would give you a perfect-looking list with titles, dates, and links. But often, the links were broken or the dates were wrong. This is called "Referential Hallucination." The AI is so eager to please that it makes up facts that sound true but are actually fiction.

The "Double-Check" Didn't Work

The researchers tried a clever trick. After the AI gave the list, they asked it: "Please double-check these numbers and links to make sure they are real."

Did it help? No. The AI just confidently double-checked its own mistakes. It's like asking a person who made up a story to verify their own story; they will just repeat the lie with more confidence.

The Takeaway: Use AI as a Guide, Not a Guru

So, what does this mean for you?

  • AI is a great brainstorming partner: It can help you think of topics or find general ideas.
  • AI is a terrible fact-checker: You cannot trust it to give you the "receipts" (the actual links, dates, and IDs) for medical or scientific research.
  • The Golden Rule: If an AI gives you a reference, you must go find the paper yourself and verify that it exists. Do not just copy-paste the AI's list into your work.

In short: Large Language Models are like enthusiastic tour guides who know a lot about the world but sometimes invent landmarks that aren't there. If you follow their map blindly, you might end up lost in a field that doesn't exist. Always bring your own compass (human verification) when navigating medical research.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →