Evaluating ChatGPT on Medical Information Extraction Tasks: Performance, Explainability and Beyond
This paper evaluates ChatGPT's performance across four medical information extraction tasks, finding that while it provides high-quality explanations and remains faithful to source texts, it underperforms compared to fine-tuned models and suffers from overconfidence and generation uncertainty.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The "Brilliant but Careless Medical Intern" Analogy
Imagine you have a new medical intern named ChatGPT.
ChatGPT is incredibly well-read. It has read almost every medical textbook, research paper, and patient chart ever written. If you ask it to explain how a heart works, it will give you a beautiful, poetic, and highly detailed lecture. It’s a master communicator.
However, this paper is essentially a "performance review" conducted by researchers to see if ChatGPT can actually do the gritty, precise work of a medical clerk—specifically, Medical Information Extraction (MedIE). This is the job of reading a messy doctor’s note and pulling out specific facts: "What is the disease? What medicine was given? What was the dosage?"
Here is the breakdown of the review:
1. The Performance: "The Smart Student Who Fails the Test"
The Finding: ChatGPT’s accuracy is significantly lower than specialized, "fine-tuned" models (AI that was trained only on medical data).
The Analogy: Imagine a student who is a genius at general literature but is taking a highly specialized surgical exam. They understand the language of the exam, but they keep missing the tiny, technical details. While a specialized medical AI is like a surgeon with 20 years of experience, ChatGPT is like a very smart student who is "winging it" based on general knowledge. It’s good at simple tasks (like identifying a basic disease), but it trips over complex tasks (like mapping out how a specific symptom relates to a specific drug).
2. Explainability: "The Confident Storyteller"
The Finding: When ChatGPT makes a decision, it can explain why it did it in a way that sounds very convincing.
The Analogy: If you ask the intern, "Why did you label this as a symptom?", they won't just say "Because I said so." They will give you a logical, step-by-step reasoning. The problem? Even when they are completely wrong, they explain their mistake with such confidence and eloquence that you might actually believe them. They are "hallucinating" logic—telling a very convincing story about a mistake.
3. Confidence vs. Reality: "The Overconfident Intern"
The Finding: ChatGPT suffers from "over-confidence." It gives high confidence scores to both its correct answers and its wrong answers.
The Analogy: Imagine an intern who walks into a room and says, "I am 99% sure this patient has Type 2 Diabetes," when in reality, they are just guessing. In medicine, this is dangerous. A good doctor should say, "I think it's Diabetes, but I'm only 60% sure; let's run more tests." ChatGPT lacks this "self-doubt" (calibration), making it hard to know when to trust it.
4. Faithfulness: "The Good Listener"
The Finding: ChatGPT is actually quite good at following instructions and staying true to the text provided.
The Analogy: Despite its other flaws, the intern is a great listener. If you give them a specific list of categories to look for, they generally stick to that list. They don't usually start making up random medical terms that weren't in the original note. They are "faithful" to the source material.
5. Uncertainty: "The Fidgety Worker"
The Finding: Because of how ChatGPT generates text, it can be inconsistent. If you ask it the same question five times, you might get slightly different answers.
The Analogy: It’s like an intern who is a bit "fidgety." If you ask them to organize a file cabinet, they might do it one way at 9:00 AM, but by 9:05 AM, they’ve slightly changed the system. This inconsistency (uncertainty) makes it hard to use them for high-stakes, repetitive medical data entry where precision is everything.
The Bottom Line
The researchers are saying: ChatGPT is a wonderful conversationalist, but a shaky medical clerk.
It is brilliant at explaining things, but it is too confident when it's wrong, too inconsistent when it's working, and not nearly as precise as specialized medical tools. For now, we shouldn't let the "intern" handle the official medical records without a very experienced "senior doctor" watching over their shoulder.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.