Do No Harm? Hallucination and Actor-Level Abuse in Web-Deployed Medical Large Language Models
This paper presents a large-scale assessment of over 6,000 web-deployed medical large language models, revealing significant risks of hallucinations, policy violations, and privacy gaps while introducing new evaluation frameworks and a dataset to guide future safety improvements.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the internet's "App Store" for AI, but instead of games or productivity tools, it's filled with thousands of digital doctors. These are custom AI chatbots (called MedGPTs) that anyone can build and publish to give medical advice, diagnose symptoms, or explain treatments.
The paper you provided is like a massive, undercover health inspection of this digital marketplace. The researchers didn't just look at the top-rated, popular doctors; they audited over 6,000 of them to see if they are safe, honest, or if they are actually dangerous.
Here is what they found, broken down into simple concepts:
1. The "Confident Liar" Problem (Hallucinations)
Imagine a tour guide who speaks with absolute confidence but makes up facts about history. In the medical world, this is called a hallucination.
- The Finding: About 25% to 30% of these AI doctors are "lying" or making up medical facts. They might confidently tell you to take a dangerous drug or ignore a serious symptom.
- The Twist: You might think the most popular doctors are the safest. But the study found that popularity is a bad lie detector. The "famous" AI doctors were just as likely to be making things up as the obscure ones. In fact, the less popular ones were often even worse.
- The User Blind Spot: The researchers found that users don't seem to notice these lies. If an AI gives a smooth, confident answer, people give it 5 stars, even if the advice is medically wrong. It's like giving a 5-star rating to a chef who serves you a delicious-looking but poisonous meal.
2. The "Bad Actor" Problem (Design Abuse)
Some of these AI doctors aren't just accidentally wrong; they are designed to be risky.
- The Finding: Nearly 50% of the AI doctors had "red flags" in how they were built.
- The Metaphor: Imagine a restaurant owner who puts up a sign saying "We are a hospital" (even though they aren't), hides the menu, and refuses to show you where the food comes from.
- The Specifics:
- Some creators named their bots things like "Instant Cancer Diagnosis" to trick people into thinking they are real doctors.
- Many of these bots have a feature called "Actions" (which lets them connect to outside websites), but 57% of them didn't have a privacy policy. It's like walking into a clinic where the doctor takes your blood sample but refuses to tell you what they will do with it.
- Even when they did have a privacy policy, nearly 70% of them were broken, vague, or didn't actually protect your data.
3. The "Open Source" vs. "Store-Bought" Showdown
The researchers also compared these store-bought AI doctors to "open-source" ones (models that developers share freely, like a community recipe book).
- The Store-Bought (MedGPTs): These were generally better at sounding smart and accurate (they got the facts right more often). However, they were less stable—sometimes they gave a great answer, and the next time you asked the same question, they gave a totally different, confusing answer.
- The Open-Source: These were more consistent (they gave the same answer every time), but they were less accurate on the facts.
- The Lesson: Neither type is perfect. The "fancy" ones sound better but wobble; the "community" ones are steady but might get the details wrong.
4. The "Trust Signals" Are Broken
In a normal app store, if an app has bad reviews or low ratings, you avoid it.
- The Finding: In this medical AI store, ratings and reviews tell you nothing about safety.
- The Metaphor: It's like a car dealership where the cars with the most test drives and the highest "5-star" ratings are actually the ones with broken brakes. The researchers found that how much people talked to the AI had almost zero connection to whether the AI was actually telling the truth.
The Bottom Line
The paper concludes that we cannot trust these AI doctors just because they are popular or because they sound confident.
- The System is Broken: The current way of checking these bots (relying on user reviews and star ratings) is failing.
- The Danger: Because patients often don't have medical training, they can't spot when the AI is lying or when the doctor is "faking" authority.
- The Solution: We need a new kind of "health inspector" that checks the AI's facts, its design, and its privacy rules before it is allowed to talk to patients. The researchers have released a new dataset and a set of tools to help build these inspectors.
In short: The digital medical marketplace is full of confident liars and risky designs, and the "customer reviews" aren't helping us spot them. We need better safety checks before we let these AI doctors into our lives.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.