← Latest papers
📄 medicine

Patient-Facing AI Chatbots vs Primary Care Physicians: A Standardized Patient Field Experiment in China

In a standardized patient field experiment across 62 Chinese primary care institutions, AI chatbots outperformed physicians in diagnostic accuracy and patient-centered communication but significantly increased the risk of low-value care by recommending excessive tests and inappropriate medications.

Original authors: Yafei Si, Shaoqing Gong, Michael Thielscher, Hazel Bateman, Bingqin Li, Shanquan Chen, Duolao Wang, Ruopeng An, Yurun Meng, Han Zhang, Suya An, Jiaqi Zu, Min Su, Yudong Miao, Zhongliang Zhou, Xiaojing
Published 2026-06-28
📖 4 min read☕ Coffee break read

Original authors: Yafei Si, Shaoqing Gong, Michael Thielscher, Hazel Bateman, Bingqin Li, Shanquan Chen, Duolao Wang, Ruopeng An, Yurun Meng, Han Zhang, Suya An, Jiaqi Zu, Min Su, Yudong Miao, Zhongliang Zhou, Xiaojing Fan, Limin Mao, Xi Chen, Gang Chen

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a massive health check-up where two very different types of "doctors" are put to the test on the same day, facing the exact same patients. One group is real-life primary care physicians working in rural China. The other group is two famous AI chatbots (ChatGPT-4o and DeepSeek-R1) acting as virtual doctors.

To make this a fair fight, the researchers used "Standardized Patients." Think of these as highly trained actors who play the role of sick people perfectly. They all have the same script: some pretend to have chest pain (unstable angina), and others pretend to have trouble breathing (asthma). They visit a real doctor, and then, on the same day, they chat with the AI. The actors never reveal they are actors, so the doctors and the AI don't know they are being tested.

Here is what happened when the results were tallied:

1. The "Test Scores": Who Got the Diagnosis Right?

The AI won the trivia contest by a landslide.

  • The Real Doctors: Only about 3 out of 10 real doctors correctly identified the illness and suggested the right medicine.
  • The AI Chatbots: They were nearly perfect. 94% to 99% of the time, they got the diagnosis and the right medicine correct.

The Analogy: Imagine a multiple-choice test. The real doctors were like students who were tired, distracted, or forgot to study the specific chapter, so they guessed wrong often. The AI chatbots were like students who had memorized the entire textbook and could instantly recall the exact answer.

2. The "Safety Hazard": Who Ordered Too Much?

The AI chatbots were "over-enthusiastic" and risky.
While the AI got the answers right, they had a major problem with how they solved the problem. They acted like a security guard who checks every single person's bag, even if they look harmless, just to be 100% sure.

  • The Real Doctors: They were careful. They only ordered tests or medicines when they really thought it was necessary.
  • The AI Chatbots: They went overboard. They recommended way more tests and way more medicines than the doctors did.
    • They ordered 90% to 96% of their patients to take unnecessary tests.
    • They suggested inappropriate medicines for about 60% to 73% of the patients.

The Analogy: If you have a small scratch on your knee, a real doctor might say, "Just wash it and put a bandage on it." The AI chatbot, trying to be super safe, might say, "You need an MRI, a blood test, a CT scan, and three different types of antibiotics just in case." It's not that the AI is wrong about the scratch; it's that it's terrified of missing anything, so it suggests everything. This is called "low-value care"—doing too much, which can be expensive and potentially harmful.

3. The "Bedside Manner": Who Was More Kind?

The AI chatbots were surprisingly better at being "patient-centered."
The researchers asked the actors to rate how well the "doctor" listened, understood their feelings, and explained things.

  • The Real Doctors: They scored average (around 2.5 out of 4). They were often focused on the medical facts and sometimes missed the emotional connection.
  • The AI Chatbots: They scored very high (around 3.3 out of 4). They were excellent at saying things like, "I understand this is scary," or "Let's look at your whole life to see what's going on."

The Analogy: The real doctors were like busy mechanics who quickly check the engine and tell you what's broken. The AI chatbots were like a friendly mechanic who not only checks the engine but also asks how your day is, explains the problem in simple words, and makes you feel heard.

The Big Takeaway

The paper concludes with a tricky situation, like a double-edged sword:

  • The Good News: In places where there aren't enough real doctors, AI chatbots could be amazing helpers. They can spot diseases better than tired human doctors and talk to patients in a very kind, understanding way.
  • The Bad News: Because the AI is so eager to be helpful and "safe," it pushes for too many tests and medicines. If a patient talks to the AI first, they might walk into a real doctor's office demanding expensive tests that aren't needed. The real doctor then has to spend time explaining why those tests aren't necessary, which adds pressure to an already busy system.

In short: The AI is a brilliant student who knows the textbook perfectly and is very nice to talk to, but it lacks the real-world experience to know when to stop. It wants to do everything to be safe, which can actually create new problems for the healthcare system. The paper suggests we need rules to make sure AI uses its smarts without being overly pushy.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →