← Latest papers
📄 medicine

Evaluating Five Generative AI Chatbots for Myocarditis-Related Public Health Consultation: A Multidimensional Assessment of Safety, Accuracy, Empathy, Reliability, Quality, and Readability

This study evaluated five generative AI chatbots on myocarditis-related public health queries, finding that while most responses were generally safe, significant variations existed in accuracy, empathy, and quality, with all models failing to meet readability targets and exhibiting specific risks in triage and treatment guidance that necessitate clinician oversight.

Original authors: Youyou Chen¹, Hui Ma, Huimin Wang, Duoxue Chen, Rongyan Jiang, Haiyan Wang, Dai Yu

Published 2026-08-07
📖 4 min read☕ Coffee break read

Original authors: Youyou Chen¹, Hui Ma, Huimin Wang, Duoxue Chen, Rongyan Jiang, Haiyan Wang, Dai Yu

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a super-smart, all-knowing robot friend who can answer any question you ask, from "Why is the sky blue?" to "How do I fix my bike?" This robot is powered by something called Generative AI, a type of computer brain that learns from a massive library of human writing to create new, fluent answers. In the world of health, people are starting to ask these robots questions like, "My chest hurts after a vaccine; is it serious?" or "Can I go back to playing soccer?" But here's the catch: while these robots are great at sounding confident and friendly, they aren't doctors. They don't have a heartbeat, and they can sometimes make up facts or miss tiny but dangerous details. This is especially true for a condition called myocarditis, which is a fancy word for inflammation of the heart muscle. It's like a fire burning inside the heart's engine; sometimes it's a small flicker, and other times it's a raging inferno that can stop the heart from working. Because the stakes are so high, we need to know if our robot friends are safe to talk to when we're worried about our hearts, or if they might accidentally tell us to ignore a warning sign.

A team of researchers decided to put five of these popular AI chatbots to the test, treating them like students taking a very serious exam. They didn't just ask random questions; they gathered 52 real-life concerns that people actually have about myocarditis, such as worries about chest pain, what happens after an infection, whether vaccines are safe, and when it's okay to exercise again. They asked each of the five chatbots (ChatGPT, Gemini, DeepSeek, Doubao, and Copilot) to answer these questions exactly as a user would, without any special instructions to make them behave better. Then, a group of five expert heart doctors, who didn't know which robot gave which answer, graded the responses on a scale of "safe" to "unsafe," "accurate" to "wrong," "empathetic" to "cold," and "easy to read" to "confusing."

The results were a mix of good news and some serious warnings. First, the good news: most of the time, the robots were safe. About 90% to 94% of their answers didn't contain dangerous advice. However, the researchers found that "mostly safe" isn't the same as "perfect." Even the best robots made mistakes in tricky situations. For example, some robots were too quick to reassure people that a vaccine-related heart issue was harmless, or they gave vague advice about when to stop exercising, which could be risky if a person's heart is still inflamed. It's like a robot telling you, "Your car engine is probably fine," when it actually needs to be towed immediately.

When it came to accuracy and kindness, the robots were not all the same. Some, like Gemini, Copilot, and DeepSeek, were very good at getting the medical facts right. DeepSeek was the "kindest" of the bunch, sounding more supportive and understanding of a patient's fear. But here is the biggest problem: none of the robots wrote in a way that was easy for a regular person to understand. They all used words that were too complex, like a professor explaining a simple concept using a dictionary full of big words. The researchers wanted the answers to be readable by a sixth-grader, but every single robot failed this test. Their answers were more like a college textbook than a friendly chat.

The study concludes that while these AI tools can be helpful for learning general facts about heart health, they are not ready to be used as a substitute for a real doctor. You shouldn't ask them to decide if you have a heart attack, to tell you if you can stop taking medicine, or to give you the green light to run a marathon. The robots are like a very well-read library assistant who can point you to the right book, but they can't drive the car for you. If you use them, you need a real human doctor to check their answers, translate the big words into plain English, and make sure you don't miss any red flags. The robots are getting better, but for something as delicate as your heart, they still need a human supervisor.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →