← Latest papers
📄 medicine

Generative AI Responses to Patient-Presented Orthodontic Misinformation and Controversial Claims: A Comparative Cross-Sectional Evaluation of Safety, Accuracy, Evidence Concordance, Misinformation Correction, Uncertainty Communication, Transparency, and Actionability

This comparative cross-sectional study evaluates five generative AI chatbots on patient-facing orthodontic queries, finding that while most avoid overtly unsafe advice, they exhibit significant differences in accuracy, evidence alignment, and corrective capabilities, underscoring the need to treat them as adjunctive tools rather than substitutes for professional assessment.

Original authors: Xiaoli Liu, Yan Zhuang, Yulong Wang, Zhaole Gong, Wu Dong

Published 2026-09-15
📖 5 min read🧠 Deep dive

Original authors: Xiaoli Liu, Yan Zhuang, Yulong Wang, Zhaole Gong, Wu Dong

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

People increasingly turn to search engines, social media, and short videos for answers about their health. These channels make information easy to find, but they also make it hard to judge who is telling the truth. This is especially true for orthodontics, the field of dentistry that straightens teeth and aligns jaws. Advice found online often mixes facts with exaggerations, personal stories, or sales pitches. A patient might read that a specific type of brace can change their face shape, or that they can straighten their own teeth at home with rubber bands. While some of this information is harmless, other claims can lead to unrealistic expectations, delayed medical care, or even physical injury. When a person encounters a confusing or dangerous claim, they need a clear, safe, and accurate explanation to help them decide what to do next.

In recent years, a new kind of tool has emerged to help people organize information: generative artificial intelligence. These systems can hold conversations and answer questions in fluent, natural language. They are becoming common in daily life, and many people now ask them for health advice. However, it remains unclear whether these systems can handle the tricky task of correcting false medical claims without accidentally giving dangerous instructions. A computer program might sound confident and polite while omitting a critical warning or failing to explain why a popular idea is wrong. To understand how well these tools perform in a real-world scenario, researchers needed to test them against the specific kinds of misleading questions that patients actually ask.

A team of researchers from hospitals in China set out to test five of the most popular consumer-facing artificial intelligence chatbots. They wanted to see how these systems responded when users presented them with false or controversial claims about orthodontic treatment. The study focused on five specific systems: ChatGPT, Gemini, Microsoft Copilot, DeepSeek, and Doubao. The researchers did not ask these systems simple questions like "what is a tooth?" Instead, they created sixty-eight specific prompts based on real-world misinformation. These prompts included false ideas about extracting teeth, changing facial features, using do-it-yourself aligners, or accelerating treatment with unproven methods. Each prompt was designed to contain a false premise that the chatbot would need to recognize and correct.

The researchers submitted each of the sixty-eight questions to all five chatbots, generating a total of three hundred and forty responses. They treated each question as a single, one-time conversation to ensure the results were consistent. To evaluate the answers, five experienced orthodontic doctors reviewed the text independently. They looked for several key qualities. First, they checked for safety: did the response encourage any behavior that could harm the patient, such as trying to move teeth without a professional exam? Second, they measured accuracy: was the information factually correct? Third, they assessed how well the answer matched established medical evidence. They also checked whether the chatbot actively corrected the false premise in the question, rather than just ignoring it and giving a general answer. Finally, they evaluated whether the response communicated uncertainty appropriately, whether it was transparent about its sources, and whether it gave the patient a clear, safe next step to take.

The results showed that all five chatbots were generally good at avoiding the most obvious dangers. Out of the three hundred and forty responses, three hundred and seven were classified as safe, meaning they did not contain advice that could lead to immediate physical harm. However, the systems were not equally good at everything else. The researchers found significant differences in how well the chatbots handled the quality of their answers. ChatGPT consistently provided the most accurate information and was best at correcting the false ideas presented in the questions. It also did the best job of explaining the limits of its knowledge and offering clear, actionable steps for the patient. Gemini performed well in some areas, tying with ChatGPT on how closely its answers matched medical evidence, but it was less consistent in other categories. The other three systems, including Copilot, DeepSeek, and Doubao, generally scored lower across the board, often failing to correct the misinformation or provide sufficiently detailed guidance.

A crucial finding of the study was that being "safe" does not mean an answer is good enough to use on its own. Many of the responses avoided dangerous advice but still failed to correct the patient's misunderstanding or provide the necessary context. For example, a chatbot might correctly state that teeth need braces but fail to explain that a specific do-it-yourself method mentioned in the question is unsafe. The study showed that while the systems rarely gave instructions that would cause immediate injury, they often lacked the depth and precision required for medical guidance. The researchers noted that the differences between the systems were not just minor variations; the top-performing models were clearly superior in their ability to handle complex, misleading claims compared to the others.

The authors concluded that these artificial intelligence tools should be viewed as helpful supplements to patient education, not as replacements for a visit to a dentist. They can help organize information and answer general questions, but they cannot replace the need for a professional examination, X-rays, or a personalized treatment plan. The study highlights that patients should not rely on these chatbots to validate claims they have seen online, especially when those claims involve changing their teeth or face. Instead, the best use of these tools is to gather initial information and then seek a qualified professional to verify the facts and discuss a safe course of action. The research suggests that while the technology is advancing, it still requires human oversight to ensure that the advice given is not only safe but also accurate, complete, and truly helpful for the individual patient.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →