← Latest papers
📄 medicine

Performance of Large Language Models in Providing Patient-Oriented Information on Gingival Recession: A Comparative Study

This comparative study found that while ChatGPT, Google Gemini, and Claude provide similarly accurate and comprehensive patient-oriented information on gingival recession, expert evaluators preferred Gemini and Claude despite ChatGPT's faster response times and greater conciseness.

Original authors: Sultan Albeshri

Published 2026-08-25
📖 4 min read☕ Coffee break read

Original authors: Sultan Albeshri

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Patients increasingly turn to artificial intelligence to understand their health, asking chatbots for advice on everything from common colds to complex dental conditions. These tools, known as large language models, are designed to converse in natural language, offering direct answers rather than a list of search links. One such condition, gingival recession, occurs when the gum tissue that surrounds the teeth pulls back, exposing the tooth root. This exposure can lead to sensitivity, decay, and aesthetic concerns, prompting many people to seek information online before visiting a dentist. While these digital assistants promise to bridge the gap between patients and medical knowledge, questions remain about whether they provide reliable, complete, and safe information, especially for specific dental issues where precision matters.

A recent study set out to test how well three leading artificial intelligence systems handle questions about gum recession. The researchers focused on ChatGPT, Google Gemini, and Claude, asking each of them the same twenty questions that a typical patient might ask. These inquiries covered the causes of the condition, treatment options, risks, and what to expect after surgery. To ensure the answers were judged fairly, five board-certified specialists in gum health reviewed every response without knowing which computer program had generated it. The experts evaluated the answers based on how factually correct they were, how complete the information was, and how well the models ranked against each other in terms of overall quality. They also measured how long it took each system to write its answer and counted the number of words used.

The results showed that all three artificial intelligence models performed remarkably well when it came to the core facts. There was no significant difference in the accuracy or completeness of the information provided by ChatGPT, Gemini, or Claude. Each system managed to generate responses that experts found to be clinically acceptable, with most answers scoring high marks for correctness. However, the models differed noticeably in how they delivered that information. ChatGPT was the fastest, producing its answers in just under six seconds on average, while Claude took nearly fifteen seconds. ChatGPT also wrote the most concise responses, averaging about 232 words per answer. In contrast, Gemini produced the longest responses, averaging 438 words, and Claude fell somewhere in between.

Despite the similar quality of the facts, the human experts had a clear preference for how the information was presented. When asked to rank the answers from best to worst, Gemini and Claude were chosen as the top response more often than ChatGPT. The experts selected Gemini and Claude as the best answer in 38 percent of the cases, whereas ChatGPT was ranked first in only 24 percent of the evaluations. This suggests that while ChatGPT was quicker and more brief, the other two models provided a style of explanation that the specialists found more thorough or better organized. The study also found that all three models were particularly good at answering questions about postoperative care and follow-up instructions, likely because these topics rely on standardized guidelines. They were slightly less consistent when discussing the specific diagnosis or candidacy for treatment, which requires more personalized judgment.

Ultimately, the study concludes that these artificial intelligence tools are becoming reliable enough to serve as helpful supplements for patient education. They can provide accurate and detailed information about gum recession that aligns with professional standards. However, the researchers emphasize that these tools should support, not replace, a consultation with a dental professional. The best approach appears to be using these fast and knowledgeable systems to get a general understanding of a condition, while relying on a human expert to interpret that information in the context of an individual's specific health needs.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →