Large Language Models as Decision-Support Tools for Laser Dentistry: A Blinded, Expert-Rated Comparison
In a blinded, expert-rated comparison of five AI platforms using 24 synthetic laser-dentistry vignettes, a medical-domain retrieval-augmented generation (RAG) system and a leading general-purpose large language model significantly outperformed other tools in accuracy, safety, and completeness, while a general-purpose RAG platform performed the worst.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Dentists have long used lasers as precise tools to treat gum disease, clean root canals, and whiten teeth. Unlike a drill that grinds away, a laser beam can target specific tissues, such as the soft pink gums or the hard white enamel, by tuning its light to interact with water or blood. However, choosing the right laser is a delicate task. If the light is too strong or the wrong color, it can burn the tooth's nerve or damage the eye. If it is too weak, the treatment fails. Because there is no single, universal rulebook for every situation, dentists must carefully calculate the power, the speed of the pulses, and the cooling methods for each patient. This complexity has led some to wonder if artificial intelligence could act as a guide, offering instant advice on how to set these machines safely and effectively.
To test this idea, researchers at the Technion and New Bulgarian University set up a controlled experiment involving five different artificial intelligence platforms. They did not ask the computers to diagnose real patients, but rather to solve twenty-four carefully crafted, fictional dental cases. These scenarios covered three main areas of laser dentistry: gum and surgical work, root canal treatments, and restorative procedures like crowns or fillings. For each case, the team asked the same six-part question: Is a laser the right tool? Which type of laser should be used? What are the exact settings for power and timing? What is the step-by-step plan? What safety measures are needed? And what is the expected result? The five platforms tested included three general-purpose chatbots that are widely known, one system designed specifically to search medical journals, and another that searches the general internet for answers.
The results of this comparison were clear and revealed significant differences in performance. The researchers enlisted three expert dentists, each specializing in one of the three areas, to grade the answers without knowing which computer generated them. The experts scored the responses on a scale of one to five, looking for accuracy, safety, and completeness. The findings showed that not all artificial intelligence is created equal when it comes to medical advice. The system that searched through peer-reviewed medical literature, called OpenEvidence, and a top-tier general chatbot named Claude, provided the best guidance. They consistently offered the most accurate settings and the safest protocols. In contrast, the system that searched the open internet, Perplexity, performed the worst. It produced the highest number of answers that contained dangerous errors or missing critical steps. In fact, more than half of the responses from the internet-searching tool contained at least one element that an expert would consider poor or potentially harmful, whereas the medical literature system never produced a single unsafe answer.
The study also highlighted that the quality of the advice depended heavily on the domain of dentistry. The differences between the best and worst platforms were most dramatic in root canal and restorative cases, where the settings must be precise to avoid burning the tooth's nerve. In gum disease cases, where there is often more flexibility in how a treatment is performed, the gap between the platforms was smaller. Interestingly, the length of the answer did not guarantee quality. One of the best systems, Claude, gave concise, direct answers that were highly rated, while another platform gave very long responses that were often filled with errors. The researchers found that the system relying on medical journals was able to ground its answers in established science, avoiding the "hallucinations" or made-up facts that can plague other systems. While the general chatbots were often good at sounding confident, they were more likely to suggest incorrect laser settings or miss safety warnings.
This experiment suggests that while artificial intelligence can be a powerful tool for dentists, it cannot yet be trusted to work alone. The study demonstrated that a system built on verified medical knowledge outperformed those scanning the open web, but even the best systems made occasional mistakes. The researchers concluded that these tools should serve as assistants to help specialists make decisions, not as replacements for them. Before a dentist uses a laser based on a computer's suggestion, a human expert must verify the settings against current medical evidence. The study serves as a reminder that in the high-stakes world of dentistry, where a wrong setting can cause real physical harm, the source of the information matters more than the speed at which it is delivered.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.