← Latest papers
📄 medicine

Evaluating the Usefulness, Quality, Reliability, and Readability of Large Language Model Responses About Pediatric Laser Dentistry: A Parent-Oriented Study

This study evaluates four paid large language models on parent-oriented pediatric laser dentistry questions, finding that while ChatGPT demonstrated the highest content quality and reliability, all models exceeded recommended readability levels and should serve only as educational supplements rather than replacements for professional dental consultation.

Original authors: Berkehan Aykanat, Emine Şuranur Ayaz

Published 2026-06-25
📖 4 min read☕ Coffee break read

Original authors: Berkehan Aykanat, Emine Şuranur Ayaz

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a parent trying to figure out if a new, high-tech "laser" tool is safe and good for your child's teeth. Instead of calling the dentist, you ask a group of four super-smart, digital "robots" (AI chatbots) for advice. This study is like a taste test where two expert dentists acted as judges to see which robot gave the best, most honest, and easiest-to-understand answers.

Here is the breakdown of what happened, using simple analogies:

The Contestants

The researchers set up a race between four advanced AI models (think of them as four different chefs):

  1. ChatGPT-5.5 Thinking
  2. Google Gemini 3.1 Pro
  3. Claude Opus 4.7
  4. DeepSeek-V4

They asked these robots 20 specific questions that a worried parent might ask about laser dentistry (e.g., "Is it painful?" "Is it safe for baby teeth?"). They asked the same questions every day for a week to see if the robots gave different answers over time.

The Judges

Two real pediatric dentists (specialists in children's teeth) read all 560 answers the robots gave. They didn't know which robot wrote which answer (they were "blindfolded" to the source). They scored the answers on three main things:

  • Usefulness: Was the answer actually helpful?
  • Quality: Was the information accurate and complete?
  • Readability: Was it written in plain English, or did it sound like a confusing science textbook?

The Results: Who Won?

🏆 The Star Performer: ChatGPT
ChatGPT was the clear winner. Its answers were like a wise, experienced teacher. It gave the most accurate, complete, and helpful information. It didn't just talk a lot; it said exactly what needed to be said to help a parent make a good decision. It scored highest on quality and usefulness.

🥈 The Middle Ground: Claude and Google Gemini
These two were like solid, reliable students. They did a good job, but they weren't quite as perfect as ChatGPT. Their answers were helpful and mostly accurate, but they sometimes missed a few small details or weren't quite as clear as the winner. They were very similar to each other in performance.

📉 The "Verbose" Underperformer: DeepSeek
DeepSeek was the most interesting case. Imagine a student who writes a 10-page essay to answer a simple question.

  • The Good: DeepSeek wrote the longest answers and used the simplest words. It was the easiest to read (like a friendly storybook).
  • The Bad: Despite being easy to read and very long, the expert dentists gave it the lowest scores for accuracy and reliability. It was like a very long, friendly story that got the facts wrong or left out important warnings. It was "fluffy" but not "substantial."

The "Day-to-Day" Check

The researchers asked the robots the same questions for seven days in a row.

  • The Finding: The robots were surprisingly consistent. They didn't change their personalities or their answers much from Monday to Sunday. They were stable, like a clock that always ticks at the same speed.

The Big Takeaway

The study found that while these AI robots are getting better, they are not perfect teachers yet.

  • ChatGPT was the best at giving accurate, high-quality advice for this specific topic.
  • DeepSeek proved that just because an answer is long and easy to read, it doesn't mean it's true or safe.
  • All of them wrote answers that were still a bit too complicated for the average parent to read easily (like a college textbook rather than a simple pamphlet).

The Final Verdict:
The paper concludes that these AI tools can be a helpful starting point for parents to learn about laser dentistry, but they should never replace a real conversation with a dentist. Just like you wouldn't let a GPS driver your car without looking out the window, you shouldn't let a chatbot make medical decisions for your child without a professional's check-up.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →