Classification Fidelity of Large Language Models on Endodontic Items from the Turkish Dental Specialty Examination
This study evaluates the classification fidelity of four large language models in categorizing endodontic examination questions across three educational taxonomies, revealing that while models like Gemini and DeepSeek achieve substantial agreement for Bloom's and SOLO taxonomies, all models struggle with Miller's Pyramid, demonstrating that LLM performance is taxonomy-dependent and distinct from answer accuracy.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a giant library of dental exam questions (specifically about root canals) from Turkey, written between 2012 and 2025. You want to know if four famous AI chatbots (ChatGPT, Gemini, Kimi, and DeepSeek) can do more than just answer these questions correctly. You want to know if they can understand the "difficulty level" of the question itself.
Think of it like a music teacher. A student might play a song perfectly (getting the answer right), but can they tell you if the song is a simple nursery rhyme, a complex jazz piece, or a full symphony? That's what this study tested.
Here is the breakdown of the study using simple analogies:
1. The Three "Difficulty Rulers"
The researchers used three different "rulers" to measure how hard a question is. The AI had to sort every question into the right category on each ruler:
- Bloom's Taxonomy (The "Thinking Ladder"): This measures how much brainpower is needed.
- Bottom rungs: Just remembering facts (like "What is a tooth?").
- Top rungs: Complex thinking like analyzing a problem or creating a solution.
- SOLO Taxonomy (The "Building Block" Tower): This measures how many pieces of information you need to connect to answer.
- Small tower: You only need one fact.
- Tall tower: You need to connect many facts and see how they relate.
- Miller's Pyramid (The "Doctor's Ladder"): This measures if the knowledge is just in your head or if you can actually do it.
- Bottom: "I know the theory."
- Top: "I can do it in real life on a patient."
2. The Human "Gold Standard"
Before asking the AI, the researchers hired human experts (specialist dentists and education professors) to sort the questions first. They did this very carefully, training the experts, auditing their work, and having them debate until they agreed. This created the "Answer Key" for what the questions actually were.
3. The AI Showdown
The researchers fed the same 200 questions to the four AI models and asked them to sort the questions using the three rulers above. They then compared the AI's sorting to the Human "Answer Key."
The Results: Who Got It Right?
The Good News (Bloom & SOLO):
- Gemini and DeepSeek were the star students. They were very good at sorting questions by how "thinking-heavy" they were (Bloom) and how many "building blocks" were needed (SOLO). They agreed with the human experts most of the time.
- ChatGPT and Kimi were okay, but they made more mistakes. They often thought simple questions were too hard, or complex questions were too easy.
The Bad News (Miller's Pyramid):
- Everyone struggled here. No AI did a great job sorting questions by "Can this be done in real life?" (Miller).
- The Twist: The rankings flipped completely!
- Gemini and DeepSeek were the best at the "Thinking" rulers but the worst at the "Real Life" ruler.
- ChatGPT was the worst at the "Thinking" rulers but the best (though still only "fair") at the "Real Life" ruler.
Key Takeaways from the Paper
- Answering vs. Understanding are different skills: Just because an AI can get the right answer to a dental question doesn't mean it understands why the question is hard or what kind of thinking it requires. It's like a calculator that can solve a math problem but doesn't understand the concept of algebra.
- One size does not fit all: You can't use one AI for every job. If you want to check if questions are getting harder (Bloom/SOLO), use Gemini or DeepSeek. If you are trying to guess if a question relates to real-life practice (Miller), ChatGPT might actually be your best bet, even though it's worse at the other tasks.
- The questions are getting harder: The study found that the dental exam questions from 2025 are significantly more complex and require more "thinking" than the questions from 2012.
- AI gets confused by the top of the ladder: All the AIs found it much harder to correctly identify the very hardest questions (the ones requiring analysis or evaluation) compared to the easy ones (just remembering facts).
What the Paper Does Not Say
The paper does not say that these AIs should be used to grade students, replace teachers, or diagnose patients. It strictly says that while AIs are getting better at sorting questions by difficulty, they are not perfect, and their performance changes depending on which difficulty ruler you use. It also notes that because the exam questions are getting harder over time, older studies on AI might have overestimated how smart the AI really is.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.