Can Multimodal Large Language Models Replace Expert Assessment in Preclinical Endodontic Education? A Cross-Sectional Agreement Study
This study demonstrates that current multimodal large language models systematically overestimate the quality of preclinical endodontic preparations and lack the reliability of expert endodontists, particularly regarding safety-critical judgments, indicating they are unsuitable as replacements for human assessment in dental education.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Technical Summary: Multimodal Large Language Models in Preclinical Endodontic Assessment
Problem Statement
The integration of Artificial Intelligence (AI) into dental education offers potential for automated feedback and scalable assessment. However, a critical gap remains in validating whether Multimodal Large Language Models (MLLMs) can reliably replace expert human judgment in high-stakes, safety-critical preclinical evaluations. Specifically, it is unclear if current MLLMs can accurately assess the nuanced psychomotor skills and safety judgments required in endodontic access cavity preparations, where errors can lead to patient harm. The study addresses the null hypothesis that no significant difference exists between expert endodontist assessments and AI model evaluations of these procedures.
Methodology
- Study Design: A single-center, cross-sectional, comparative agreement study conducted at Istanbul Aydın University during the 2025–2026 academic year.
- Participants and Specimens: 25 third-year undergraduate dental students performed access cavity preparations on standardized 3D-printed teeth. This resulted in 200 specimens across eight tooth-type groups (maxillary and mandibular central incisors, canines, premolars, and molars).
- Data Acquisition: Each specimen was evaluated using a standardized four-image dataset: occlusal, buccal, lingual/palatal photographs, and a periapical radiograph.
- Assessment Groups:
- Expert Group: Three experienced endodontists independently scored all specimens using a blinded, five-criterion analytic rubric (0–10 scale). The criteria were: (1) outline form, (2) location and centering, (3) extent and conservatism, (4) straight-line access, and (5) iatrogenic damage and safety. A penalty rule capped scores at 5/10 if critical errors (e.g., perforation) were detected.
- AI Groups: Four configurations of state-of-the-art MLLMs were tested: ChatGPT-5.2 (Standard and Reasoning Modes) and Gemini 3 (Standard and Reasoning Modes). Models were prompted to act as expert endodontists, process the four images, and output structured JSON scores.
- Statistical Analysis: Inter-rater reliability was assessed using Intraclass Correlation Coefficients (ICC). Group differences were analyzed using one-way ANOVA with Tukey post-hoc testing. Effect sizes were calculated as eta squared ().
Key Results
- Expert Reliability: Expert raters demonstrated excellent inter-rater reliability across all measurements (ICC range: 0.916–0.981), establishing a robust gold standard.
- AI vs. Expert Discrepancy: The null hypothesis was rejected. Significant differences () were found between expert and AI assessments across most measurements and tooth groups.
- Systematic Overscoring: All AI configurations consistently assigned higher scores than experts.
- Safety Criterion Failure: The most critical failure was in the "Iatrogenic Damage and Safety" criterion. While experts showed high reliability (ICC: 0.924) and variability reflecting clinical judgment, AI models clustered near maximum scores (≈1.87–2.00) with near-zero standard deviations. AI ICCs for this criterion ranged from 0.165 to 0.572, indicating poor to moderate agreement.
- Model Performance: Gemini 3 and Gemini 3/R showed the greatest divergence from expert scores (exceeding expert means by up to 4.5 points in some groups). ChatGPT-5.2 and ChatGPT-5.2/R showed closer alignment but still lacked expert-level precision.
- Reasoning Mode: The "Reasoning Mode" (chain-of-thought) did not consistently outperform "Standard Mode," suggesting that increased computational complexity did not translate to better clinical scoring.
- Tooth Type Variability: The Upper Premolar group was the only category where AI performance approached expert agreement on specific dimensional measurements (Outline Form, Extent, Straight-Line Access). Conversely, complex morphologies like Lower Canines and Lower Premolars yielded the largest score divergences.
- Internal Consistency: While AI models showed reasonable internal consistency across repeated runs (e.g., Gemini 3/R had high consistency), this consistency did not correlate with accuracy. A model could be internally consistent yet systematically diverge from expert judgment.
Key Contributions
- Empirical Validation of Limitations: The study provides empirical evidence that current MLLMs are not suitable as standalone replacements for expert assessment in preclinical endodontics, particularly for safety-critical judgments.
- Safety-Critical Failure Identification: It highlights a specific, dangerous limitation where AI models fail to discriminate between safe and unsafe preparations, defaulting to high safety scores regardless of actual iatrogenic damage.
- Distinction Between Consistency and Accuracy: The research demonstrates that high internal consistency within an AI model does not equate to accuracy or alignment with human expertise, cautioning against the use of consistency metrics as a proxy for validity in clinical education.
- Mode Comparison: It challenges the assumption that "Reasoning Mode" or chain-of-thought processing inherently improves performance in visually grounded clinical tasks.
Significance and Claims
The paper concludes that while AI models show promise for evaluating objective, dimensional parameters (such as outline form in simpler tooth types), they currently lack the reliability required for high-stakes competency evaluation, especially regarding patient safety. The authors argue that the failure to accurately assess "Iatrogenic Damage and Safety" renders these tools unsuitable for replacing expert judgment in preclinical education.
The authors propose a hybrid human-AI framework as the most pragmatic pathway forward. In this model, AI could support education by providing immediate, preliminary feedback on lower-stakes, quantifiable metrics, while expert human evaluation remains the essential standard for safety-critical judgments and complex clinical reasoning. The study emphasizes that before any AI tool is adopted for clinical competency evaluation, rigorous, criterion-level validation—specifically on safety metrics—is mandatory.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.