Performance of Advanced Large Language Models in Caries Risk Assessment and Preventive Decision-Making: A Multidimensional Evaluation of Five Chatbots.
Caries research, ss.1-17, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası:
- Basım Tarihi: 2026
- Doi Numarası: 10.1159/000553557
- Dergi Adı: Caries research
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, BIOSIS, CINAHL, EMBASE, MEDLINE, Academic Search Ultimate (EBSCO), Health Research Premium Collection (ProQuest), Pharma Collection (ProQuest)
- Sayfa Sayıları: ss.1-17
- Çanakkale Onsekiz Mart Üniversitesi Adresli: Evet
Özet
Introduction: Large language models (LLMs) have recently been integrated into dental practice to support clinical reasoning and preventive decision-making. This study compared the performance of five advanced chatbots-ChatGPT-5, Claude 4.5 Sonnet, Gemini 2.5 Pro, LLaMA 3.1, and Mistral 7B-in providing evidence-based responses for caries risk assessment and preventive management in pediatric cases.
Methods: Twenty-five validated, case-based questions were developed in accordance with internationally recognized pediatric and preventive dentistry guidelines. Responses were evaluated by six pediatric dentistry experts for accuracy, completeness, relevance, clarity, and usefulness using Likert-type scales. Response time, word count, and linguistic readability characteristics (Flesch Reading Ease Score and Flesch-Kincaid Grade Level) were additionally analyzed to compare textual complexity across chatbot-generated responses. Data normality was assessed using the Shapiro-Wilk test; parametric tests (ANOVA with Bonferroni correction) or non-parametric tests (Kruskal-Wallis with Dunn's post hoc) were applied as appropriate.
Results: Statistically significant differences were observed across all qualitative criteria, including accuracy, completeness, relevance, clarity, and usefulness (p < 0.001). ChatGPT-5 consistently ranked among the top-performing models, showing balanced and high-quality responses across domains, while Claude 4.5 Sonnet achieved the highest accuracy and completeness scores. Gemini 2.5 Pro produced the fastest responses (p < 0.001), whereas Claude 4.5 Sonnet generated the longest and most linguistically complex outputs. Readability metrics also differed significantly among models (p < 0.001), with Mistral 7B and LLaMA 3.1 showing the highest readability.
Conclusions: All evaluated chatbots generated generally relevant responses for caries risk assessment and preventive counseling; however, substantial inter-model differences were observed in qualitative performance, linguistic complexity, and response characteristics. Occasional inconsistencies and outdated content highlight the need for cautious interpretation and further externally validated evaluation before broader clinical implementation.