Artificial Intelligence in Caries Risk Assessment: Evaluating the Current Status of Caries Management by Risk Assessment and Cariogram with Large Language Models


Çakıcı Ş., Akkoç S.

Acta Cytologica, ss.1-15, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1159/000553200
  • Dergi Adı: Acta Cytologica
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, CINAHL, EMBASE, MEDLINE
  • Sayfa Sayıları: ss.1-15
  • Anahtar Kelimeler: Artificial intelligence, Caries risk assessment, ChatGPT, Early childhood, Google Gemini
  • Kütahya Sağlık Bilimleri Üniversitesi Adresli: Evet

Özet

Abstract – Introduction: Caries risk assessment (CRA) is essential for individualized caries management in early childhood, but existing tools are difficult to standardize in routine practice. The ability of large language model (LLM)-based artificial intelligence (AI) systems to recognize and correctly apply validated CRA tools remains unclear. This study aimed to evaluate the ability of ChatGPT and Google Gemini, LLM-based AI systems, to apply and interpret CRA tools in early childhood under different prompting conditions. Methods: Thirty standardized clinical vignettes of children aged 0–5 years were evaluated using Caries Management by Risk Assessment (CAMBRA) and the Cariogram by ChatGPT-5.1 (OpenAI, San Francisco, CA, USA) and Google Gemini-3.0 Pro (Google LLC, Mountain View, CA, USA) under three prompting conditions: unguided, guideline-informed, and guideline-only. An expert panel provided reference classifications. AI outputs were assessed for categorical agreement, mean absolute error (MAE), quality, accuracy, and readability. MAE differences were analyzed using the Friedman test and mixed-design repeated-measures analysis of variance (ANOVA), while quality, accuracy, and readability were analyzed using repeated-measures ANOVA. Reliability was assessed using intraclass correlation coefficients (ICCs). Results: Inter- and intra-rater reliability were high (ICC = 0.89–0.93). Reference CAMBRA and Cariogram classifications showed a moderate correlation (ρ = 0.509; n = 30). MAE differed significantly across AI model-prompting condition combinations for both tools (p < 0.001). No significant difference was observed between AI models for CAMBRA, whereas ChatGPT showed lower MAE than Gemini for the Cariogram (p = 0.002). Guideline-informed and guideline-only prompting significantly reduced MAE and improved quality and accuracy compared with the unguided condition (p < 0.001). Conclusion: LLM-based AI systems can support CRA, particularly when guideline-based prompting is used. However, performance depends on the CRA tool and AI model, and numerical or algorithmic components, such as those in the Cariogram, remain challenging. To prevent tooth decay in children, dentists first need to estimate how likely a child is to develop cavities. This process is called caries risk assessment (CRA), which means evaluating a child’s risk of getting tooth decay based on their habits, oral health, and clinical findings. Tools such as CAMBRA and the Cariogram are commonly used to support this decision. In recent years, artificial intelligence (AI) systems that can generate text, called large language models, have become widely available. These systems can answer medical and dental questions, but it is still unclear how well they can apply structured clinical guidelines such as those used in CRA. In this study, we examined whether two AI systems, ChatGPT and Google Gemini, could correctly assess caries risk in children using hypothetical clinical cases. We tested the models under three different conditions: without guidance, with partial guidance, and with detailed guideline-based instructions. Their answers were compared with an expert-developed reference framework. We found that AI performed better when clear guideline-based instructions were provided. However, the results also showed that some parts of CRA are still difficult for these systems, especially when numerical or rule-based decision steps are required, such as those used in the Cariogram. Our findings suggest that AI may be helpful as a supportive tool in CRA, but it cannot replace professional clinical judgment. Careful use and clear guidance are essential if such systems are to be used safely in pediatric dental practice.