Performance of large language models in fluoride-related dental knowledge: a comparative evaluation study of ChatGPT-4, Claude 3.5 Sonnet, Copilot, and Grok 3
- Author:
Raju BISWAS
1
;
Atanu MUKHOPADHYAY
;
Santanu MUKHOPADHYAY
Author Information
- Publication Type:Original article
- From: Journal of Yeungnam Medical Science 2025;42(1):53-
- CountryRepublic of Korea
-
Abstract:
Background:Large language models (LLMs) are increasingly used in medical and dental education to enhance clinical reasoning, patient communication, and academic learning. This study evaluates the effectiveness of four advanced LLMs— ChatGPT-4 (OpenAI), Claude 3.5 Sonnet (Anthropic), Microsoft Copilot, and Grok 3 (xAI)—in conveying fluoride-related dental knowledge.
Methods:A cross-sectional comparative study was conducted using a mixed-methods approach. Each LLM answered 50 multiple- choice questions (MCQs) and 10 open-ended questions on fluoride chemistry, clinical applications, and safety concerns. Two blinded experts rated the open-ended responses on accuracy, depth, clarity, and evidence. Interrater reliability was assessed using Cohen’s kappa and Spearman’s correlation, and statistical analyses were performed using analysis of variance, Kruskal-Wallis, and post-hoc tests.
Results:All models showed high MCQ accuracy (88%–94%). Claude 3.5 Sonnet achieved the highest scores in open-ended responses, especially for clarity (p=0.009). Minor differences in accuracy, depth, and evidence were not statistically significant. Overall, all LLMs performed strongly, with high interrater agreement supporting result reliability.
Conclusion:Advanced LLMs show strong potential as supportive tools in dental education and patient communication on fluoride use. Claude 3.5 Sonnet demonstrated superior linguistic clarity, enhancing its educational value. Continued evaluation and clinical oversight are crucial for their safe and effective integration into dentistry.
