1.Accuracy of Orthodontic Malocclusion Detection Using Multiple AI Models: A Comparative Study
Hillda HERAWATI ; Joko KUSNOTO ; Indrayadi GUNARDI ; Anggit WIRASTO ; Tri Erri ASTOETI
Healthcare Informatics Research 2026;32(2):166-178
Objectives:
This study aimed to evaluate and compare the accuracy of multiple artificial intelligence (AI) models (ChatGPT 5.2 Pro, Gemini 3 Fast, Claude 4.5 Sonnet, and Microsoft Copilot) in detecting orthodontic malocclusion features in standardized multiview intraoral photographs. The reference standard was assessment by an orthodontist.
Methods:
A cross-sectional observational study was conducted using five standardized intraoral photographs (frontal, right lateral, left lateral, maxillary occlusal, and mandibular occlusal) obtained from 50 children aged 9–12 years. The following eight malocclusion parameters were assessed: anterior crowding, diastema, overjet, overbite, molar relationship, canine relationship, crossbite, and dental arch symmetry. Diagnostic accuracy and agreement between each AI model and the orthodontist were evaluated using Cohen’s kappa (κ) and the area under the receiver operating characteristic curve (AUC).
Results:
Agreement between the AI models and the orthodontist ranged from poor to moderate across all orthodontic domains, with Cohen’s κ values ranging from -0.15 to 0.63. Visually prominent alignment features, including anterior crowding and diastema, demonstrated comparatively higher agreement (κ, 0.00–0.63) and discriminatory performance, with AUC values ranging from 0.56 to 0.85. In contrast, parameters requiring precise spatial interpretation, such as sagittal relationships, overbite, crossbite, and arch morphology, showed consistently low agreement (κ, -0.15 to 0.38) and poor to near-random classification performance, with AUC values predominantly ranging from 0.41 to 0.70 and, in some cases, approaching 0.50.
Conclusions
Current multimodal AI models demonstrate limited, parameter-dependent accuracy in detecting orthodontic malocclusions from intraoral photographs. These findings emphasize the limitations of general-purpose AI systems for orthodontic decision support and highlight the need for task-specific models trained on clinically annotated datasets.

Result Analysis
Print
Save
E-mail