Accuracy of Orthodontic Malocclusion Detection Using Multiple AI Models: A Comparative Study

Open

Hillda Herawati, Joko Kusnoto, Indrayadi Gunardi, Anggit Wirasto, Tri Erri Astoeti

2026 Healthcare Informatics Research Vol. 32 Issue 2 Article Cited by 0 Quartile

Abstract

Objectives: This study aimed to evaluate and compare the accuracy of multiple artificial intelligence (AI) models (ChatGPT 5.2 Pro, Gemini 3 Fast, Claude 4.5 Sonnet, and Microsoft Copilot) in detecting orthodontic malocclusion features in standardized multiview intraoral photographs. The reference standard was assessment by an orthodontist. Methods: A cross-sectional observational study was conducted using five standardized intraoral photographs (frontal, right lateral, left lateral, maxillary occlusal, and mandibular occlusal) obtained from 50 children aged 9–12 years. The following eight malocclusion parameters were assessed: anterior crowding, diastema, overjet, overbite, molar relationship, canine relationship, crossbite, and dental arch symmetry. Diagnostic accuracy and agreement between each AI model and the orthodontist were evaluated using Cohen’s kappa (κ) and the area under the receiver operating characteristic curve (AUC). Results: Agreement between the AI models and the orthodontist ranged from poor to moderate across all orthodontic domains, with Cohen’s κ values ranging from-0.15 to 0.63. Visually prominent alignment features, including anterior crowding and diastema, demonstrated comparatively higher agreement (κ, 0.00–0.63) and discriminatory performance, with AUC values ranging from 0.56 to 0.85. In contrast, parameters requiring precise spatial interpretation, such as sagittal relationships, overbite, crossbite, and arch morphology, showed consistently low agreement (κ,-0.15 to 0.38) and poor to near-random classification performance, with AUC values predominantly ranging from 0.41 to 0.70 and, in some cases, approaching 0.50. Conclusions: Current multimodal AI models demonstrate limited, parameter-dependent accuracy in detecting orthodontic malocclusions from intraoral photographs. These findings emphasize the limitations of general-purpose AI systems for orthodontic decision support and highlight the need for task-specific models trained on clinically annotated datasets. © 2026 The Korean Society of Medical Informatics.

Affiliations

Doctoral Program in Dental Sciences, Faculty of Dentistry, Universitas Trisakti, Jakarta, Indonesia; Department of Orthodontics, Faculty of Dentistry, Universitas Trisakti, Jakarta, Indonesia; Department of Oral Medicine, Faculty of Dentistry, Universitas Trisakti, Jakarta, Indonesia; Informatics Study Program, Faculty of Science and Technology, Universitas Harapan Bangsa, Purwokerto, Indonesia; Department of Dental Public Health and Preventive Dentistry, Faculty of Dentistry, Universitas Trisakti, Jakarta, Indonesia

Research at a Glance

Premium content — register to unlock

Research at a Glance

Register to unlock

Topics & SDG Alignment

Premium content — register to unlock

Topics & SDG Alignment

Register to unlock

Collaboration

Premium content — register to unlock

Collaboration

Register to unlock

Author Profile (Selected)

Premium content — register to unlock

Author Profile (Selected)

Register to unlock

References Overview

Premium content — register to unlock

References Overview

Register to unlock

Journal & Source

Premium content — register to unlock

Journal & Source

Register to unlock

Metadata & Integrity

Premium content — register to unlock

Metadata & Integrity

Register to unlock