Manami Yamaguchi, Masato Tsutsumi, Yasuhiro Kuroda, Yoshiki Soeda, Tetsutaro Yamaguchi
Background/Objectives: Accurate landmark identification underpins reliable cephalometric analysis. This study evaluated a ResNet50-based, single-stage regression convolutional neural network for direct automatic localization of 15 landmarks on lateral cephalograms. Methods: This retrospective study included 669 lateral cephalograms (669 patients): 619 for training and 50 randomly selected for an internal holdout test set. One of five orthodontists annotated each cephalogram, and coordinates served as the reference standard. One orthodontist independently re-annotated all test images after more than 2 weeks to assess reproducibility. Model performance was evaluated using Euclidean localization errors and success detection rates (SDR). Results: Across 750 landmark predictions, mean localization error was 1.25 ± 1.39 mm (95% confidence interval, 1.15-1.35 mm) and median error was 0.72 mm. SDRs within 1.0, 2.0, and 4.0 mm were 64.5%, 81.5%, and 94.4%, respectively. A statistically significant overall difference was observed among the 15 landmarks (Friedman χ2(14) = 28.84, p = 0.011). Point A had the numerically largest mean error (1.66 mm) and the mandibular central incisor the smallest (1.04 mm). The mean difference between original and repeated annotations was 0.52 ± 0.92 mm. Conclusions: In this single-center internal holdout test set, mean localization error was below the 2.0 mm benchmark. However, 18.5% of predictions exceeded 2.0 mm, and external validity remains unconfirmed. Artificial-intelligence-generated landmarks should be verified by orthodontists, and external validation is required before broader clinical use.