科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Diagnostics (Basel, Switzerland)2026-08-26

Cross-Domain Generalization of Deep Learning Architectures for Cephalometric Landmark Detection: A Dual-Dataset and Multi-Device Benchmark.

Mustafa Özcan, Ferdi Allaf

原始摘要(英文原文)· Original abstract
Background/Objectives: Deep learning models for cephalometric landmark detection report near-ceiling accuracy on single benchmarks, yet most are trained and tested on the same dataset. Whether the best in-domain architecture remains best out-of-domain has not been systematically quantified. Methods: Four architecture families (heatmap CNN, two-stage cascade, coordinate-regression Vision Transformer, pretrained ResNet-50) were each trained on two independently sourced datasets-ISBI 2015 (400 images, one device) and Aariz (1000 images, seven devices)-and evaluated on both, over their 19 shared landmarks in millimeters. All axes used three seeds (mean ± SD): the cross-dataset matrix, leave-one-device-out shift, balanced joint training, a pretrained-versus-scratch ablation, and a landmark-level breakdown. Results: In-domain mean radial error (MRE) was 2.47-3.04 mm (Aariz) and 4.12-5.55 mm (ISBI); cross-dataset error rose steeply (Generalization Drop-the relative increase in error out-of-domain-155-426%). The most accurate model in-domain (a from-scratch heatmap CNN, 2.47 mm) showed the largest drop, and every pretrained estimate fell below every from-scratch estimate (family means 248% vs. 380%): in-domain ranking did not predict cross-domain ranking. Leave-one-device-out revealed reproducible device-specific shift (held-out MRE 1.89-7.94 mm). Inter-observer variability was 0.53 mm, so cross-domain errors were 15-44× the human band. Balanced joint training reduced the cross-domain gap for all four architectures (both domains ≈ 2.1-3.5 mm) without harming in-domain accuracy. Pretraining more than halved cross-domain error from Aariz to ISBI (7.73 vs. 16.36 mm) but not in reverse, supporting the mechanism in one direction rather than uniformly. A-point, B-point, Nasion, and Menton all exceeded the 2 mm clinical threshold out-of-domain, including points localized to sub-millimeter accuracy in-domain. Conclusions: Single-dataset accuracy substantially overstates clinical generalizability, and the best in-domain architecture is not the most transferable, so in-domain leaderboards are an unreliable basis for selecting a model for clinical deployment; balanced multi-source training recovers most of the loss across the architectures tested.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Cross-Domain Generalization of Deep Learning Architectures for Cephalometric Landmark Detection: A Dual-Dataset and Multi-Device Benchmark. — 科研速览 Science Skim