Thaísa Pinheiro Silva, Fernanda Bulhões Fagundes, Maria Fernanda Andrade-Bortoletto, Christiano Oliveira-Santos, Deborah Queiroz Freitas, Matheus L Oliveira
The evaluated VLMs showed limited agreement in estimating chronological age from panoramic radiographs. Therefore, their application in clinical and forensic contexts cannot currently be recommended.
OBJECTIVE: To explore the performance and limitations of three vision-language models (VLMs) in estimating chronological age from panoramic radiographs.
MATERIALS AND METHODS: A sample of 140 panoramic radiographs from individuals aged 7 to 20 years was uploaded to three VLMs: ChatGPT-4o, Microsoft Copilot Fast, and Microsoft Copilot Deep. Each VLM was independently prompted to estimate chronological age based on dental eruption patterns. The results were compared with the individuals' chronological ages. Statistical analysis included agreement within ± 1 year percentage calculation, and one-way analysis of variance to compare estimated and chronological ages (α = 0.05).
RESULTS: All three VLMs demonstrated low agreement in age estimation, with overall performance below 35%. Slightly better agreement was observed in individuals aged 9-10 years. The models showed greater deviations from chronological age at the extremes of the age range (7-8 and 17-20 years). No correct estimations were recorded for individuals aged 7, 17, 18, 19, and 20 years for ChatGPT-4o, 15 and 20 years for Microsoft Copilot Fast, and 20 years for Microsoft Copilot Deep.
CONCLUSIONS: The evaluated VLMs showed limited agreement in estimating chronological age from panoramic radiographs. Therefore, their application in clinical and forensic contexts cannot currently be recommended.
CLINICAL RELEVANCE: Understanding the potential applications and limitations of VLMs as an auxiliary tool in image interpretation is imperative to its use in clinical scenarios.