Hossein Mohammad‐Rahimi, Rishi Sanjay Ramani, Frank Setzer, Falk Schwendicke, Ruben Pauwels, Ali Nosrat
Artificial intelligence (AI) is emerging as a powerful tool in dentistry, where endodontics can benefit significantly from its potential applications. From enhanced diagnostic accuracy to treatment planning and decision-making, optimisation, and outcome prediction, AI may significantly improve the clinical practice and the teaching of endodontics (Mohammad-Rahimi et al. 2024; Ourang et al. 2024). However, this still remains a largely unfulfilled promise, because many studies suffer from significant limitations or methodological flaws. Several factors contribute to this condition. First, as often seen with the advent of novel technology, there is a lack of scientific expertise in AI methodology by editors and reviewers, which may result in the eventual publication of studies that display limitations or, sometimes, even fundamental flaws in design, implementation, or validation. Conversely, technical publications may present sophisticated algorithms that solve problems with limited clinical utility or fail to address the nuanced challenges of real-world endodontic practice, mainly due to a lack of endodontic domain experts who could bridge the gap between AI engineers and the field of endodontics. Consequently, literature becomes populated with low-quality studies with oversold conclusions, but no conceivable clinical applicability. In this paper, the authors critically present the common pitfalls observed in current AI research in dentistry, and more specifically, in endodontics and then offer practical solutions to address these pitfalls in future research. The goal is to provide constructive guidance that can help establish rigorous standards for a field that is transitioning from research to real-life application. While we aim to present ideal scenarios that highlight these challenges, it is essential to acknowledge that every research study has limitations and shortcomings. These observations are not a critique of the researchers' intent but a call for consideration of these factors to enhance future work in the field. Foundational research has been conducted, and basic AI applications have been developed for endodontic applications. The next steps in endodontic AI research will have to address clinical translation and relevance which may often have been overlooked. This pitfall includes two interconnected challenges, but in distinct ways: developing AI models that solve problems of limited clinical significance and conducting studies under conditions that cannot translate to the real world. Some AI studies focus on tasks that, while technically interesting, have limited practical value in endodontic practice. Thus, the fundamental question researchers must ask is: Does this AI application solve a problem that clinicians face in daily practice, and would its solution meaningfully improve patient care or clinical workflow? Without this foundation, even technically sophisticated and highly performant models may risk being solutions in search of problems. Another core problem in AI studies is using non-representative training datasets. Some studies developed AI models under idealised or artificial conditions and then reported results as if they were clinically meaningful. This gap between experimental success and clinical utility may lead to a false sense of success. It impairs translation to practice, as models that perform well in controlled settings can fail or mislead in real clinical settings (Ibrahim et al. 2021). For example, there are AI models developed on ex vivo samples, phantom images, or synthetic entities (e.g., artificial lesions or artificial root cracks) that simplify the task. Training and testing an AI model on the detection of artificial root cracks in radiographic images of extracted teeth does not imply that the model can detect vertical root fractures in radiographic images of real patients. While convenient for initial experimentation, like pilot or proof-of-concept studies, such data lack the complexity and variable features of in vivo conditions. However, problems arise when researchers overinterpret their results or fail to acknowledge the fundamental limitations. Since these models often overestimate the capabilities of AI, we recommend that researchers test and validate their models on real clinical samples to verify their performance. Specifically, CONSORT-AI, a guideline for reporting AI clinical trials, explicitly noted that AI interventions should be evaluated in their intended clinical setting with patient-centric outcomes, not just technical metrics (Liu et al. 2020). Moreover, the American Dental Association (ADA) standard for dental AI validation (ANSI/ADA 1110-1) emphasises that AI radiology products must be tested on data that accurately represent the intended clinical cohort (2025). The validity of a supervised learning artificial intelligence system in endodontics is ultimately based on the reliability of its annotations. Labels provide the ground truth that supervised models rely on during training and evaluation. If the labels are poorly defined or inconsistently applied, that flawed representation of the dataset embeds itself into the model and can limit its performance. A frequent shortcoming is the reliance on a single annotator, with a lack of reporting on consensus or assessment of intra- or inter-rater reliability. Without such measures, the degree of subjectivity inherent in identifying features such as periapical radiolucencies or resorptive defects cannot be quantified. Even with measuring agreement, dentists only show moderate agreement with tasks such as caries identification on radiographs. These cases make establishing reliable reference annotations difficult (Devlin et al. 2021). Multiple annotations with experts who have differing judgements create fuzzy labels that can bias AI models if combined through simple majority voting (Zheng et al. 2021). Deep learning models are highly sensitive to noisy labels, and their performance drops as label noise increases (Karimi et al. 2020). There are especially strong effects when errors occur in minority classes with fewer examples, such as odontogenic tumours. Training with multiple annotators exposes inconsistency. The same image often receives conflicting labels, which shows that disagreement is not random but systematically influenced by lesion subtlety, annotator experience, and bias (Büttner et al. 2023). Probabilistic aggregation methods such as multi-annotator competence estimation (MACE) and the Dawid–Skene method better estimate true labels by modelling annotator reliability and bias compared to simple consensus (Venanzi et al. 2014). A recent study that tested various annotation aggregation strategies found that MACE worked best for single-modality datasets (e.g., radiographs), while Dawid–Skene excelled in multimodal combined datasets (e.g., images and text) (Klein et al. 2025). This highlights the importance of structured training and calibration of annotators before starting large-scale labelling. Another recurring weakness lies in the absence of detailed annotation protocols. When definitions of pathological findings, outcome measures, or anatomical boundaries are vague, annotators may apply inconsistent criteria, producing noisy labels that are very variable, heterogeneous, and difficult to replicate. AI papers often fail to describe how ambiguous cases, such as faint bone discontinuities or borderline artefact lines, were handled. This is particularly troubling in endodontics, where anatomical structures such as root apices, the location of apical foramina, and canal boundaries are subtle and sometimes difficult to accurately define. Similar concerns have been raised in reviews of dental AI research, which emphasise the importance of transparent and standardised annotation methods to ensure reproducibility and clinical trustworthiness (Schwendicke et al. 2021; Uribe et al. 2025). The problem becomes even more complicated when the expertise and calibration of annotators are not reported. This causes uncertainty as to whether images were assessed by highly trained specialists, junior clinicians, residents or students. Finally, the data source is often under-documented. This refers to systematic documentation of how data, especially imaging data, were acquired and processed, including scanner/imaging machine specifications, exposure settings, or artefact prevalence are rarely described in detail. Each of these factors may influence the appearance of endodontic features and thereby bias the annotations (Norgeot et al. 2020; Liu et al. 2020). The cumulative effect of these deficiencies constrains the achievable accuracy of the model and lowers confidence in its clinical applicability. Recent standards, such as International Organisation for Standardisation (ISO) 18374:2025 (2025), the first international standard specifically governing AI in dentistry, provides requirements for data generation, annotation, and processing in AI-based dental radiograph analysis to enhance reproducibility and trustworthiness. Moreover, the ADA standard for dental AI (ANSI/ADA 1110-1) underscores that annotation should follow standardised definitions and record metadata about the annotator experience (2025). Even though some studies don't adopt this standard, its existence signals the community's recognition that annotation is a first-order problem that needs to be addressed. Future directions include smarter label aggregation, semi-supervised learning, and reference modalities (e.g., micro-CT) to improve annotation quality and reduce human inconsistency (Klein et al. 2025). Various studies have tried to suggest new models specific to endodontic tasks. However, they do not benchmark and compare the “novel” method against state-of-the-art models. Without comprehensive comparisons, there is a possibility of presenting new, complex models as superior when simpler or more established approaches might yield similar results. This may mislead the scientific community, as these models may be prematurely adopted despite offering little or no real improvement, ultimately hindering scientific progress (Isensee et al. 2024). Similarly, Bassi et al. (2024) suggested that “unfair comparisons” are one of the challenges in studies to propose a new AI model, which leads to biases and should be treated cautiously in real-world situations. Additionally, some studies introduced minor tweaks as major innovations (e.g., swapping out a model component for another well-established one or applying a simple post-processing step), and prioritise algorithmic sophistication over practical applicability. This practice may mislead and promote a culture of exaggerated claims and false novelty under the guise of advanced AI development. For example, Schneider et al. (2022) showed that more complex models did not necessarily correlate with better performance in their teeth segmentation benchmarking study. In the 3D medical image segmentation task, Isensee et al. (2024) demonstrated in their benchmarking study in 2024 that their original model, introduced in 2021 (Isensee et al. 2021), still outperformed many more recently popularised models in the field. The authors suggested that there was a bias toward new AI architectures, which indicates a need for stricter validation standards. Moreover, when comparative models are applied to clinical tasks, the selection often appears arbitrary (Schneider et al. 2022). Various studies reported unfavourable outcomes for comparative models, while using slightly modified models (such as architecture, learning pipelines, and initialisation), better outcomes could be achieved. The significance of these limitations is the potential to underestimate the true capabilities of AI models. For instance, some studies may report that humans outperform AI in specific tasks based on comparisons with a single, potentially suboptimal model configuration. However, without systematic ablation studies to explore alternative architectures, hyperparameters, or training strategies, such conclusions may underestimate AI's true capabilities. Therefore, a systematic, comprehensive, and “hypothesis-driven” approach for the selection of AI models and their configurations was suggested (Schneider et al. 2022; Schwendicke et al. 2021). The reasons for choosing the selected approach or model should be justified (Weikert et al. 2021). In general, the selection of AI model types should be driven by factors including, but not limited to, the dataset size and distribution, complexity of the task, interpretability, and computational and validation costs (De Hond et al. 2022). Another critical challenge in model development is the claim that AI models are explainable because of post hoc visualisation tools such as saliency maps. While these techniques can offer insights into model decision-making, they do not achieve true explainability by design. True explainability requires that the model architecture or learning process inherently produces interpretable outputs or structured reasoning, rather than relying on post hoc processing (Rudin 2019). Overstating the level of explainability can be misleading to reviewers and end-users and compromise trust. Beyond data preparation, the credibility of endodontic AI research rests on the strength of its validation and analysis. A recurring methodological weakness is the practice of reporting a single result from a single model run without accounting for natural variation in sample data. Such single-point estimates risk overstating the model's capability. It is common for studies to report apparently impressive performance metrics such as accuracies exceeding 95%. Such results may be misleading if derived from a single validation set. To mitigate this instability, more rigorous approaches such as cross-validation should be employed (Bradshaw et al. 2023). This ensures that each case contributes to both training and testing, which provides a distribution of performance estimates rather than a single point value, and reduces the risk that results are driven by a favourable split. When cross-validation results are reported, it is often unclear which model is ultimately presented in the publication. In many cases, authors appear to have selected the “best” fold based on validation performance. This information is not always clearly stated and could lead to a misrepresentation of the cross-validation process. The lack of transparency of the cross-validation process further introduces the possibility of the lack of true testing on an unseen dataset. This necessitates the need for test data to be split from the dataset (i.e., held-out test set) before cross-validation is carried out. Inappropriate data partitioning is common, with images or slices from the same patient appearing in both training and test sets. In endodontic imaging, a scenario could arise where cone-beam computed tomography (CBCT) volumes from a single patient are divided into multiple slices and then distributed across both training and test sets. This is called ‘data leakage’ and can inflate the model's performance by allowing it to memorise patient-specific feature patterns rather than learn generalisable features (Apicella et al. 2025). Such practices result in overfitting, which may yield deceptively high-performance reporting, despite notable reductions in performance when the model is tested on new, unseen patient data or images (Aliferis and Simon 2024). Methodological assessments of medical imaging AI studies have repeatedly identified “metrics as a major source of bias et al. 2021; Liu et al. 2020). In many accuracy or under the is even in the of or when the clinical question is more the or level (Büttner et al. 2024). This inconsistency in reporting metrics it to compare model performance across This could result from a lack of in or a to that have been in The variable of AI models based on the from to even their selection of performance For instance, metrics such as and are for segmentation tasks where the model is and an in the same an or is more for a where the model the of the image et al. 2024). the reporting of such as accuracy for types of dental AI tasks, is in the is comparative claims are sometimes without confidence or testing, the of validation on datasets is often A lack of testing across imaging and may make the reported model appear highly a single or fail when with the of real clinical (Liu et al. 2019). these practices risk producing claims and models that are poorly for translation into real-world clinical reviews have reported that dental AI studies lack reporting (Mohammad-Rahimi et al. 2022; et al. 2024). There are methodological and inconsistent reported outcomes in these In practice, common reporting can include reporting of data and image settings, the model architectures, and limited reporting of performance The methodology and results must be reported in a and to ensure medical and scientific et al. 2024). To standards for reporting the of studies, it is highly that studies to reporting et al. 2024; et al. 2020; and 2022). The AI papers in the field of dentistry, including endodontics, often fail to follow established reporting for AI in it difficult to the validity and relevance of (Mohammad-Rahimi et al. 2022; and 2022). A comprehensive of is by Uribe et al. and et al. Moreover, Schwendicke et al. a for and reporting of AI in (Norgeot et al. is the suggested guideline for AI in medical image analysis. The of et al. in is a This was specifically for diagnostic accuracy studies with dataset algorithmic and For clinical AI is the reporting guideline (Liu et al. 2020). It is that the of an for AI it should be to the editors and reviewers that to these needs to be and that this refers to the of information in the paper, not the quality of the work Moreover, an of AI research is the for results to be or which requires essential such as or trained et al. 2022; Uribe et al. 2025). However, studies do not provide to or trained models and fail to to et al. 2025). In even when an is in outputs can arise from that is without transparency across the as can performance by biases et al. 2024). Another critical but common is not the between researchers and Such can be and can lead to more but they can result in significant of both in of study and reporting of outcomes, which must be to report these scientific In cases of AI it is the that the model based on data and labels by the If this is the model testing should be and reported by the or by an these methodological limitations the challenges in the design, and reporting of AI studies in endodontics. However, in to these core there are that as the field to clinically applications. AI research is the of basic research and of studies toward and the development and testing of clinical applications et al. 2024). must be to ensure that the development of endodontic AI research in a that is both by scientific validity and clinical These the focus from identifying technical to is for translation from to practice. First, the question of how an AI system would a clinical requires performance on even if and does not establish whether an application will meaningfully or in real clinical should how clinicians with the the degree to which it the and of the clinicians, practical to in practice settings, outcomes, and Without this the relevance of an model remains must be addressed. While guidance is by ADA and standards, clinical ultimately on (e.g., or and by which are and requirements for are such as for models, and are to the of medical AI and will have to be of a when an AI tool is for clinical there is a need for dataset and datasets may suffer from limitations such as and datasets with must be between imaging and variation of can significantly influence model performance. Therefore, between or the of data have to be et al. that datasets with and transparent documentation can provide reliable and improve while pitfalls is researchers should more guidance on a rigorous for the next of AI studies in endodontics. A structured that clinical problem data generation, annotation model validation and reporting is to establish common standards. as AI research there is a need for research and dental to a research that where AI may provide clinical value and methodological across the of model reliability should performance should be for calibration the quality of and the approaches to uncertainty as these may ultimately patient and clinical when AI models are to in ambiguous or the of these definitions would claims clinical applicability. a systematic assessment of bias should be into study and imaging and case can each model and performance should be reported to the of and to explicitly highlight these that the scientific and trustworthiness of future work can be The this from specific methodological pitfalls to the of dataset model and the need for a more structured and clinically approach to AI research in endodontics. The challenge researchers face this new is to bridge the gap between and dental A and between the two that AI engineers and clinical can help to address the pitfalls in this However, the lies with editors and reviewers, that future work is to a Future research should follow and structured strategies to in the and to address clinical data and annotation, validation of and reporting of methods to ensure data original and data original and and and original and data and The authors have to The authors no of The data that the of this study are from the