Nidanur Sinanoglu, Azada Ismayilova, Antiga Muradova, Rovsahana Hajiyeva, Natig Ahmadov, Yusif Hajiyev, Aynur Aliyeva
Current AI and LLM tools in OHNS demonstrate promising but insufficient accuracy for unsupervised clinical deployment. Structured governance frameworks, mandatory clinical validation pipelines, and bias-audited datasets are urgently required.
BACKGROUND: Artificial intelligence (AI) and large language models (LLMs) have rapidly entered otolaryngology-head and neck surgery (OHNS). Despite accelerating publication output, critical translational and safety challenges remain undercharacterized.
METHODS: A scoping review was conducted, and PubMed/MEDLINE, Cochrane Library, Embase, Web of Science, and Scopus were searched from January 2020 to June 2025 using pre-specified terms encompassing AI, LLMs, machine learning, and deep learning in OHNS.
RESULTS: Of 3,648 screened records, 68 met the final inclusion criteria. Six principal challenge domains were identified: (1) accuracy and validity (LLM correct-answer rates: 53-75% across studies); (2) hallucination and reference fabrication, including a 61.6% reference-to-prompt irrelevancy rate across the platforms evaluated in one study; (3) the 'AI Chasm' translational gap (99.3% of deep-learning studies remained in silico); (4) Black-Box/explainability failure; (5) algorithmic bias and demographic disparities; and (6) data privacy, regulatory compliance, and legal accountability. GPT-4-class models consistently outperformed GPT-3.5, and domain-specific models (e.g., ChatENT) achieved error reductions of 26-58%.
CONCLUSIONS: Current AI and LLM tools in OHNS demonstrate promising but insufficient accuracy for unsupervised clinical deployment. Structured governance frameworks, mandatory clinical validation pipelines, and bias-audited datasets are urgently required.