科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ European archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : affiliated with the German Society for Oto-Rhino-Laryngology - Head and Neck Surgery2026-09-16

Turing problems in otolaryngology: a scoping review of the principal challenges of artificial ıntelligence and large language models.

Nidanur Sinanoglu, Azada Ismayilova, Antiga Muradova, Rovsahana Hajiyeva, Natig Ahmadov, Yusif Hajiyev, Aynur Aliyeva

一句话结论 · In one sentence

Current AI and LLM tools in OHNS demonstrate promising but insufficient accuracy for unsupervised clinical deployment. Structured governance frameworks, mandatory clinical validation pipelines, and bias-audited datasets are urgently required.

原始摘要(英文原文)· Original abstract
BACKGROUND: Artificial intelligence (AI) and large language models (LLMs) have rapidly entered otolaryngology-head and neck surgery (OHNS). Despite accelerating publication output, critical translational and safety challenges remain undercharacterized. METHODS: A scoping review was conducted, and PubMed/MEDLINE, Cochrane Library, Embase, Web of Science, and Scopus were searched from January 2020 to June 2025 using pre-specified terms encompassing AI, LLMs, machine learning, and deep learning in OHNS. RESULTS: Of 3,648 screened records, 68 met the final inclusion criteria. Six principal challenge domains were identified: (1) accuracy and validity (LLM correct-answer rates: 53-75% across studies); (2) hallucination and reference fabrication, including a 61.6% reference-to-prompt irrelevancy rate across the platforms evaluated in one study; (3) the 'AI Chasm' translational gap (99.3% of deep-learning studies remained in silico); (4) Black-Box/explainability failure; (5) algorithmic bias and demographic disparities; and (6) data privacy, regulatory compliance, and legal accountability. GPT-4-class models consistently outperformed GPT-3.5, and domain-specific models (e.g., ChatENT) achieved error reductions of 26-58%. CONCLUSIONS: Current AI and LLM tools in OHNS demonstrate promising but insufficient accuracy for unsupervised clinical deployment. Structured governance frameworks, mandatory clinical validation pipelines, and bias-audited datasets are urgently required.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Turing problems in otolaryngology: a scoping review of the principal challenges of artificial ıntelligence and large language models. — 科研速览 Science Skim