Federico Leone, Filippo Omenetti, Alessandro Bianchi, Giannicola Iannella, Andrea Collettini, Saverio Nicoletti, Maria Elvira Distefano, Yuliya Germini Moysey, Edoardo Rosciano, Giorgio Imbrogno, Andrea Preti, Federica Vultaggio, Franco Minnella, Riccardo Zanzi, Fabrizio Salamanca, Federico Ambrogi, Francesco Mozzanica
DISE interpretation remains complex and operator-dependent. Pattern classification appears more reproducible than grading. While a non-task-specific AI model captured clinically relevant patterns, it did not match expert performance in complex decision-making. AI may serve as a supportive tool, particularly for less experienced clinicians.
PURPOSE: To evaluate the inter-rater reliability of DISE interpretation using the VOTE classification, assess variability in therapeutic decision-making, and explore the potential role of artificial intelligence (AI) in DISE evaluation.
METHODS: Twenty DISE recordings were retrospectively and independently evaluated by 12 raters (7 junior and 5 senior) and by a general-purpose AI model without task-specific training. Inter-rater agreement among human raters for obstruction grading and pattern classification was assessed using percentage agreement, Fleiss' kappa, and weighted kappa. Agreement with a reference standard (unanimous expert consensus) was analyzed. Treatment proposals were evaluated using exact match ratio and Jaccard similarity.
RESULTS: At baseline, agreement among human raters ranged from 43.6% to 68.0% for obstruction grading (weighted κ, 0.255-0.425) and from 60.8% to 79.3% for collapse pattern (Fleiss' κ, 0.287-0.370). Senior raters showed higher baseline agreement than junior raters across all anatomical levels for collapse pattern and at three of four levels for obstruction grading. Agreement with the expert consensus ranged from 70.0% to 90.0% for senior raters and from 35.0% to 85.0% for AI in baseline grading, and from 75.0% to 95.0% and 50.0% to 80.0%, respectively, in baseline pattern classification. For therapeutic recommendations, exact agreement was 70.0% for senior raters, 55.0% for junior raters, and 50.0% for AI; mean Jaccard similarity was 0.867, 0.833, and 0.742, respectively.
CONCLUSIONS: DISE interpretation remains complex and operator-dependent. Pattern classification appears more reproducible than grading. While a non-task-specific AI model captured clinically relevant patterns, it did not match expert performance in complex decision-making. AI may serve as a supportive tool, particularly for less experienced clinicians.