Abdurrahman Koç, Abdullah Enes Ataş, Şebnem Yosunkaya, Hülya Vatansev
Contemporary large language models demonstrate promising yet imperfect capacity for real-world polysomnography report interpretation in obstructive sleep apnea, with performance varying by model and case complexity. These findings support selected models as potential adjunctive decision-support tools under specialist oversight.
PURPOSE: The diagnostic and therapeutic capabilities of large language models in interpreting real-world polysomnography reports for obstructive sleep apnea remain insufficiently characterized, with prior studies limited by small samples, single-model designs, and unidimensional outcome measures. This study compared the clinical performance of nine contemporary large language models in diagnosing obstructive sleep apnea and generating treatment recommendations from polysomnography reports, benchmarked against expert consensus.
METHODS: Two hundred twelve polysomnography records from a university sleep laboratory were retrospectively classified as simple (n = 153) or complex (n = 59) based on AASM criteria. De-identified reports, originally in Turkish, were submitted to nine large language models using an English prompt within a one-week frozen evaluation window. Model outputs were independently scored by two sleep medicine specialists using a four-dimensional rubric encompassing diagnostic accuracy, recommendation quality, safety, and parameter coverage. Generalized estimating equations accounted for within-case clustering across 1,908 model-case assessments.
RESULTS: Fully correct diagnostic rates ranged from 75.0% to 86.3%, with Claude 4.1 Opus, ChatGPT-5, and Gemini 2.5 Pro forming a statistically indistinguishable top tier. All models demonstrated significant performance degradation on complex cases (11.8-25.7 percentage point decline). Safety rates exceeded 88% across all models. Moderate obstructive sleep apnea was the most challenging diagnostic category. Hypoxemia disproportionately impaired diagnostic accuracy in lower-ranked models.
CONCLUSION: Contemporary large language models demonstrate promising yet imperfect capacity for real-world polysomnography report interpretation in obstructive sleep apnea, with performance varying by model and case complexity. These findings support selected models as potential adjunctive decision-support tools under specialist oversight.