Mattia Nigro, Andrea Aliverti, Alessandra Angelucci, Apostolos Bossios, Hilary Pinnock, Jeanette Boyd, Pippa Powell, Michal Shteinberg, James D Chalmers, Stefano Aliberti, AIR-BE Task Force
AI software, particularly ChatGPT, provided accurate, comprehensive, and understandable answers to our 15 bronchiectasis-related questions.
BACKGROUND: The popularity of Generative Artificial Intelligence (AI)-powered chatbots is growing, with more patients using AI to understand respiratory conditions, including bronchiectasis. The quality of AI's responses to bronchiectasis-related questions has not previously been evaluated.
OBJECTIVES: To evaluate the reliability, accuracy, comprehensiveness, and understandability of responses generated by different AI-based chatbots to bronchiectasis-related questions formulated by patients.
DESIGN: Cross-sectional international study.
METHODS: People living with bronchiectasis from the EMBARC/European Lung Foundation (ELF) patient advisory group formulated 15 bronchiectasis-related questions, categorised into 3 difficulty tiers. These questions were submitted to three AI-based chatbots: ChatGPT, Bard, and Copilot. An international group of 28 experts from the European Respiratory Societies (recruited through Assemblies and the CONNECT clinical research collaboration) and 33 patients evaluated the answers' reliability, accuracy, comprehensiveness, and understandability.
RESULTS: Of 45 outcomes, 37 were deemed reliable. Median accuracy scores ranged from 6.0 to 9.0, with ChatGPT scoring higher than Copilot and Bard (8.0 [7.0-9.0] vs 7.0 [6.0-8.0] vs 7.0 [6.0-8.0], p-value < 0.001). Median comprehensiveness scores ranged from 6.0 to 9.0, with ChatGPT providing more comprehensive answers compared to Bard and Copilot (8.0 [7.0-9.0] vs 8.0 [6.0-9.0] vs 7.0 [6.0-8.0], p-value < 0.001). Median understandability scores ranged from 7.0 to 10.0, with ChatGPT and Bard delivering more understandable answers than Copilot (9.0 [8.0-10.0] vs 9.0 [8.0-10.0] vs 9.0 [7.0-10.0], p-value < 0.001). Differences in understandability were noted in Bard's responses across difficulty tiers (easy 9.0 [8.0-10.0] vs medium 9.0 [7.0-10.0] vs difficult 8.0 [7.0-10.0], p-value = 0.012).
CONCLUSION: AI software, particularly ChatGPT, provided accurate, comprehensive, and understandable answers to our 15 bronchiectasis-related questions.