Kejun Li, Beilin Fu, Yinuo Chen, Ziteng Zhu, Daohui Wang, Xiangting Ge
Background Online video platforms such as YouTube and Bilibili have become important sources of health information for the public. However, the quality and reliability of online health education videos vary substantially. Artificial intelligence (AI), particularly large language models (LLMs), has recently shown potential for automated evaluation of health information, yet evidence regarding the agreement between AI-based and expert evaluations remains limited. Methods A cross-sectional study was conducted to analyze asthma-related health education videos retrieved from YouTube and Bilibili. Two medical experts independently evaluated video quality using the Global Quality Scale (GQS) and modified DISCERN (mDISCERN). AI-assisted evaluation was performed using the GPT-4 large language model based on video transcripts under standardized prompts. Interrater agreement between experts was assessed using intraclass correlation coefficients (ICCs) and Spearman correlations. Agreement between AI and expert ratings was evaluated using Spearman correlation and Bland–Altman analysis. Multivariable linear regression was performed to identify factors associated with differences between AI and expert scores. Results A total of 200 asthma-related health education videos were included, with 100 videos from each platform. AI ratings showed moderate correlations with expert ratings for both GQS ( ρ = 0.55, p < 0.001) and mDISCERN ( ρ = 0.49, p < 0.001). AI scores were slightly higher than expert scores, with mean differences of 0.24 for GQS and 0.60 for mDISCERN. Multivariable regression analysis identified platform (YouTube), source type (organizational source), and PEMAT actionability as significant factors associated with differences in mDISCERN scores, whereas no significant predictors were identified for GQS score differences. Conclusion AI-generated evaluations demonstrated moderate positive correlations with expert assessments in evaluating the quality of asthma-related health education videos. AI may serve as a scalable tool for preliminary screening of online health information, although systematic differences remain across platforms and content characteristics. Expert evaluation remains essential to ensure the accuracy and reliability of medical information assessment.