科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ JMIR medical education2026-09-24

Evaluating Large Language Model-Based Automated Scoring in a Voice-Based Virtual Standardized Patient Platform for Medical Students: Cross-Sectional Agreement Study.

Xiaoxing Gao, Xiaoming Huang, Rongrong Hu, Li Zhang, Huiting Liu, Bingqing Zhang, Chong Wei, Wei Qiu, Mengyu Zhang, Xuefeng Sun

一句话结论 · In one sentence

LLM-based scoring in a VSP showed moderate agreement with faculty ratings, performing better for information gathering than for communication. Due to rater and case heterogeneity, ceiling effects, and proportional bias, this method is suitable for formative use and enhanced sampling in programmatic assessment but not for independent, high-stakes summative decisions.

原始摘要(英文原文)· Original abstract
BACKGROUND: Large language model (LLM)-powered virtual standardized patients (VSPs) enable scalable clinical skills practice, but the validity of AI-generated scores relative to faculty ratings remains unclear. OBJECTIVE: This study aimed to assess agreement between LLM-generated and faculty ratings of history-taking and communication performance and to examine the influence of rater and case heterogeneity. METHODS: In this cross-sectional study, 92 fourth-year medical students completed one of three 15-minute voice-based VSP cases (fever, diarrhea, and cough). Ten blinded faculty raters scored performance (0-100 points total; 0-50 points per domain). AI scores were generated by DeepSeek-V3 using a calibrated prompt. Agreement was evaluated using mixed-effects models, intraclass correlation coefficients (ICC [2,1]), Spearman correlations, mean absolute error (MAE), Bland-Altman analysis, and variance partition coefficients (VPC). RESULTS: Median total scores were similar for AI and faculty (median 93.0, IQR 89.0-95.0 vs median 94.0, IQR 91.0-95.0). Rater variability accounted for 37% of residual variance in faculty total scores (VPC=0.37). AI total scores were positively associated with faculty total scores (β=0.37, 95% CI 0.26-0.48; P<.001; Spearman ρ=0.50, 95% CI 0.34-0.65). Absolute agreement was moderate (ICC[2,1]=0.51, 95% CI 0.34-0.65), with MAE of 3.11 points. Mixed-effects Bland-Altman analysis showed a small, not statistically significant mean bias (1.26 points, 95% CI -0.48 to 3.01; P=.16) and 95% limits of agreement from -4.95 to 7.48 (width=12.43 points), with proportional bias (β_proportional bias=-0.55; P<.001). Agreement was stronger for information gathering (β=0.46; ρ=0.49; ICC=0.54; VPC=0.23) than for communication (β=0.27; ρ=0.28; ICC=0.29; VPC=0.52). A sensitivity analysis in the lowest quartile showed attenuated but consistent agreement (ICC=0.38). CONCLUSIONS: LLM-based scoring in a VSP showed moderate agreement with faculty ratings, performing better for information gathering than for communication. Due to rater and case heterogeneity, ceiling effects, and proportional bias, this method is suitable for formative use and enhanced sampling in programmatic assessment but not for independent, high-stakes summative decisions.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Evaluating Large Language Model-Based Automated Scoring in a Voice-Based Virtual Standardized Patient Platform for Medical Students: Cross-Sectional Agreement Study. — 科研速览 Science Skim