Yi-Ching Wang, Ya-Chen Lee, Yu-Jeng Ju, Tzu-Ting Chen, Sheau-Ling Huang, Chih-Wei Yang, Ching-Lin Hsieh
Both the ChatGPT-rated and expert-rated GKCSAF showed evidence of responsiveness for detecting changes in clinicians' communication skills.
OBJECTIVE: The Gap-Kalamazoo Communication Skills Assessment Form (GKCSAF) is widely used to assess clinicians' communication skills. While expert-rated administration of the GKCSAF is seen as objective and thorough, manual scoring is time- and resource-intensive. The ChatGPT-rated GKCSAF is a large language model-based rating approach that uses the same GKCSAF items and scoring criteria. It shows acceptable reliability and validity while reducing rater burden. However, the responsiveness of the expert-rated and ChatGPT-rated GKCSAF has remained unknown, limiting the interpretation of change scores in communication skills. The study aimed to compare the responsiveness of the expert-rated and ChatGPT-rated GKCSAF.
METHODS: Eighty occupational therapy students completed two recorded clinical interactions with different patients. Between interactions, students received training via the Communication Skills Measure for Therapists, a formative measure that assesses communication skills and yields actionable feedback. Transcripts of the interactions were independently scored by trained expert raters and by ChatGPT. Responsiveness was examined using paired t-tests, Cohen's d, and standardized response mean (SRM). Differences in responsiveness indices between the rating approaches (i.e., the expert-rated and ChatGPT-rated GKCSAF) were estimated using bootstrap confidence intervals.
RESULTS: For both rating approaches, mean post-training scores were significantly higher than mean pre-training scores (both t = 3.8, p < 0.001). The responsiveness indices were nearly moderate (expert-rated GKCSAF: d = 0.45, SRM = 0.42; ChatGPT-rated GKCSAF: d = 0.47, SRM = 0.43), with no significant differences between the two rating approaches (bootstrap 95% confidence intervals for the differences included zero).
CONCLUSIONS: Both the ChatGPT-rated and expert-rated GKCSAF showed evidence of responsiveness for detecting changes in clinicians' communication skills.
PRACTICE IMPLICATIONS: Considering the shorter scoring time and reduced rater burden of the ChatGPT-rated GKCSAF, it may be a promising outcome assessment for clinical and research use.