Christine P Beltran, Shriya Nallamaddi, Jeffrey A Wilhite, Khemraj Hardowar, Kathleen Hanley, Lisa Altshuler, Sondra R Zabar, Colleen Gillespie
Overall communication scores may be comparable across modalities, but some individual communication items may function differently in virtual versus in-person formats. Such differences may reflect modality-related variation in item proficiency, differences in rater scoring, or both. Threshold analysis offers a nuanced approach to evaluating comparability of OSCE formats by examining whether an observed rating reflects the same level of underlying communication proficiency across modalities.
BACKGROUND: Prior studies comparing virtual and in-person OSCEs have focused mainly on overall performance scores, with less attention to whether specific communication skills function differently across modalities.
OBJECTIVE: To compare the level of communication proficiency required to achieve specific checklist ratings in virtual versus in-person OSCEs.
DESIGN: We identified Internal Medicine case (IM) OSCE cases conducted in both modalities. In each case, standardized patients (SP) rated resident performance in three communication domains (information gathering, relationship development, and patient education), using a behaviorally anchored scale of "not," "partially," or "well done." We compared overall communication scores and rating distributions across modalities and used a graded response Item Response Theory model to estimate item-level thresholds, representing the communication proficiency needed to move from one rating category to the next. Virtual and in-person threshold estimates were compared using Wald tests.
PARTICIPANTS: 126 PGY-1 IM residents participated across six cases between 2019 and 2023 (in-person = 82, virtual = 54).
MAIN MEASURES: Mean communication domain scores, rating distributions, and item-level threshold estimates for each checklist item across modalities.
KEY RESULTS: Most residents demonstrated sufficient skill to receive "partly" or "well" done ratings on most communication items. Overall communication performance appeared broadly similar across modalities. At the item level, one behavior-"using words patient understood and explaining jargon"-showed a statistically significant difference: residents required a lower level of proficiency to receive a "partly done" rating in virtual versus in-person encounters (z = 3.92, p_adj = 0.001).
CONCLUSIONS: Overall communication scores may be comparable across modalities, but some individual communication items may function differently in virtual versus in-person formats. Such differences may reflect modality-related variation in item proficiency, differences in rater scoring, or both. Threshold analysis offers a nuanced approach to evaluating comparability of OSCE formats by examining whether an observed rating reflects the same level of underlying communication proficiency across modalities.