Jessica L Chee-Williams, Thomas J Sitzman, Adriane Baylis, Kelly Nett Cordero, Katherine Dillon, Simone Fischbach, Sara Kinter, Paula Klaiman, Megan Donner, Jamie L Perry
Findings suggest that nasopharyngoscopy ratings made in a clinical setting may be systematically different than ratings made in a research setting; thus, research results that depend on nasopharyngoscopy ratings may not be generalizable to clinical care. This may be caused by a variety of patient-specific and practice-specific factors, in addition to a lack of standardization. Findings suggest that high reliability for nasopharyngoscopy ratings can be achieved if raters receive consistent training and are following the same rating scale definitions.
PURPOSE: This study aimed to evaluate intrarater reliability for nasopharyngoscopy ratings of velopharyngeal closure across two contexts: ratings made in a clinical setting and ratings made in a research setting using standardized rating definitions.
METHOD: Five speech-language pathologists participated in this study. Two assessments of intrarater reliability were performed: between ratings from clinical reports and re-ratings made in a research setting and videos re-rated twice in a research setting. Nasopharyngoscopy ratings analyzed included velopharyngeal closure pattern, total percent closure, and extent of velar movement. Reliability was calculated using Cohen's kappa and the intraclass correlation coefficient (ICC).
RESULTS: Intrarater reliability between clinical reports and re-ratings made in a research setting varied for closure pattern from weak (k = .24) to perfect (k = 1.00), for total percent closure from poor (ICC = -.09) to good (ICC = .81), and for extent of velar movement from moderate (ICC = .57) to good (ICC = .80). Compared to intrarater reliability between clinical and research settings, intrarater reliability for videos re-rated twice in a research setting was higher for all raters. For closure pattern, intrarater reliability within a research setting ranged from substantial (k = .68) to perfect (k = 1.00), for total percent closure from moderate (ICC = .70) to excellent (ICC = .99), and for extent of velar movement from good (ICC = .80) to excellent (ICC = .91).
CONCLUSIONS: Findings suggest that nasopharyngoscopy ratings made in a clinical setting may be systematically different than ratings made in a research setting; thus, research results that depend on nasopharyngoscopy ratings may not be generalizable to clinical care. This may be caused by a variety of patient-specific and practice-specific factors, in addition to a lack of standardization. Findings suggest that high reliability for nasopharyngoscopy ratings can be achieved if raters receive consistent training and are following the same rating scale definitions.