Jenny Roth, Karthikeyan Baskaran, Ava Dashti, Petros Moustardas, Antonio Filipe Macedo, Neil Lagali
Independent-image selection gives poorer reproducibility of CNFL and DCD than multiple graders using the same-image sets. The same-image sets yielded excellent agreement, suggesting that studies relying on a single selection of five or fewer images should be interpreted with caution, as a single selection is more prone to rater bias. Because independent-image selection reflects real-world practice, we recommend that future studies use multiple graders, each independently selecting and grading images, with more than five images per participant where feasible.
AIM: The aim of this study was to measure the impact of image selection on quantification of corneal nerve fibre length density (CNFL) and dendritic cell density (DCD) using in-vivo confocal microscopy images.
METHODS: We used images from 80 participants (mean age 51 years; 55 female), averaging 1030 images per participant. CNFL and DCD were quantified using independently selected sets and same-image sets. Inter-grader agreement was computed using 10 or 5 images per participant (patient-level data point). Grader agreement was quantified by intraclass correlation coefficients (ICC).
RESULTS: Using independently selected images, the ICC for CNFL was 0.82 with 10 images, which was significantly higher by 0.07 than with 5 images (p = 0.009). For DCD, the ICC was 0.92 with 10 images and was not significantly higher than with 5 images. Using same-image sets, the ICC for CNFL was 0.94 with both 5 and 10 images. For DCD, the ICC improved by 0.03 when 10 images were used compared with 5 images (p = 0.02).
CONCLUSION: Independent-image selection gives poorer reproducibility of CNFL and DCD than multiple graders using the same-image sets. The same-image sets yielded excellent agreement, suggesting that studies relying on a single selection of five or fewer images should be interpreted with caution, as a single selection is more prone to rater bias. Because independent-image selection reflects real-world practice, we recommend that future studies use multiple graders, each independently selecting and grading images, with more than five images per participant where feasible.