Patrycja Szczepańska-Ciszewska, Wojciech Michał Glinkowski, Andrzej Śliwczyński, Magdalena Głowacka, Anna Erkiert-Polguj, Agata Kmita, Aleksandra Rybak, Barbara Algiert-Zielińska, Maria Koserczyk
Background/Objectives: Clinical cellulite grading relies on inspection and palpation and is susceptible to examiner variability. A single experienced assessor previously evaluated a five-grade infrared thermographic classification, but its reproducibility across raters was unknown. To assess inter-rater and intra-rater reliability of a standardized thermographic grading workflow applied by trained novice raters. Methods: This secondary reliability study included 78 anterior FLIR P620 scans. Five raters analyzed each scan twice, one week apart, using standardized FLIR Tools settings and a ΔT-based 0-IV scale. The prespecified principal agreement analysis used Gwet's AC2 with quadratic weights across the full 0-IV scale; a sensitivity analysis recalculated AC2 using only the observed categories. Confidence intervals were estimated by scan-level bootstrap. Results: The scans yielded 780 ratings, all within grades 0-II. Full five-rater exact agreement was 26.9% and 33.3% in rounds 1 and 2, respectively, while within-one-grade agreement was 89.7% and 92.3%. Observed-category AC2, which restricts quadratic weighting to the grades represented in the data (0-II), was 0.786 (95% CI 0.730-0.840) and 0.845 (0.794-0.885). The corresponding prespecified full-scale AC2 values calculated across the theoretical 0-IV weighting range were 0.949 (0.936-0.962) and 0.963 (0.951-0.972). Intra-rater exact agreement ranged from 66.7% to 97.4%, with 100% agreement within one grade for all raters. Conclusions: The workflow demonstrated reproducible ordinal grading within the lower severity categories represented in the dataset, although exact five-rater agreement remained limited. The findings do not establish reliability for grades III-IV, diagnostic accuracy, or clinical utility.