Yaqiong Zhang, Kai Zhang, Meijia Wang, Peng Bai, Shengqiang Wang, Xuanye Hu, Ting Ma, Feng Hu, Peng Li, Guisheng Liu
Our results demonstrate that the SimCLR framework enables high-performance quality control in capsule endoscopy, achieving accuracy comparable to - or even surpassing - supervised learning methods, while requiring substantially fewer annotated images.
BACKGROUND: It is possible to control the quality of capsule endoscopic images using artificial intelligence, but it requires a great deal of time for labeling. Verifying the performance of self-supervised learning for quality control for capsule endoscopic images.
METHODS: A simple framework for contrastive learning of visual representations (SimCLR), is capable of acquiring the inherent image representation with minimal annotation, but the feasibility is not studied. A total of 62 840 images were collected to train models in internal cross-validation (more training data and less testing data) and reversed cross-validation (less training data and more testing data). Random forest and eXtreme Gradient Boosting (XGBoost) were used to complete the quality control after SimCLR extracted the features from images.
RESULTS: Random forest and XGBoost reported that the mean area under the receiver operating characteristic (AUROC) curve exceeded 0.98 and 0.97 using SimCLR-derived features. Moreover, XGBoost surpassed supervised convolutional neural network (CNN). Extra 12 032 images were gathered for prospective validation and the AUROC of SimCLR surpassed 0.93 (95% confidence interval = 0.9271-0.9548), which is close to supervised CNN (0.9645) in cross-validation. Moreover, the AUROC of random forest and XGBoost (trained with SimCLR-derived features) surpasses 0.96, which is better than supervised CNN (0.8374) in reversed cross validation.
CONCLUSION: Our results demonstrate that the SimCLR framework enables high-performance quality control in capsule endoscopy, achieving accuracy comparable to - or even surpassing - supervised learning methods, while requiring substantially fewer annotated images.