科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Medical Image Analysis2025-11-17· Foundation (evidence)

Scaling up self-supervised learning for improved surgical foundation models

Tim J. M. Jaspers, Ronald L. P. D. de Jong, Yiping Li, Carolus H. J. Kusters, Franciscus H. A. Bakker, Romy C. van Jaarsveld, Gino M. Kuiper, Richard van Hillegersberg, Jelle P. Ruurda, Willem M. Brinkman, Josien P. W. Pluim, Peter H. N. de With, Marcel Breeuwer, Yasmina Al Khalil, Fons van der Sommen

原始摘要(英文原文)· Original abstract
• Demonstration of effectiveness of SSL for surgical computer vision using the largest dataset reported to date. • Strong generalization and robust evaluation are shown across six surgical datasets, four procedures, and three tasks, outperforming current SOTA foundation models. • Providing insights into large-scale SSL for surgical computer vision in terms of scaling, pretraining time, dataset composition, and model architecture. • Release of the models and a curated dataset of 2.1 million surgical video frames, establishing a critical resource for advancing surgical foundation model training Foundation models have revolutionized computer vision by achieving vastly superior performance across diverse tasks through large-scale pretraining on extensive datasets. However, their application in surgical computer vision has been limited. This study addresses this gap by introducing SurgeNetXL, a novel surgical foundation model that sets a new benchmark in surgical computer vision. Trained on the largest reported surgical dataset to date, comprising over 4.7 million video frames, SurgeNetXL achieves consistent top-tier performance across six datasets spanning four surgical procedures and three tasks, including semantic segmentation, surgical phase recognition, and critical view of safety (CVS) classification. Compared with the best-performing surgical foundation model, SurgeNetXL shows mean improvements of 4.0%, 8.9%, and 11.4% for semantic segmentation, phase recognition, and CVS classification, respectively. Additionally, SurgeNetXL outperforms ImageNet1k by 16.1%, 8.0%, and 4.3% for the respective tasks. In addition to advancing model performance, this study provides key insights into scaling pretraining datasets, extending training durations, and optimizing model architectures specifically for surgical computer vision. These findings pave the way for improved generalization and robustness in data-scarce scenarios, offering a comprehensive framework for future research in this domain. All models and a subset of the SurgeNetXL dataset, including over 2 million video frames, are publicly available at: https://github.com/TimJaspers0801/SurgeNet .
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Scaling up self-supervised learning for improved surgical foundation models — 科研速览 Science Skim