Noé Constans, Mihai Mitrea, Nicole Neumann, Nabil Chakfe
Using a single camera and no physical markers, this framework offers a cost-effective alternative to hardware-heavy simulators, enabling democratized, interpretable feedback.
PURPOSE: The transition toward competency-based medical education requires scalable, objective surgical skill assessment. While sensor-based approaches remain burdensome, video-based deep learning models often lack the interpretability required for formative feedback. We introduce a fully automated, markerless framework for assessing vascular open surgery suturing skills providing meaningful interpretation.
METHODS: Our pipeline leverages 3D hand tracking coupled to an LSTM-attention network for precise temporal segmentation. By merging explicit kinematic metrics with latent deep features, we develop a voting ensemble classifier categorizing surgeons into three skill levels. We further evaluated the impact of ground-truth reliability by comparing models trained on single-assessor versus multi-assessor consensus labels.
RESULTS: The framework achieved 96.4% segmentation accuracy and 72.2 ± 0.5% accuracy (F1 = 0.73) under leave-one-video-out cross-validation (LOVCV). To go deeper in assessing cross-surgeon generalization, we additionally report a strict leave-one-subject-out (LOSO) protocol, under which accuracy decreases to 52.3%. This gap quantifies the surgeon-specific component captured by video-level splits and reframes the system as a tool for longitudinal, per-trainee skill tracking rather than zero-shot scoring of unseen surgeons. Phase-stratified analysis uncovers that experts exhibit significantly higher jerk and velocity, reflecting decisive ballistic motor planning. Training on consensus labels yields a 15-20% performance increase over single-assessor baselines.
CONCLUSION: Using a single camera and no physical markers, this framework offers a cost-effective alternative to hardware-heavy simulators, enabling democratized, interpretable feedback.