Yongchao Huo, Xiaodong Zhang, Yuanyuan Ni, Fei Zhao, Xiangyang Yan
Accurate segmentation of phonocardiogram (PCG) recordings into the first and second heart sounds (S1 and S2) is essential for computer-aided auscultation, but remains challenging because cardiac-cycle timing, event duration, and waveform morphology vary across and within subjects. We formulate heart sound segmentation as a one-dimensional temporal tracking problem and propose an online closed-loop prediction-detection-update framework. S1 and S2 are tracked independently using a state vector that contains a reference-point location and left and right temporal extents. A Kalman filter predicts the next search region, an adaptive correlation filter localises the reference point with confidence-guided recovery, and a Mann-Kendall-Sneyers change-point detector refines the onset and offset. The resulting measurements update the state estimate, while accepted segments update the waveform template. On the PhysioNet 2016 and CirCor DigiScope datasets, the proposed method achieved F1 scores of 98.08% and 96.41%, respectively, exceeding the evaluated baselines. Under the strongest oscillatory interference, the corresponding F1 scores remained 97.67 ± 0.24% and 96.13 ± 0.26%. Ablation experiments confirmed the complementary contributions of temporal prediction, confidence-guided recovery, adaptive waveform matching, covariance updating, and boundary refinement. The proposed framework provides robust and interpretable online PCG segmentation for downstream cardiac analysis.