Yikang Zhang, Yanan Zhou, Yuwen Li, Zhimin Zhang
Automated heart sound auscultation is crucial for early cardiovascular disease screening, yet existing multi-task learning approaches face challenges in cross-modal feature fusion, signal quality coupling, and heterogeneous data optimization. This paper proposes a robust multi-feature fusion framework that simultaneously performs murmur detection, clinical outcome prediction, and signal quality assessment. We extracted raw waveforms, Mel-frequency cepstral coefficients (MFCCs), and statistical features and deeply integrated them via a bidirectional cross-modal attention mechanism. Crucially, to enable joint training on heterogeneous datasets with missing labels, we introduced a task masking mechanism alongside a homoscedastic uncertainty-based dynamic weighting strategy to resolve multi-task loss conflicts. When evaluated on the CirCor DigiScope and a multi-source signal quality dataset, the proposed framework demonstrated highly competitive performance. It achieved a murmur detection weighted accuracy of 0.766±0.026 and reduced the clinical screening cost to 10490±491, which was far below the PhysioNet Challenge 2022 average. Concurrently, the model yielded a cost of 10553±1134 for clinical outcome prediction and secured a macro-F1 score of 0.845±0.015 for signal quality assessment. By explicitly co-optimizing diagnostic tasks and signal quality, our method provides an efficient, low-cost solution tailored for primary healthcare and noisy real-world clinical settings.