Asifa Tassaddiq, Aiman Albarakati, Rabab Alharbi, Carlo Cattani, Dalal Khalid Almutairi, Ruhaila Md Kasmani
We then evaluated the frozen 1.33-million-parameter InceptionTime-CNN-BiGRU-Transformer model and validation-derived thresholds on 15,931 Ningbo ECGs without retraining, recalibration, or external threshold adjustment.
Background: Automated interpretation of 12-lead electrocardiograms (ECGs) remains challenging because multiple abnormalities may coexist and appear in selected leads or brief waveform segments. We developed a compact and interpretable framework for five-superclass multi-label ECG diagnosis. Methods: We evaluated PTB-XL records using the official fold protocol, with folds 1-8 for training, fold 9 for validation monitoring and class-specific threshold selection, and fold 10 for independent internal testing. We then evaluated the frozen 1.33-million-parameter InceptionTime-CNN-BiGRU-Transformer model and validation-derived thresholds on 15,931 Ningbo ECGs without retraining, recalibration, or external threshold adjustment. We also examined calibration, demographic subgroups, computational efficiency, and complementary ECG-domain attribution methods. Results: Macro-AUROC reached 90.83% on PTB-XL fold 10 and 88.85% on Ningbo, indicating generally consistent diagnostic ranking with a modest reduction during external evaluation. Sensitivity analysis showed that CNN-only outperformed the frozen primary model on five of six endpoints. Attribution analyses highlighted qualitatively plausible lead and temporal patterns in representative examples, while calibration and subgroup analyses further characterized model behavior under dataset shift. Conclusions: Our framework integrates leakage-aware development, threshold-controlled testing, frozen external validation, and multimethod interpretability. These findings support its further prospective, locally calibrated evaluation as a potential aid for multi-label ECG interpretation.